In short
Denoised Podcast Episode Notes: ComfyUI Explained: How AI Image Generation Actually Works (Step-by-Step)
Episode Overview Hosts: Addy Ghani (Media Industry Analyst) & Joey Daoud (Media Producer and Founder of VP Land) Release Schedule: Twice-weekly (Tuesdays and Fridays) Episode Theme: This episode dives into the mechanics of AI image generation, focusing on ComfyUI, discussing core concepts such as latent space, diffusion models, and noise reduction in image recognition.
---
Key Concepts and Discussions
- AI Image Generation Basics
- Neural Networks: Central to image generation and computer vision, advanced neural networks use billions of nodes to understand and generate images based on textual input.
- Mathematical Operation: The basic function of a neural network node can be simplified to the formula: `x * y + b` (linear algebra).
- Training Models: Models like MidJourney and Stable Diffusion are trained with millions of images, alongside captions for accurate recognition and generation.
- Understanding Latent Space
- Definition: Latent space is a multidimensional mathematical representation where concepts and images are clustered together (e.g., beaches, urban landscapes).
- Operation: When a text prompt is given, the model accesses relevant sections of latent space to form an image based on learned associations.
- Noise in Image Generation
- Noise Addition: AI models begin with random noise, gradually ‘denoising’ it to create a recognizable image.
- Training Process: The training involves adding noise to images and teaching the model to recognize and reduce that noise effectively.
- ComfyUI Overview
- User Accessibility: ComfyUI is a user-friendly interface for AI image generation available on multiple platforms (Windows, Mac, Linux).
- Workflow Visualization: The interface allows users to visualize the steps taken from text prompt to image generation.
- Practical Applications of ComfyUI
- Image Generation Workflows:
- Text-to-Image: Basic workflow involves loading checkpoints and denoising images based on text prompts.
- Image-to-Image: Users can modify existing images by inputting a base image to guide the generation process.
- LoRA (Low-Rank Adaptation): A lightweight model that can be attached to enhance larger models with specific styles or character representations.
- Technical Considerations
- Checkpoint Loading: Models must be selected based on the desired output, and parameters must match the original training data.
- Sampling Methods: Different algorithms (e.g., Euler) and settings (e.g., CFG - Classifier Free Guidance) influence the creative liberties of the model.
- API Integrations: ComfyUI can connect with cloud-based tools like Runway and OpenAI, allowing for more extensive generation capabilities without needing local hardware.
---
Key Takeaways
- Control vs. Convenience: ComfyUI offers more control over the image generation process compared to simpler platforms but requires a steeper learning curve.
- Hybrid Approaches: Many creators are likely to use a combination of ComfyUI for detailed work and other tools for quicker, less granular tasks.
- Future Developments: As AI tools evolve, the balance between robust control and user-friendly interfaces will continue to shift.
---
Conclusion This episode of Denoised provides an in-depth look at the workings of AI image generation through ComfyUI, highlighting the importance of neural networks, noise management, and latent space. The discussion emphasizes the potential of AI in creative sectors while recognizing the need for both control and accessibility in tools used by creators.
For further insights and tutorials on ComfyUI and related tools, listeners are encouraged to follow the podcast and engage with the hosts through comments and questions.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28All right, welcome back to Denoised. ConfUI is the universal tool across different AI companies, different creators. Yeah. Once you sort of get to a certain level and you're like, I need to either have more control or automate things. Yeah. ConfUI usually comes to the conversation. But I mean, we've had this conversation before of, you know, with tools like runway references or FlexContext, how useful are LoRa's and stuff like that. We're going to have ourselves. Let's start high level. We'll talk about that a little bit later in the episode. Yeah. This is going to be a fun one. And please engage us with questions and comments and we're going to get to them as well.
0:59Yeah.
1:03All right, so let's talk high-level first, AI image generation. AI image generation. How does this work? What is happening? It's actually a miracle that it even works. If I go under the hood and sort of explain how it actually works. So one thing that all image generation, video generation relies on is a neural network. And a neural network is what separates computer vision, machine learning from AI. So when people say, well, we've been using AI for 20 plus years. Not exactly neural networks. That is a new thing, especially neural networks at scale. When we talk about an image generation model like the one Mid Journey has or stable diffusion, you're talking about billions of nodes inside of a neural network.
1:48And a neural network node is nothing complicated. It's actually a very simple mathematical function. It's x times y plus b. So it's essentially linear algebra. So you have a slope, which is the x, and then you have a bias on the slope, which is a plus b. The thing that makes it incredibly useful is when you do this millions and millions and millions of times, you have millions of nodes, and then you go to billions of nodes, then it starts to develop a source of understanding of big concepts and not just concepts of what an image should look like but associating that with words and that is the magic here like for example if you have a notion of a beach regardless of language you could say it in mandarin in tagalog in american english but a beach that concept is associated in that neural network to that of malibu that of you know can film festival like beaches everywhere in the world are associated with the word beach and the notion of the word beach so how is this all built it's actually built from images and the way we train them i have an actual tactile demo okay so this is noise and it's hard to believe that a neural network can turn this into an image that we can recognize so it goes from this to this or it can go in the reverse order it can actually take this go back into noise and it happens by training so the way it does this is it adds steps of noise so it'll take an input image and then it'll slap on some noise and you can still kind of see the image there but this is what it's using for training to understand the denoising process and it's the training of the denoiser that's the magic so it's taking a nice clean image adding the noise and sort of memorizing each step of like yeah this level of noise is for this image and it's doing this billions of times yeah and the neural network is able to train itself on what is noise and what is image which is the crucial part and as the noise gets thicker and thicker like here you can see the job gets harder and harder Like you could barely make out the text and our faces there.
4:13Yeah. But it is using recognition of features in there and it is doing it in a mathematical space called latent space. Latent space. All right. Yeah. Which is quite different from how we traditionally compute stuff like edge detection or video compression. That stuff happens in the frequency domain, typically using convolutions and cosines and sine waves. This is happening in a space that is multidimensional. It's happening in the multiverse. Exactly. So each vector can have 20 different commas on it, right? Because it's 20 different axes in this multidimension. and uh you know if you imagine like the way i imagine latent space is you have like this gladiator arena this giant space and then in there there's tiny little neighborhoods like you have your santa monica you have your echo park you have your valley and each neighborhood represents an idea of a thing that we associate with the real world like echo park is just all about, you know, lamps and lights and table lamps and floor lamps and lights.
5:33Santa Monica is all about the beach and the sand and the beach umbrella and those things. So like they're clustered together. So when a text prompt comes in and says, you know, I want to generate a light that's sitting on the sand and a beach, then it's pulling the latent space is pulling from Echo park it's pulling from santa monica and guiding that noise into an image based on the training data based on when you trained it with the image you also were had proper labeling and identification to know that correct these is the two guys on a podcast image or this is a picture of a beach and so it's trained on that way so then when you want to generate something yeah a real network is only as good as the training data and uh like you nailed it in the head the captioning process and the data grooming process is as important as the training process.
6:24So you're talking about millions, if not billions of images. That's why scale AI is worth billions of dollars. Right. And we're only getting started. So if you follow the big AI companies, you'll see an exponential growth in scale. So for example, ChatGPT4 to ChatGPT3, I think, had like 100x scale in the size and complexity of the neural network. And they are looking to one-up that one again and again. So we're going to get to a point where the notion of things not only exists, but the granularity in which they exist is even finer and finer. So then I could determine a beach in Santa Monica versus a beach in Tahiti versus a beach in south of France.
7:04And those things will have individual, very granular changes to them. Okay. So to summarize training, we take a high quality image with good data captioning, and then it gets noisied up step by step. the models are trained on like this level of noise at this step kind of produces this image so then later on when we're like okay i want to generate an image we start with a very noisy you just feed a random noise you feed a random noise right and that's where the seed also the seed is a thing that comes in so like you need to start with a base noise image empty latent image yeah called in nice and comfy ui i'm checking because i know comfy you know i gotta add an empty latent image yeah just which looks like one of the static images that you just showed and then it so you give this to the case sampler in comfy ui and then based on your number of steps it starts denoising that image to create based on whatever your text prompt was correct correct high level explanation yeah and the training is so computationally expensive joey you're talking about millions of images but it doesn't go through the neural network once it'll go through it thousands of times.
8:12And each time it'll like put different weights on each neural network node differently. And then there is a loss function. So then when it generates something and we determine it as error or another machine determines an error, it goes back in and tries to correct itself and gets better in the next one and the next one. A lot of training. All right. So let's look at this in practice in CompFUI. And CompFUI is kind of nice too, because you sort of, I mean, it can get very high level and technical, but this also gives you more control. But this also is a good way to visualize what's happening under the hood with a lot of AI models that you are just kind of going to the website and typing in.
8:48This is, in a lot of cases, happening behind the scenes. Yeah, ComfyUI is a free platform available on Windows, Mac, and I think Linux as well. We were talking about this. We don't know how they make their money because this stuff is amazing. We still have not quite figured out. Yeah, it is free. They've also made it a lot more user-friendly because originally you sort of either had to install it on terminal or from like GitHub or, and now there is an installer. Yeah. There's a Mac installer. There's a Windows installer. I know we have a video about how to do it, but it's literally just download the installer package and it sets up any Python stuff you might need.
9:19So initially when I try to get into this, I'm not the most technical terminal person. So it was a little bit scary and complicated, but now it's a lot easier to use. You and I, I think we're both visual learners. For me, node-based stuff just is so intuitive. And our viewers out there, if you've use Nuke, if you use Resolve or Fusion, if you've used Houdini, even Unreal Engine Blueprints, you've already got it down. You know Node-based workflows, so you should have no problems. Or even any basic flowchart builder and whimsical or Miro or something that's the same idea. Exactly. So I loaded up one of their default text-to-image templates.
9:55I got a lot of good templates. This is a good one because it kind of shows all the basic building blocks that happen for taking a text prompt and turning it into an image. First up in our workflow, we got the load checkpoint. Yeah. So what are the checkpoints? Yeah, checkpoint is the model that you're loading. So if you're generating an image, you could pick Flux, Dev, Flux, Schnell, Stable Diffusion. There's a ton of models to pick from. Yeah, the one that default here is the Stable Diffusion XL, which I know a lot of people still like. Yeah, it's the most popular one. It's also open source. It's open source.
10:27You can run it locally. Yeah. Anyway, I should say too, we'll talk about APIs in a second, but the other advantage of Comfy that people like a lot is you can download the models and run it on your local machine. So if you want to experiment or generate stuff, you are generating it on your computer. There's no credits. There's no pain for each generation. So it's good for messing around and experimenting on your hardware. Just one note, you do need an NVIDIA GPU here. You do need that, yes. Because these things run on CUDA. Yeah. there are some i mean i've run it on my mac yeah some stuff it'll run slower i almost burnt my legs right with my laptop on my lap and spinning it but yeah it is possible but it is more designed for run on nvidia cards yeah and i'll just say uh if you notice the load checkpoint stuff it's loading a save tensor file so what that is is is a pre-packaged pre-compiled version of the model that is uneditable.
11:25And that's how the load checkpoint node likes it. So there are two types of ways you can download a model, a diffuser file and a safe tensor file. Okay. Diffusers are more, they're more open to being modified. Okay. You can also load those here as well, but the safe tensor is the one that's a single file. Safe tensor, single file, not modified. You can modify it in your chain that you're building with in Comfy, but that model itself is locked down. Everyone's got the same model. Yeah. if you downloaded what is because i've seen this on different models too like fp8 fp16 yeah that's a floating point 8-bit or floating point 16-bit it's the precision of the model itself so all the parameters either holding it in a 16-bit float point which is just a fancy way of saying a really long decimal number okay or 8-bit would be half the length of the decimal number so i think 8-bit is like 0.001 up to that accuracy would you pick one based on your hardware because like wouldn't And FP16 require more B3 cards.
12:25I would say go with FP8 for ideation and stuff like that, and FP16 for the highest quality that you can get. Okay. All right. So then we're going out of our checkpoint, our model, and then we've got two clips, the LIP text and codes, one for our positive prompt and one for negative prompting. Okay. So we got the clips. We got the two clips because also this model supports positive prompting and negative prompting saying what you don't want. But not every model does support that. But in this case, it does. Correct. And this is not the same clip that you and I think clipping like clipping an image.
12:56Not clippy? Not clippy. This is a completely different clip. I think it stands for contrastive language image pre-training. It's exactly what I just talked about. It's the Silver Lake and the Santa Monica and the latent space. So not just words, but notions of objects and places and experiences can be stored in the latent space that are associated with words. right so the actual word actual language is not as relevant as the fact that you're trying to pull an idea from the latent space and so this is actually going to take your simple text prompt here and it's going to encode it into a vector that goes into the latent space and then the solve will then use that vector to generate the image okay and i guess at high level this is taking the text prompt that you wrote out and turning that text into something that the they can understand and turn into a visual image yeah into latent vectors yeah okay and then these are all getting plugged into the case sampler and the case sampler sort of been described as sort of the heart absolutely image generation and so we've got both the model our checkpoint loading of the case sampler our positive and negative prompts loading to the case sampler and then a latent image, which is basically just a blank empty image at the resolution that we're going to generate it.
14:20Here is 512 by 512, which is pretty small. This one's really important because a lot of models have specific latent image requirements. They sound like what they were trained on or what they understand. Exactly, and what they are inherently doing under the hood. So, for example, SD 3.5 Stable Diffusion Latest is a 1024 by 1024 model. So if you give it a weird latent image size, it's going to struggle a little bit. So give it its native resolution. So don't do what I first tried to do, where it's like, I want a 19 by 20 by 1080 image. Yeah, so there's remedies for that. You can add upscaler models at the end to get to that.
14:56Right, but you'd add that afterwards, you wouldn't change your initial latent image because it's not really trained on. Exactly, to do images of that size. Okay, so in our case sampler, we do have the seed. And so we talked about this before. This is a random number generated, but this is how it generates the very first base layer of noise. Yeah. Think of it as a unique identifier for the noise. Theoretically, theoretically, if you have the same seed and the same prompt and you push it through again, you should get the same image. Right. Theoretically. Every other settings are the same. Right.
15:27You should. I mean, I've seen people keep track of seeds too, if they find a seed works out well for some scene and they might try to use the same seed for other shots in the scene. theoretically possibly helps get more consistent shots it should yeah i mean it should get you 80 more consistency than just random seed random seed random seed and then we got a bunch of other settings uh i will say so steps can you talk about steps yeah so that is the number of steps it is denoising at so imagine that thick fog of noise uh broken up into 20 different layers and And then each layer is being denoised 20 times.
16:04And in theory, the more steps you have, the better image you should get because it's doing more steps. It's diminishing returns. Up to a point, I think I've seen like 20 to 30 be the normal here. And then anything beyond that is just going to be computationally expensive. Like if you do 50, you're not going to get a better image. You're wasting electricity. You're melting your card more. Yeah. So CFG is classifier free guidance. And the higher this number is, the more you're giving the AI model leeway to generate what it wants and almost pull away from the text prompt, from the clip. So if you have a high CFG value, then let the model be more creative, right?
16:44Come up with the thing that it wants to come up with. But if you have a low CFG, then it's going to really stick to what you prompted. The sampler name, this is the mathematical algorithm that's used under the hood. Euler is the most common. it's the most computationally efficient I've seen ancestral Euler I think that's in there yeah ancestral I see that too yeah so I see some of that but generally I don't see this being changed yeah I mean I see like 20 different options here like if you're an AI scientist you're using all of the you're asking for more but Joey and Addy is using Euler okay this is one where it's pretty much yeah you don't need to mess with that too much yeah scheduler scheduler I actually am not familiar with I usually leave schedule and do noise the nice thing is if you do hover they do have pretty good uh tool captions tips pop up the scheduler controls how noise is gradually removed to form the image yeah probably when you just you know unless you know what you're doing just leave it alone yeah and i just want to say that this entire block the case sampler block the mathematical the mathematical computation is happening in the latent space so this is not happening in xyz coordinates this is not a cg render this is not anything that you and i can kind of comprehend in our brain but it is happening in a high multi-dimensional space that only exists inside of a computer like not yeah we don't break the computer open like in zoolander the files are in the computer but the read so why do we use latent space the short answer is it's a compression so if imagine storing a billion images inside of a model and then shipping that model in a safe tensor file you would be in the petabytes right even with image compression so latent space allows you to preserve all the notions all the ideas of that image and each image comes down comes down to like kilobytes or bytes because it's no longer in pixels it's just in this multi-dimensional vector that it exists a vector is going to be way lighter than a pixeled rendered rastered image right so that's why we bother with latent space is that you can you can get to scale and then you can still ship the models over the internet you know a 20 gigabyte model is honestly not a big deal versus like a petabyte okay yeah yeah okay and so then we go and then we have denoise which which is the name of our show they they are promoting our podcast so that's very nice thanks thanks no this would be coming to play more if you're doing like an image to image yeah i believe so yeah just generally here you know the the steps take care i guess it would be the amount of denoise per step uh when the amount of denoise and applied lower values will maintain the structure of the initial image allowing for image to image sampling right so in this case because you're starting from nothing we want we don't want the noise to stay but if you're starting from an image to image workflow, you would want, it would control like how much of the original image do you want versus how much do you want it to change?
19:40Great, great. Yeah, we're going to get into image to image in a little bit. Yeah. And then after our case sampler, we've got the VAE decode. Yes, the VAE stands for variational auto encoder and decoder as well. So this is where you go from the latent space back to pixel space, stuff that we understand. And so this is taking all of the latent stuff, turning it into an image, and then it saves our image. we have a save node which will save our image yeah now you're back to like a jpeg or whatever yeah png and this is the other nice thing about comfy is you can set up batches and it just saves the image to your computer uh so you don't have to you know if you're working on one of the web portals yeah they're used to and then you kind of have to like flag which stuff you like and then download them individually later right comfy you're running it on a computer and you designate a folder file name prefix and everything you generate gets saved you know what i found out recently too um any image you generate on comfy if you take that image i think maybe i think it's only png files but and if you drag that image into comfy it saves the entire workflow oh and so you can pull up if you have an image or someone sends an image yeah you can pull up the entire workflow that was used to make that image which is really awesome because i know why uh so under the hood this whole note-based thing is really just a text file.
20:54And it's a JSON file. Those in VFX, very familiar with JSON files. You can attach the entire JSON file as a metadata to the image. And Comfy will read that metadata and recreate the notes. But yeah, I did not know that it did out of every image. So if you are sharing images or something and you want to just see what the workflow was. It's all in there. It's there. Or if you make the image and you're like, what did I use for that thing? And you didn't save the workflow, you can just drag it in. So I just learned that. That's also very cool. The other cool thing here is Comfy keeps a history of all the input and output images in the local folder that you install it.
21:27So even if you're not saving images, it's saving it for you. Yeah. So, yeah, it's very good for just like pulling back stuff back up again. Yeah. Okay. So now that we've covered basic image. Text to image. Yeah. Let's cover some other options. So the two other one common workflows would be like image to image and Laura's. First off, we talked about this before, but Laura recap. What is Laura? Yeah. Laura is not the name of my girlfriend. L-O-R-A stands for Lower Ranking Adaptation Model. So it's essentially, you train a model, a very lightweight model, typically like 100 megabytes, and you can attach it to a larger model, which are in the 20, 30 gigabytes, right?
22:07And what I imagine visually, like the large model is like the Queen Mary 2 in Long Beach, right? It's a giant cruise ship or any of your carnival cruise ships in Miami. and the Laura is a tugboat. It's like a tiny little ship, but it has the power to pull that giant thing to where it needs to go on the dock. And the Lauras are never modifying the actual models. All Lauras are doing is it's attaching itself to the correct attachment points in the big model to influence it significantly. And this could be a style. This could be a face. Yeah. Logo. Characters. Characters. That's probably the number one use is like, I want to create Joey every time.
22:52Yeah. You have tons of photos of you. So what you would do is you would train a Laura model on Joey, call it Joey, and then attach the Joey Laura to Flux, the big actual Flux. And then you type a man wearing a suit and then you tag it with a special tag word, Joey. Whatever I named the Joey. Yeah. And then it knows that now you're invoking the Laura. so it'll generate not just a generic man but you right right yeah okay and you train laura as you just use a like replit it's a rep not replit that's the coding one replicate is probably the most uh there is for flux you have flux gym for stable diffusion you have automatic 1111 and uh koya so there's tons of places for you to train laura's usually pretty easy web interfaces where you just give it the images and it'll train it and give you the laura file it's not going to be free i think it's there's a cost but it's dirt cheap yeah i think i can trade one for yeah a few dollars yeah under five dollars you could train a laura and you can train a laura on comfy y as well but i don't recommend it because the internet tools are so much um just just easier yeah yeah all right so uh we'll cover the laura one first this is just a general laura template inside comfy and so a lot of the basics are the same so we have our load checkpoint loader model but now before we go into our case sampler we have load laura correct yeah so you're loading the checkpoint which is the laura model not the main model so you're going to have two checkpoints here the load checkpoint is our main main okay correct so then you're going into the load laura which is gonna take the main model attach the laura to it and then send that to the case sampler which is what's happening here in the load laura node so you see um under on the load checkpoint there's a clip white yellow line there.
24:42So this model also has the clip loaded in it, but a lot of times it could be separate from the model. So you can have a load checkpoint and a load clip, and you might have to load both. Oh, you might have to load a clip? So like, as in, how would that be different than our clip here, which is like just our text prompt that we want? So a lot of times the models generally come with a clip loaded on top. So the safe tensor file has this in there and you don't even notice it. But for more advanced workflows, you can actually put in a much higher quality clip encoding mechanism. And so the triple clip loader is what I see a lot of times professionals use is it'll have a large, it'll have a clip L, clip G and a TPXXL clip model.
25:32And those three clip models work together to give you three different vectors for that text prompt. So what is this doing exactly and how is this different than our clip text and code here, which is just our prompt? It's giving you the vector in three different ways. So the prompt is giving you three different results here. And it's just giving the case sampler a more variable environment to build off of. So it's just giving more meat for the generation. Okay, but you would still use this in addition to your text prompt that you want to generate. So right now, if you see the load checkpoint, that yellow line, the clip here, so you're getting the clip from the safe tensor file there, and that's passing from the LoRa into the text, into the case sampler.
26:17But if you want it to be a little bit more advanced and give the generation just more to work with, you can get a triple clip loader and then attach that yellow line. Yep. Right, right there. Oh, I put it before the LoRa. Before the LoRa. Yeah. And then you would go about it that way. Okay. You don't actually need that because you're not going to use the clip from the checkpoint. Got it. Okay. Interesting. Good tip. Where would you get these clips? Is it something you have to download? Is it just on the... Great question. So in general, I would say 99 % of the stuff that you actually need to download from ComfyUI exists on a repository called Hugging Face.
26:54Okay. So Hugging Face is like the national database, the international database, if you will, of all of the models. You could get safe tensors there, diffusion models there. And that's not the only place. There's a couple other places as well. If you go to civet.ai. Oh, yeah, they have stuff too. So civet.ai specializes in LORAs. So if you're looking for a specific style, like let's say you want an anime style transfer or you want a character that's consistently a certain ethnicity, you go to civet.ai and grab that LORA. But that LORA is going to be specific to an image model, right? You can't just use a Flux Laura on a stable diffusion model.
27:34Right. And so here, actually, I pulled up because I was curious. So Comfy does have a good model manager built in as well. Yeah. And I searched for clip models. And so here, Stable Cascade from Stability AI, a dedicated clip model. So I assume that would go with stable diffusion models. Yeah. So these are probably referencing back to a hugging face repository. But the nice thing about Comfy here is it'll download directly into your Comfy local folder. So you don't have to go to the internet and find the thing. Text encoder for Flux. fp16 text encoder for flux fp8 yep so okay to what you're saying yeah and then the other neat thing here is if you're starting with comfy you've just downloaded a workflow if you go to the manager uh there is an install missing nodes button yeah which yeah which is cool because like a lot of times there are nodes that are just doing mathematical functions right it's like you know cropping an image clipping an image whatever and for that it'll just find it on the comfy server and then put it up here.
28:29Yeah. Okay. And then as far as everything else on this workflow, it's exactly the same as what we just covered. Yeah, it goes into case sampler, it goes into VAE to code, and then it saves your image. Yeah. So yeah, we just covered all of this. Now, the fact that you have the LoRa in here is going to significantly impact your image generation. And in your text prompt, you're going to have to add the tag word for your LoRa. So like, if we added Joey, it's looking if the Joey is not in there, it's not going to invoke the laura all right okay and then image to image real quick image to image so now let's say you want to have an existing image and you want to modify that image yeah so this is where i think a lot of professionals will end up is um you'll find that text to text to image is not going to get you where you need to go most of the time in that instance you already have a good idea of the framing the blocking where you want the character where you want the background to be what objects you can do something like sketch it up.
29:23Yeah. Or you can go into Unreal and you can block it with a camera and take that render. Bring that in as an input to the generation. So here we have an image to image. And our workflow is pretty much the same because we have this blue box. Yeah. Which is to load our image. So we got a load image node. Correct. And see, all this is happening is it's taking the image and it's encoding into latent space with the VA variational auto encoder. And then that is going to go ahead and influence the K sampler. So instead of giving it a static empty image, you're giving it an actual image. And also we can see here that the noise setting that we talked about before, it was set at one when we were doing text to image because we didn't want any noise, but now it's set to 0.87, which would modify it a lot, but keep some of the basics of the image.
30:05Yeah. And all that happens under the hood is like that image actually ends up having structure. So like if you picture a portrait of a man, you know, you have the shoulders and the head, there's like structure to that image and it's going to use it in the in the guidance i also noticed here too the the sampler for this is not uh it is dpm pp 2m it's that i would say if you load one of the default comfy workflows and settings are changed leave it as is and yeah to start with and then maybe modify it you know to see what they do but um they usually set these for whatever the optimal parameters are for the workflow that you're loading.
30:45Yeah, and full disclosure to our viewers, like I am not the comfy UI expert here. I'm just scratching the surface. Yeah. And I think you can say the same. Oh yeah, I'm just catching up. The AI researchers of the world are the ones using this to its full potential. Yeah, that's where you're really into the mathematics behind this and stuff that I'm not quite sure about. But we're bringing you the knowledge that we know. Yeah. Everything. Yeah, this is a good leaving off to a point too because one of the interesting uses that I found for comfy lately is, this is everything we're talking about. This is all running locally.
31:14And so you also like it, this is kind of either going to kill or not even doable to run. Like I haven't run any of these as demos because I'm on a MacBook and it would kind of kill it. But even if you did have your own hardware or if you just had a MacBook, what I found useful is they have API nodes. And so the API nodes at first I thought, oh, that seems like a little bit overkill, but they're actually really cool. And so they have a whole variety of, I'll show you in the templates. It'll connect to runway. It'll connect to Google. It'll connect to any of the main cloud tools that have APIs. And so they've got connections to LLM, so OpenAI or Gemini.
31:46They've got connections to image APIs, so Flux, which you could run locally. Or if you don't have the hardware, you could go connect to their cloud service and use their cloud servers to generate. Runway, stability, OpenAI, ideogram. And then they have video APIs as well with the main, you know, Pika runway. Yeah. So how does billing and pricing work? So the building is cool. So at first I was like, oh, this seems cool. But then you don't have to like sign up for the APIs for each and developer accounts for each tool. But not the case. Comfy has a cool basically just connect your account and you can just buy credit.
32:17So they're reselling. yeah and they're not marking it up they're just charging you whatever the base value is of the nodes and then the cool thing too is that'll tell you what like this is a runway text to image node and i think because runway's pricing is variable it doesn't have a price here but on some of these nodes if it's like you know 0.08 cents a generation it'll tell you up here exactly how much it's going to cost for each generation before you hit the button before you run it and you don't have to connect any API. You don't have to, you just load this node and it just takes the money out of your account.
32:49So it's awesome. Really handy. And then the advantage of using this is you can do, you can connect this to your comfy workflow, but you have the advantages of everything you generate automatically saved to your computer. So you don't have to like use a web portal and download each image. So if you're trying to do a lot of image generation using these models, that helps speed it up a lot. And I've been using a new node too, where you can do a batch generation. So you can give it a text file of like a bunch of different prompts and then have it generate each prompt and so that's kind of handy for yeah just getting a bunch of stuff out to kind of like look around and explore and brainstorm uh rapidly so i think uh because like you said you're in a macbook you know you're not an nvidia gpu you can still have the flexibility of building your own stuff but then running it on the cloud exactly i just use cloud access cloud hardware right to uh to to generate stuff yep so that's i've been finding this to be pretty handy that's awesome and a good I've never actually used API, but I'm sure this is the way to go of the project that you're on right now.
33:44Yeah, exactly. And then we've talked about Comfy and building out these models. It's still kind of a base. If you were to compare running this workflow of a text image using Stable Diffusion or even using Flux, compared to the more accessible tools like we have on Runway, like we have Flux Context on the cloud with references, how useful or how necessary is it still to train a laura or build out a complicated workflow on runway when we have like the chat gpt interface where you can just give it an image of someone and be like put this person in a car yeah or runway references and you can just do the same thing conversational single image you don't have to train a whole laura yeah how useful is comfy in these workflows with the advancement of the other tools that we're seeing yeah great question I think it all comes down to your use case and what you're trying to do.
Read the full transcript
34:38Every AI creator is going to end up doing a hybrid approach where they're going to do some of the stuff on Comfy and some of the stuff on the commercially available tools like Runway or Luma. And then they're going to combine the results in their own unique ways. For me, I find that LoRa's that you build, you're on your own. You have really specific control over the captions, over the training set. And so you directly influence the quality of that LoRa. And then when you're doing image-to-image workflow, you're using depth map, canny, edge detection, all those things. So you have direct influence and control over the ChatGPT stuff.
35:18It's good enough for impressing somebody or doing a social media post. But when you're getting paid, when you're a commercial creator, right? some brand is paying you to make something you want the highest level of control you feel like you feel like you are obligated because of this looks more complicated that we have to use this because if we're because it because we got more nodes no i i don't think like i mean i know i mean i feel like like yeah sometimes it's like oh man it feels too easy on chat gpt yeah like this should have been more this should have been more complicated for like you know okay i'll put it this way i'll put it this way a chat gpt i still keeps like when i want to give it you know a character reference image and just like hey may put this person in a car it still does really good job does a good job i would say for 90 of the way it's there and uh sooner or later the stuff that we're building here is going to get gobbled up by the big commercial services right like flux context is now essentially making lauras for you right and now they do have a version that you can download and run locally correct which last time we talked about that that wasn't available yet but now it is available yeah i would say with the comfy stuff if you want to be absolutely cutting edge and have the most level control you're like six months to a year ahead of the commercially available tools so if you're just not getting the result that you want on chat tpt image you're in then come back to comfy and try it here you need to do volume or like yeah a lot a lot of stuff as much as i was at another workflow idea too with the api nodes you could get the node you could have a prompt and then connect it to sync one prompt connected to like every image generator node yes and instantly spin up and see like okay with this prompt how am i getting from each tool and test it out and to be fair there are services now that let you do that so weavy is one of them free pick we talked about we did yeah and then ltx studio flora ai yeah uh there's like five or six but node-based yeah the thing i kind I wish Flora was a little bit more like Comfy because Comfy could build out your workflow and then you hit run and it moves through the entire workflow.
37:29Flora was a bit more like you kind of had to run each node step by step. Everybody is after the Comfy business. This is what I mean by that. Like, first of all, we should make their tools free for the Comfy business model. They know that professionals want control and consistency and repeatability. and Comfy gives you that. And the no base stuff could be, like you said, it's just a little bit spaghetti for most people. So Invoke, Weavey, LTX, Flora, they're all making Comfy-like features that are perhaps a little bit of a simpler user experience. I mean, I feel like for a good example of like the top level of what you can do with Comfy and some of the crazy workflows is to look at McMuppets, who we've talked about before, where he's built crazy, complicated, massive, comfy workflows for...
38:19Character consistency. Character sheets, creating one-shot character sheets. Actually, that is a pretty good use case of comfy that no other tool really does right now. Styling in Blender, adding LORIS to characters. I think Mick does a really good job of combining traditional 3D with new AI rendering. Yeah. So that's something that is quite difficult to chat GPT, right? Like, yeah, you can upload an image from Unreal and da-da-da-da-da. But then Unreal has something called Comfy UE now. So you can take screenshots from Unreal right into the Comfy instance. Oh, okay. Because I know Blender's had that workflow for a bit.
38:55And I've seen other people connect like Cinema 4D and other tools. It's all about speed and throughput. As an artist, you're going to be in charge of like 100 shots or whatever. You can't just go to runway every time. It's going to be too cumbersome. And also cost too. Because if you have the hardware and you can just spin this up on your local machine, that can save more money versus having to pay for API credits every time you want to generate. 100%. Yeah. All right. I think good place to wrap it up. Links in the video for the Comfy install. We'll put that at denoispodcast.com. But yeah, let us know if you have used Comfy or you kind of land on Comfy in this debate.
39:29Because this is a conversation to keep going back and forth with of like how useful is Comfy versus how much better the tools are getting at just being like, I want this. And you tell it what you wanted. Yeah. And six months from now, like this could be completely useless of a video. Like that's how fast this is moving. And also like, let us know if this is the type of content you're looking for. We'd love to be able to go another step a little, maybe a little bit more advanced into this, give you even more of our knowledge here. Yeah. Yeah. Especially as we experiment more and find out more workflows that are useful.
39:58All right. Thanks everyone. We'll catch you in the next episode.
From the publisher
Curious how AI actually turns text into images? In this episode, Addy and Joey break down the inner workings of AI image generation and explore ComfyUI. We'll explain the core concepts of latent space, diffusion models, and how noise becomes a recognizable image. From basic text-to-image workflows to advanced techniques with LoRAs and image-to-image transformations, discover when to use ComfyUI versus web-based tools like Runway or ChatGPT for your creative projects.
---
The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.




