In short
Podcast Notes: Leveraging AI - Episode 192: Create AI Images Like a Pro!
Episode Overview In this episode of the Leveraging AI podcast, host Isar Meitis talks with visual AI educator Luka Tisler about creating AI-generated images. They discuss various tools and techniques for generating visuals, focusing on the importance of consistency in branding and how to use both proprietary and open-source AI tools effectively.
Key Topics Discussed
- Introduction to AI in Visual Content
- AI-generated visuals are becoming essential in business.
- Consistency is key for branding when creating images.
- There are tools like ChatGPT and MidJourney that are accessible but may lack the consistency required for professional use.
- Different Paths to Visual AI
- Conversational Path: Using AI like ChatGPT for image prompts.
- Classic Tools: Established platforms like MidJourney for image creation.
- Open Source Frontier: Advanced tools that allow greater flexibility and control over image generation (e.g., Stable Diffusion, ControlNet).
- Understanding Open Source Tools
- Luka shares insights on using open-source tools to achieve brand-specific imagery.
- These tools offer a steep learning curve but provide powerful capabilities for those willing to invest time.
- Exploration of ControlNet and Image Generation
- ControlNet allows users to guide the denoising process of images by inputting specific references (e.g., poses, styles).
- This provides a more controlled output compared to traditional prompting methods.
- Proprietary vs. Open Source
- Proprietary tools like MidJourney offer speed and user-friendly interfaces but may limit the control.
- Open source tools provide extensive customization options but require more technical knowledge.
- Future of Image Creation
- The merging of design tools with AI capabilities is anticipated.
- Users may see integration of AI-generated features in applications like Canva, which can revolutionize design processes.
Key Takeaways
- Consistency in Branding: When generating images for business, maintaining brand identity through consistency is crucial.
- Diverse Tool Options: Understanding the strengths and weaknesses of different AI tools can help users select the right one for their needs.
- Open Source Complexity: While powerful, open-source tools demand a greater commitment of time and technical skill.
- Future Trends: The evolution of AI in design suggests a shift towards more intuitive and conversational interfaces for image generation.
Conclusion The episode provides valuable insights into the multifaceted world of AI-generated images, emphasizing the balance between accessibility and control, and the importance of adapting to new tools and methods for business success. Luka Tisler’s expertise in using open-source tools shines a light on the potential for creating professional-grade visuals with AI.
Additional Resources
- AI Business Transformation Course: Learn more about the course starting May 12. [AI Business Transformation Course](http://multiplai.ai/ai-course/)
- YouTube Channel: Watch full episodes on [YouTube](https://www.youtube.com/@Multiplai_AI/).
- Connect with Isar Meitis: Follow on [LinkedIn](https://www.linkedin.com/in/isarmeitis/).
Call to Action If you found value in this episode, consider leaving a five-star review and sharing your insights about what you learned!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hello, and welcome to another live episode of the Leveraging AI podcast, the podcast that shares practical ethical ways to improve efficiency, grow your business, grow your business, grow your business, and advance your career. This is Isar Maitis, your host. And we've been talking a lot on the podcast on one of the top use cases of AI today, which is creating visual content for business use cases. Now, if you are creating images with AI for either fun or presentations, it has actually become really easy to do. You can do this across multiple tools, but there's still one main issue, which is consistency.
0:32And when you're creating images in a business perspective, consistency is key because you You have either your brand guidelines or your logo or the actual look of your product that has to stay consistent. Otherwise, you cannot use the images that you're generating. And so creating consistent imagery is a problem that requires skills and knowledge and the right tools in order to use. Now, while mainstream tools such as ChatGPT or MidJourney or Gemini are good enough on the entry level, there are other open source tools that are actually providing a whole universe of tools of capabilities that make different things that are problematic with the available standard tools out there.
1:17And they solve for that in really beautiful ways. So the open source world is very much more of a geeky kind of solution. People are like, oh, I don't know how to use it. So today we're going to demystify a lot of the open source tools, specifically when it comes to creating imagery, specifically when it comes to creating consistent brand-relevant, product-relevant imagery. Our guest today, Luca Tisler, has been in the world of graphic design and visual content for businesses for many years across multiple roles in multiple companies. But in the past two years, he's been focusing on helping businesses, training businesses, and consulting to businesses on how to implement AI-based solutions for company-specific branded content.
2:04which makes him a perfect guest to show us through this process. I'm personally played a little bit with open source visual tools, but not a lot. And so I'm personally very excited and curious to see what Luca has to share with us. And so Luca, I'm really happy to have you. Welcome to Leveraging AI.
2:23In the next few years, AI technology will change our world dramatically. Whether you are a business executive trying to catapult your business forward, or just somebody who refuses to be left behind and want to advance your career, this is the show for you. I'm your host, Isar Maitis, a serial entrepreneur and an AI enthusiast. You'll hear invaluable practical tips from innovative business leaders, AI practitioners, and some of the brightest AI minds in our world today on how you can leverage AI in ethical ways to advance your career and grow your business.
3:04Hey, Sartre. Thank you for having me again. Really excited to be here and talk about imagery and creation and open source. Thank you. Yeah. Before we get started, first of all, if you are joining us live, either on Zoom or on LinkedIn, so thank you so much for joining us. I'm sure all of you have other stuff that you can do on Thursday noon, if you are listening to this after the fact. So first of all, just to let you know, we do this every Thursday at noon on Eastern Time, so you can come and join us. There's always amazing people like Luca that are going to share use cases. And then you can chat with people in the chat and get to know other people, network, as well as be able to ask questions, which you cannot do if you're listening to the podcast after it's being released as a recording.
3:42Also, one last thing, we are running the AI Business Transformation course for the spring version already. It's been running for two weeks. It's been amazing. There are 25 people in this specific cohort and they're learning things such as image generation, video generation, content creation, data analysis, data manipulation, writing, basing prompts, et cetera, et cetera, as well as how to implement AI successfully business-wide, how to come up with a business strategy for AI implementation. The next course, we just announced the dates, is starting in the beginning of August. So we do the courses at least once a month, but most of them are private to specific organizations.
4:20So if your organization needs training, you can reach out to me. But if you're looking for a course, you can just sign up for yourself. then we just announced the dates for the next course. You can go to our website or click on the link in your show notes and get to the signup for the course. But that's it from me. Now let's give the stage to Luca and let's talk about AI image generation. If you have any questions and you're here with us live, feel free to ask them in the chat. If you are listening to this after the fact and you're like, oh my God, I want to see this as well. So it's going to be available on YouTube as well.
4:52There's going to be a link in the show notes for you to go to see this on YouTube. but we will explain everything that we're doing and everything that's on the screen so you guys can follow us even if you're driving your car or walking your dog or doing whatever it is that you're doing running on the treadmill that you're doing while you're listening to your podcast so Luca the stage is yours thank you so much Isar so just a brief introduction my name is Luca of course and I've been in business business for the last almost 20 years I started as video production but then I quickly moved to video post-production.
5:23Then I moved to compositing, digital compositing, VFX, animation, motion design. So everything visual, everything that's moving, it's my domain. So about almost three years ago, I found out this little thing on a Discord called MidJourney. And I knew immediately that this is going to change everything. Half of a year later, I resigned in my company, in the company. I quit because there was not enough time for me to learn. So I spent a lot of time learning and discovering. And of course, mid-journey was not enough. I wanted more and more and more. So the next progression was, of course, table diffusion.
6:03Table diffusion is the first open source AI image creation platform, kind of. It's not as relevant as it used to be because there has been some advancement on that area. And also serious leadership issues that almost got them bankrupt and a lot of other stuff. So on the business side, they had an amazing technological platform and they weren't doing great decisions on the business side of things. Absolutely. I agree. I wish them well. I hope their business is back on track. But they did a huge thing to the open source community because they were first that opened up the weights for 1.5 and SDXL models.
6:43So this image model is something like large language models. People teach how to, well, people teach them. And they teach them by inputting an image and the description of that image. And they put both as a pair into the magic box. And they do this billions and billions of times. And all of a sudden, now we have a machine that knows what is a car, what is a banana, what are the glasses, and can also draw them. and it's getting better and better and better. I remember the first tries were very awkward. The pictures resembled the stuff you prompted, but it wasn't that, right? But I knew that this is going to progress.
7:25So it slowly, slowly started to progress from 1.5 to 2.0, now SDXL, and then 3 and 3.5. Just before they released 3.5, there was a new kit on the block called Flux. People went berserk because it was truly a model that understood natural language. So you could start prompting as you speak. And this was amazing. This was incredible because no image generator understood natural language. We had to talk to it with tags. So like city, cyberpunk, dawn, car, reflections, good quality, and so on. And, you know, we just had our fingers crossed to get those images as best as possible. But now, I mean, you can talk to all of them in natural language.
8:17Do this, do that. Give me this, give me that. So the progression is amazing. And there has been also the proprietary models raising quality. So now we have like Imogen, we have Ideogram, we have many, many, many image models that are absolutely gorgeous and they produce amazing images. And the best thing is that with the proprietary models, you have control to an extent. So you can control your imagery, but most by prompting. But open source has something more. And it started with stable diffusion, and it's called ControlNet. So those are kind of differently trained models to guide the imagery, how you want it to look like.
9:04So, for example, you have to know how latent diffusion technology works. I'll keep it very short and simple. So it starts with noise, just nonsense, complete nonsense. And then it starts putting that noise away. And with each step, it takes a bit of noise away and inject some of your ideas that is conditioning, aka prompting. So imagine like you are lying on the field and staring into the sky watching clouds, and you can see a shape of a turtle in a cloud. and you decide, okay, that's a turtle, right? Well, latent diffusion works in a similar way. So it kind of says, okay, I can see that inside.
9:43I'm going to shape that into the image that I already know when I go to my magic box and see what's, for example, sunglasses. So, aha, okay, we have sunglasses. Now I will shape this noise with conditioning, aka prompting, and we will get a nice result. So stable diffusion opened up the models so we can create our own models. We can teach our own models. We can fine tune them and we can create loras. About this a bit later, but now this noise has constraints. Now all of a sudden you can control it by, let's say, with the pose. So if you prompt, the man has his left leg and right hand in the air.
10:27If you prompt this, you probably won't get that specific result, right? Or you will have right hand or left hand. Maybe your hand will have seven fingers and so on. But with control nets, we guide this noise. So we say to the model, okay, here is the image of the man within this position. And via control net, we transfer the pose to our image. So we get the result that we actually want. And pose is just one of those many controlments. We have lots of controlments. And they're different for each use case. Some of them are good for, like I said, poses. Some of them are good for architecture. Some of them are good for, well, a lot of use cases.
11:10And they're getting better and better and better. Yeah. What else? Yeah. So to add my two cents to what you just said, to connect a few dots of the things you mentioned, one is how these models are getting trained. The way they're getting trained is they took a gazillion images and then noised them step by step by step. So what the model is getting is getting the final photo that is an existing image of anything you can imagine. And then it gets another level with 2 % noise and another level with 6 % noise and another level with 8 % noise all the way to 100 % noise. So that's how they teach the model to basically then reverse the process, to denoise from 100 % noise to an outcome because it has seen multiple noising levels of any image that you can imagine.
11:57So this is how it works. And the control nets allows you to guide the process beyond the prompt itself by giving it references that it understands, whether about position, lighting, reflections, outlines of things, like literally anything you can imagine that is a part of creating an image. You can then use, think about it as an additional reference for the denoising process or for the creation process. It doesn't really matter how the creation happens. You're just giving it more references than the written reference that you had in your prompt. Exactly. That's a very important part. And I'm sorry I skipped it, but it is noising and denoising process.
12:35It goes both ways. So you're 100 % right. But not just open source has a control. Lately, we've seen the development also in proprietary models. For example, Midjourney has this amazing thing called OmniReference. And you can actually put in the platform the image you want to recreate, and it will recreate it very, very good. You can input your characters. So the consistency is kind of soft. We'll probably never have 100 % consistency, but, I mean, it depends what you're working with. Most of AI lands on digital media. and on digital media, we usually have smaller screens, right? And the human eye just says, for example, Robbie Margo.
13:22Robbie Margo is very, very famous in AI community because everybody is testing with her face. And if you see a person similar to Robbie Margo, you will just tag her as Robbie Margo and that's it. It doesn't have to be 100 % consistent. It's good enough for our brain, but on the larger scale, it's just not good enough. Then we have to use Photoshop. then we have to use in-painting. Well, we get about 80 % to 85 % done with AI, but the last 10 % or 15 % we have to work on manually. And this goes for all imagery, actually. If you want to bring your imagery to another level, you will have to add the human touch in the process of the creation of the image.
14:04So Photoshop is not going anywhere. For now, I'll say something about the whole Photoshop thing, and then we can dive into the actual examples. But Photoshop is not going anywhere for now, I agree. I think what Canva did to graphic design on the lower levels, AI will do to Canva because my gut feeling tells me that these two universes are going to merge. And you might still use Canva, but you will use AI in Canva versus dragging and dropping and using templates. You will have an idea in your head and you will request it and it will be generated on the fly. And I think what's going to happen is features that you currently have in Canva will migrate into ChatGPT, Gemini, Me Journey, and so on, where you will understand layers and you will understand text.
14:52You will understand grabbing a component of the image and filling up the background where you moved the thing from in a very intuitive way. And there is very small doubt in my mind that's going to happen in the next few months, meaning this merging of design and an AI image generation to one unified environment. And I do think that the only gap between that and the professional world is some additional tools that just don't exist in the simple tools. And I think it's just a matter of time until they're there as well. So I'll add two things that are very obvious to people who are in the professional side.
15:29One is upscaling. So like you said, if you want to print a billboard that will cover the side of a building, the resolution that MidJourney gives you is just not good enough. but there are already amazing upscalers today. And once it's going to be built into mid-journey, then that problem is solved. And the other is really small, fine things on specific textures, specific fonts versus just random fonts and things like this, masking of specific aspects, those of you know what I'm talking about. And so I assume these universes will merge together and there's going to be more professional tools with AI built into them like Photoshop, which already has their version of AI built into it.
16:07And there's going to be the more basic tools that are either going to be Canva or working the actual tools themselves, such as Gemini, Me Journey, et cetera. Absolutely. And we can also see the rights of agents. So you don't prompt anymore. You talk to a machine. So do this, do that, change this, change that. And I think this is one of the biggest things that is happening in AI right now. Because we won't be producing imagery and videos with our mouse. We will guide the model with our language, with our speech. And all the changes will be instant. Yeah. Shall we jump into examples? Yeah, yeah.
16:49What do you want to see? I think it will be interesting to see, first of all, for people to see like ControlNet and what exactly and how it works. So just to give examples of things. And I think doing the same thing in mid-journey would be useful as well. so showing kind of like the omni reference and how that works and i think both these things will show people stuff that they may or may not know how exactly it works and maybe they haven't experimented with that before absolutely so for example here is the mid-journey and you can well i just created this mascot this image for a language school that i'm working with and I used, let me see, I used.
17:32So for those of you who are not seeing as Luca is searching, for those who are just listening, we're looking at Me Journey, which is one of the better AI image generation tools. And we're looking at like a yellow frog or lizard that is teaching in a classroom as the mascot. And it's wearing like a blue hoodie with the logo of the school. And I think what we're going to see is how it was actually created. I don't have this image by hand right now. Let me just close the Lighthouse Academy website. Okay. Subtle, subtle. Yeah. But let's try to create something with Margot Robbie Scratch. First of all, no one prompts from their head anymore.
18:17So everybody's using LLMs. So give me an image prompt for a woman in red dress posing in front of Eiffel Tower. and we're going to make it editorial. So again, for those of you not seeing, we're writing this prompt in ChatGPT and it is going to give us the prompts, better prompts with a lot more details and qualifiers than if we wrote the prompt ourselves. By the way, there are multiple custom GPTs already created and available that are very good at that, that are built to give you highly detailed prompts for image generation. But as you'll see in a minute, even just writing what Luca just wrote will give you a very long four sentences worth of details in a prompt that you can then paste into whatever it is that you're using.
19:21All right, so our prompt is, I won't read it because it's just too long. But anyway, we will set the settings to, let's say 16 by, oh, let's go with three to four. Let's raise stylization and variety. Everything is okay. Version, ooh, I experimented with mid journey version three. Okay. Stylization a bit higher and we are good to go and press. Oh no, we will import our image that I took from internet of Margot Robbie. We didn't say anything about Margot Robbie. We just said a woman, a stylish woman. And we dragged and dropped OmniReference into OmniReference. And OmniStrength, you can use it from 0 to 1000.
20:16I will use the low value, about 300. And now we can also insert image prompts and style references. So if we don't know how to describe the style, we can just drag and drop the image of a style we like or use image prompt. So we kind of get the composition that we want from another image. Okay. So again, just to explain what these are, going back to our initial conversation, when you are now prompting these tools, in addition to the written prompt, you can drop in images and use them for different purposes. So you can use an image as a reference for a person, a reference for a style, a reference for the composition of the actual image itself, a reference.
21:05And when we say style, this could be very broad, like cartoonish versus realistic, but it could go into way more detail, like the color palette that is going to use and so on. And you can do all of that. Definitely mid-journeyed that now has it broken up into different levels and tools. And here we are. we get a woman that is similar to Robbie Margo, but she is not Robbie Margo. Yeah. Because our reference was set very low. But now we can use everything and we can up the Omni Strength. So our woman will be similar a bit more to Marco Robbie. So again, for those of you not watching this, there's a slider next to the image that you upload and you can move it left or right between zero and a thousand.
21:58You can also enter the parameter as a parameter, but the mid journey now moved it to make it more user-friendly where you can just move the slider around and you can control how much weight the image that you uploaded will have in the output. And if you bring it closer to a thousand, it's going to be very similar to the image you uploaded. Again, just to explain to people are not watching, the image that we uploaded is just the face. You can't see the entire person. The prompt is for a woman in her address, hence we see the entire woman. And it still knows how to pick the head of the person from the other image and apply it to the full body shots that we're doing right now.
22:41So the thing with mid-journey is after the prompt, it sets a couple of parameters and you have to be very careful. So if you change the parameters, you always have to delete your parameters at the end and use the new settings that you're giving it. Because right now we created four more images, but the Omni reference weight was left at 300. So because preferences behind the prompt are stronger than the preferences that we use through the website. So each time when we change our preferences, we have to delete those from the prompt and set everything else through the website. And now Omni reference is high enough and we will get a woman that looks like Margot Robbie because we created two sets of images but I didn't set the Omni weight high
23:43enough. And while you're waiting, we can actually check out chat GPT because chat GPT has introduced images in March, I think. And now we can ask it to create the image prompt. Yeah. So we will use the same prompt. So here is the prompt and I will ask GPT create an image based on this prompt. And I need a widescreen aspect ratio.
24:25so gpt images are using different kind of architecture if we check out how oh there it is and yeah she looks much more like margot robbie yeah so as we said at the beginning the latent diffusion works with adding noise yeah well chat gpt uses different kind of structure and it's, I don't know how it works, but it's adding the details from up to bottom. Yeah. You can't communicate with MidJourney like communicating with chat GPT because MidJourney is not large language model. It knows images, but the images that it's creating are much, much higher quality than images that we're getting from GPT because GPT's first priority is language or words.
25:19They also did a good job training it to create images, but it's not as good as mid-journey quality-wise. There's an interesting question while we wait for ChatGPT to generate the image. You know what? I'll touch before I jump to the question. I want to kind of go back to what you said as far as the differences between ChatGPT and mid-journey. And then we're going to go to like ControlNet and StableDiffusion or maybe ComfyUI not to scare people too much away, but pros and cons of the different systems. from a professional approach perspective, Midjourney still provides better results. Yes. From a day-to-day usage perspective, ChatGPT generates good enough results in many cases.
26:01And to get to these results are easier. Why? Because it understands context and you can provide it a lot more information that you can because it's a complete conversation. It's not just, here's the prompt to create the image and now I want to change something in the image. I need to go and write a completely new prompt that will start from scratch, basically. It understands the conversation, understands the context. You can upload your brand guidelines to ChatGPT as an attachment, and you will know how to use them in the image, which Midjourney does not know how to do because it does not know how to read PDF documents.
26:32And so things like that are things that are benefits of using ChatGPT. Another advantage of Midjourney is speed. Midjourney generates four images in about 10 seconds. ChatGPT creates one image in about a minute. So if you want to iterate a lot, doing it with Me Journey is going to work a lot faster. And as I mentioned, Me Journey has built a lot of tooling around its image generation with different sliders and bars and controls of different things and being able to reference previous images and being able to reference previous prompts and being able to add all these different things where that does not exist on the ChatUPT side.
27:06If you're a beginner, I think starting with ChatUPT is easier. If you're a more advanced user, you can get better results with using MidJourney. Either way, both of them, I think, for the average user for generating images for presentations and or basic social media stuff, both tools are definitely adequate. Yeah, absolutely. And you're spot on with adding GPT, different pieces together, and it knows how to create images with those pieces, right? You can upload, let's say, an image of a chair, image of a wardrobe, an image of a bed, an image of a picture, a painting, and you drop those images in ChatGPT and say, create a room out of these images.
27:54And it will be spot on. It will be perfect. But as you said, it has its limitations. And for professional usage, it's not good enough because we need to create images fast and very, very high quality. So chat GPT is just not good enough. Also, chat GPT nor Mejourney cannot batch render images. So you can just say create 100 images based on this prompt and I will select the perfect one, while open source tool can do that. So let's really jump to that. Let's jump to maybe ConfUI just to show people what it is. And I know it might scare people in the first minute, but I think we can explain what it is and how it works and maybe demystify it a little bit.
Read the full transcript
28:38Yeah, this is my workflow. It's a bit more advanced. We can check out the, where is it? So again, for those of you who don't see the screen, what ConfUI is, think about a flow chart of a process where every step of the flow chart, you have multiple tools that you can connect that impact how the image will be generated. So it's not post-processing. It's actually how the image will be created. And those of you who can't see, you can see there's like, I don't know, 20 different boxes, each one with different components and lots of lines combining them that looks like a web of different things. But each and every one of these boxes adds another layer of control over the process on how the image is going to get generated.
29:24And by combining them together, you can be a lot more specific with how the output would look like going back to consistency when you can control the lighting the angle the pose the graphics the style the entry images the size the like literally every aspect of how the model actually works at the model level you will get a lot more consistent outputs which going back to what lucas said if you want to now mimic a photo shoot where you're going to bring a model, let's take your example of a girl in a red dress in Paris in front of the Eiffel Tower, you're not going to take two pictures if you're doing an actual photo shoot.
30:04You're going to take 250 pictures and then pick the two best ones and then work on them in Photoshop and then pick the final one based on that. And you can do that in Comfy UI with no problem with a much higher level of consistency and with a lot more control of exactly how it's going to look like. So think about a professional photographer. You don't just take random pictures. You set up the camera to the right setup with the right aperture value, with the right lighting, with the right camera, with the right ISO parameters, like all the stuff that you want to control in the image, you will set up in your camera.
30:38You don't just randomly shoot and click. And this is kind of like what you can do in the digital world with ConfiUI and open source tools. That was so beautifully said. I can give another comparison. So if you just want to use your computer, you will probably buy Mac. Yeah. But if you want to know every nitty gritty detail, you will buy this computer by parts and you will assemble it for yourself because you have control over each part, what part, what's the compatibility, where you will put it. And you know this thing by heart, not just opening a computer, start working, but you know everything that's happening inside every process.
31:20So, Confi is the same. For example, here is a node where you define the checkpoint or the model you will be using. This is the LoRa. So, the node for specially trained small models that you can train yourself. We have clip loader, we have VAEs, we have resolutions, we have so many things that you can toy with and see what are the results. So I would say that confi is for people that want to know more and most of all are very, very curious. Because if you don't have curiosity in your jeans under your skin, then this will be just too overwhelming. Understand people who don't want to use confi, they just want to produce.
32:03But, well, I need a bit more. I need control. I need to know how I will construct the image. And I can change the schedulers. I can change the samplers. I can change anything and I can influence my image based on all these parameters. Now, in MidJourney, you have a couple of parameters. Here, you have thousands and you can combine them and you can do, well, a lot of things. But what we see right now is a fairly complex, not too complex, but fairly complex workflow. So maybe it would be better if I would show you one a bit more. Exactly. Yeah. So, for example, let's take control network flow. And I have a couple of missing nodes.
32:50That's okay. Not nodes, but yeah. So, let's just switch a couple of stuff. VAE. Not pruned, but yeah, that's okay.
33:11so again for those of you who are not watching now we're looking at a process that has four or five different nodes versus the 30 that was on the screen before nodes are basically building blocks so think about them like legos so you can build a simple lego with a bunch of parts you can build a really complex lego with multiple parts and the more you know what the parts do the more you can use more legos to build more sophisticated stuff so you don't have to start with like the 30-step process you can start with a four-step process and build something that will be a lot easier to construct but still will provide you a lot more control than doing it in let's say mid-journey on definitely on chat gpt correct so what we did now is okay everything is ready we can't run it so i inserted the image reference and i expect my new image will have a person that stands exactly the same as the woman on the photo but i have to give it so so again this is a tool that is not for somebody who's just want to play around and create images this is a professional tool that if you do this for a living you can do this.
34:23Or like Lucas said, if you're just really curious about image generation and you want to experiment with doing things that are beyond what the image generation tools can do, you can do that as well. And again, the cool thing here is that you literally control every single thing. And one of the cool things that we talked about several times in this episode already is control nets, right? So as an example, you can create, use a model of a person or a cat or an anything, and then have the output resemble it in whatever way you want. So either resemble it in the texture or resemble it in the fur that it has or resemble it in the pose that it's standing or walking or jumping in and so on.
35:02And it will know how to do that because there's one component that you control, which is the pose of the person or the texture of the image and so on, where it will follow what the control net tells it to do. Right now, control net is not working. I don't know why. Yeah, the beauty of live demos. So let's do a quick summary of everything we talked about, and then we'll see if you have any final thing to add. There are really three main channels today, right? One is the more professional but yet readily available proprietary models like MidJourney. there is the open source universe that has tools like stable diffusion and flux that have multiple ways to use them including through confui which is another layer of control over just using the model itself but you have to have an open source model to do that because literally what confui does is it controls a huge number of parameters within the model itself that literally what it allows you to do.
36:04And then there are, if I may interrupt you, there are two steps. First of all, you need a powerful computer. You need a graphic card that has a lot of VRAM. So your basic graphic card is not good enough. It has to have at least eight gigabytes of VRAM or more, the more the better. And the second one is the steep learning curve. So it's not something that you just toy with. You have to invest time to learn how to use those tools. But when you use them, when you start using them, you become invincible. You can create anything you want. If, of course, things are working. Yes. Well, I think they're working.
36:42It's exactly the same thing as a complex machine, right? The chances of getting one thing wrong, and then the output is not exactly what you wanted. The more complex the machine, the more chances you're going to get something that is not what you meant. And then really the other component is going to a tool like Chachepet or Gemini, which both can create good enough images for many day-to-day things, but with a lot less control, a lot less parameters, a lot less capabilities to know exactly what's going to happen. For many daily use cases, like creating presentations or an image for a social media post, that is the easiest way to go.
37:15If you are a professional designer or if you need to create stuff that is more consistent and so on, going the open source path, we're just going to give you a lot more capabilities. Look, if people want to follow you, learn from you, know more about what you do, how you do it, hire you, what are the best ways to do that? Well, in social media, I'm not spread around. I mainly use LinkedIn and Instagram here and there. So you can meet me on LinkedIn and send me a message, connect to me. And if you have any questions, I will be happy to answer them. Awesome. Luca, thank you so much. I think this was very educational and valuable to people.
37:54I think most people, at least people that I know, and I know a lot of people who are playing with AI do not know a lot about the visual side in general and definitely not about the open source models and how they differ and what benefits they provide. So I'm sure this was very helpful to people. Thank you so much. And to everybody else who has joined us, I appreciate you being here. I appreciate spending the time with us and being active in the chat. And until next time, have an awesome rest of your week.
38:27Thank you.
From the publisher
👉 Fill out the listener survey - https://services.multiplai.ai/lai-survey
👉 Learn more about the AI Business Transformation Course starting May 12 — spots are limited - http://multiplai.ai/ai-course/
AI-generated visuals have become table stakes. But what if you're still playing checkers while others are playing 4D chess?
In this session, we’re breaking open the real world of visual AI — the one most people don’t even know exists.
You’ll see the 3 major paths to creating powerful visuals with AI:
- the conversational path (think: ChatGPT),
- the classic tools (Midjourney, etc.),
- the open-source frontier, where precision, layering, and total control change the game.
Guiding us through this visual odyssey is Luka Tišler — a visual AI educator, workshop leader, and founder of an academy teaching professionals how to wield tools like ComfyUI, ControlNet, and more. Luka doesn’t just play with prompts — he builds pipelines. If you want consistency, control, and pro-level results, he’s your guy.
You’ll walk away understanding the trade-offs between tools, how to make the right choice for your business needs, and why there’s a whole new level of image generation out there waiting for you.
About Leveraging AI
- The Ultimate AI Course for Business People: https://multiplai.ai/ai-course/
- YouTube Full Episodes: https://www.youtube.com/@Multiplai_AI/
- Connect with Isar Meitis: https://www.linkedin.com/in/isarmeitis/
- Join our Live Sessions, AI Hangouts and newsletter: https://services.multiplai.ai/events
If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!



