In short
Podcast Notes: Leveraging AI - Episode 204
Episode Overview Title: The Ultimate AI Showdown: Comparing Image Generation Tools Host: Isar Meitis Guest: Rory Flynn (AI Image Expert and Founder of Systematic AI) Release Date: TBD Episode Description: This episode dives into the capabilities of AI image generation tools, discussing their effectiveness, ethical implications, and practical applications for business professionals. Rory Flynn shares his expertise in using these tools to create brand-worthy images without the need for a designer.
Key Themes and Discussions
Importance of AI in Image Generation
- Visuals in Business: The significance of high-quality visuals in marketing, social media, and presentations.
- Cost and Time Efficiency: AI can drastically reduce the time and budget allocated to creating visuals.
- AI as a Creative Tool: The episode emphasizes leveraging AI tools responsibly to enhance creative processes rather than hinder them.
Insights from Rory Flynn
- Indistinguishable Quality: AI-generated images are becoming nearly indistinguishable from real photography.
- Tool Comparison: Discusses various AI image generation tools, including:
- Midjourney
- ChatGPT
- Imagen
- Flux
- Choosing the Right Tool: Key considerations for selecting the appropriate AI tool based on specific business needs (e.g., social posts vs. product catalogs).
Rory's "Visual Building Blocks" Framework
- Prompting Techniques: A structured approach to prompting AI tools effectively.
- Clear and Direct Prompts: Emphasizes that detailed prompts yield better results.
- Iterative Process: Encourages users to refine their prompts based on initial outputs to achieve desired results.
Creative vs. Coherent AI Models
- Creative Models: Tools that produce visually stunning outputs with minimal instruction (e.g., Midjourney).
- Coherent Models: Tools that require more detailed prompts but provide high levels of control over the output (e.g., Flux, ChatGPT).
- Training Tools: Models that can be tailored to specific needs or styles.
Practical Applications
- Reverse Engineering Brand Visuals: Methods for analyzing existing brand images to create consistent and on-brand visuals using AI.
- Automation in Content Creation: How AI can streamline the creation of multiple assets rapidly (e.g., through tools like Wey).
Tips for Effective AI Image Generation
- Detailed Prompts: The more specific the prompt, the better the output.
- Understanding Photo Elements: Knowledge of photography basics can enhance the quality of AI-generated images.
- Different Models Suit Different Needs: Understanding the strengths and weaknesses of various AI tools is crucial for optimal results.
Future Insights
- Upcoming Episode: Preview of part four of the Ultimate AI Showdown, focusing on AI video generation tools with guest Tianyu Xu.
Conclusion This episode provides valuable insights into using AI for image generation, outlining strategies to leverage these tools for effective business communication and marketing. Rory Flynn's expertise serves as a guide for professionals looking to incorporate AI responsibly and effectively into their creative workflows.
---
Connect with the Hosts and Guests
- Isar Meitis on LinkedIn: [Isar Meitis](https://www.linkedin.com/in/isarmeitis/)
- Rory Flynn on LinkedIn: [Rory Flynn](https://www.linkedin.com/in/roryflynn/)
- AI Business Transformation Course: [Course Link](http://multiplai.ai/ai-course/)
- YouTube Full Episodes: [YouTube Channel](https://www.youtube.com/@Multiplai_AI/)
Feedback If you enjoyed this episode, consider leaving a five-star review on your favorite podcast platform and share your insights!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Welcome to part three of the Ultimate AI Showdown. If you've missed part one and part two of the Ultimate AI Showdown, you should start with part one in which we covered data analysis and comparing AI tools on how to do that in the most effective way. And part two in which we covered vibe coding tools and what are the pros and cons in each and every one of them. So you don't want to miss those. But now into part three in which we're going to cover how to create amazing images and what are the different tools and how to pick the right tools for the right use case. in the next few years AI technology will change our world dramatically whether you are a business executive trying to catapult your business forward or just somebody who refuses to be left behind and want to advance your career this is the show for you I'm your host Isar Matis a serial entrepreneur and an AI enthusiast you'll hear invaluable practical tips from innovative business leaders AI practitioners, and some of the brightest AI minds in our world today on how you can leverage AI in ethical ways to advance your career and grow your business.
1:10And now to something completely different. I always wanted to say that on the show, so now I did. The next topic is image generation. Now we all can benefit from good visuals, like whether it's as simple as creating a PowerPoint presentation or posting on social media or creating a blog or writing proposals or creating brochures and websites and so on. Now, some of us have people doing this for us. There's a design team. There's a marketing team. Some of us don't. But even those of us who have a marketing team or a design team or whatever, sometimes you just want to create a presentation, a PowerPoint.
1:41And you're like, okay, you can wait for them to create it. But if you know how to create it yourself, A, it's going to come out exactly what you wanted. And B, it's going to save you a lot of back and forth and it's going to save you a lot of time. Now, AI image generation is probably the most advanced AI capability right now from the perspective of being indistinguishable from the real thing, right? So yes, writing becomes very, very good, and it's very hard to distinguish AI writing from real writing. But especially if you're a beginner, you could still spot it out. Like, you can still see that this was generated by AI.
2:11I can spot a LinkedIn AI message a mile away, and I'm sure I see Rory smiling. I'm sure he does the same thing. but images are not the case. Like images today, if you know what you're doing, it's literally indistinguishable from real life. And when I say real life, it doesn't necessarily have to be a realistic image, but the thing that you wanted to create, whether it's a cartoon or the style of your company or whatever the case may be. But just like in everything else that we're doing in this episode today, there's so many tools, right? And you can do it in multiple different ways. You can do it in a chat.
2:40So in ChatGPT or Imagine or ImageGen, depending on how you pronounce it from Gemini, you can do it in a chat interface, or you can do it in more professional tools like Me Journey, or you can do it in some kind of an aggregator like Freepik or CREA or Weevy, or you can do it in open source like Flux, like there's so many freaking options to do this. And so which one do you use? Well, lucky for us, we have the number one AI image expert on the planet, or at least on LinkedIn with us today. So Rory Flynn has been creating images with AI since it was possible to create images with AI, but he probably creates every day more images than all of us on this live create together.
3:18That will be, I think, a fair guess, Rory, right? Yeah, ish. Yeah, probably. And so in addition to the fact that he's doing it, he's sharing literally everything that he's doing and everything that he's learning on LinkedIn. And I've been following him for a while and learning from him a lot on how to create images. He's literally the master Jedi when it comes to image generation. And he is also working in partnerships with a lot of the big brands and the tool companies, which gives him even more insights than just us common people to exactly which tools do what's better and how to use them, which literally I couldn't think of anybody better to teach us about the different tools and what are the pros and cons and which one and how to use them for different use cases.
3:55He's also the one that has been on the show more than anybody else ever. So Rory, welcome back to Leveraging AI. Awesome. Thanks for having me, man. I didn't realize that I have the title there of being on the show. You hold the title. You hold the title. As long as you keep coming back when I call you, I think that's going to be ongoing. But I appreciate you. I appreciate everything you're doing. and I know that there's a lot of people waiting for this particular session because it's, A, it's useful to a lot of business aspects, but B, it's a lot of fun too. Yeah. Yeah. And I mean, I use it now to like relieve stress, as weird as that sounds.
4:25I just like to open my mind and sort of brain dump into the image generators and see what my head has been thinking about all day. So yeah, it's really cool, but I think I went a little overboard on the presentation, so I'll try not to go over for everyone. I got really excited and I started going into it. Awesome. I'm excited as well. Let's dive in. Let's do it. All right. So let me pull this thing up. Make sure this is good. We can see this working. All right, cool. So, you know, I like to say the future is generative. It's just a relative, you know, term that I like to use. I just don't see in any other way how it's not going to be part of what we do in the future, whether that's for super high power content, super low power content, anywhere in the middle.
5:02I just feel like it's got to be part of a process in some way, shape or form, whether that's internal or external or just for yourself. So, you know, it's kind of how I'm thinking about everything at this point. You know, my name is Rory Flynn. I'm the founder of Systematic AI, a company you've probably never heard of. But essentially, you know, what we do is we look into people's businesses, we find holes, and then we plug those holes with the conventional AI tools, oftentimes just to simplify workflows and just make things easier, you know, on your teams and the people that are actually working on this stuff.
5:28So, you know, it wasn't always something that we set out to do. But, you know, at the time when we started using this back in 2020, late 2022, early 2023, working at a digital marketing agency, primarily in paid media, email marketing. And we had 90 plus clients and we were really awesome at sales and we were probably not the best at operations. So we could get a lot of clients in, but then once we had them, we had a lot of work to do and we could not fill that work fast enough. So if you've ever worked in an agency setting, I'm sure some of you have, we had basically insane creative needs, minimal assets and just like no bandwidth.
6:00So in the performance marketing side of things. If I'm running a Facebook ad, I might run 50 pieces of creative each week, three might work. So we have to scrap 47 and then build 47 new ones for the next week. So you can imagine something that would produce volume like AI would come in handy for us. So that's how we got really good, really quick at it. Now, again, we use this to really just amplify our productivity, reduce our costs per image that were created, and then also just generally created happier clients because we were able to provide better optimization and performance for their advertising.
6:28So a little bit different now. I work with a team. I partnered up with a team at Superside. If you're not familiar with them, please go check them out. 800 person team worked for some of the biggest brands in the industry. Goal there was really to operationalize AI for mass scale, like mass production of it. Push the boundaries, right? So that's what we've been doing. I'll show you a few things that we sort of worked on process throughout the time there. But again, stuff about me. I'm really here to talk about what you guys, who you are, entrepreneurs, innovators, just AI enthusiasts, probably using these tools in some way, shape, or form.
6:58But here's the problem, as Sar mentioned, like everything is moving at this crazy pace. It's moving faster than ever. I had this slide in here probably two years ago that was moving fast. It's moving faster now. So every day seemingly there's like a new model, new tool, new update. It's hard to sort of stay on top of it all. I have a problem staying on top of it all. Now, the problem with that also, too, is the tools are getting easier and everyone is starting to become more accustomed to it. So I think we're heading towards a point right now in the image space where you're going to see everything that looks pretty much the same, right?
7:28You have this sort of like big data play that everyone is generating off of. I see a lot of lazy creation out there because it is so easy now. You don't have to try. You can say, give me a picture of a car and it'll produce a picture of a car that looks relatively good. So I sort of mirror it to what's going on in the automotive industry. If you look over here on the right at these images, these are all different cars. To me, they all look exactly the same because they're basically designed on mass data preference. They come in four colors. You never see anything that looks different, right? You don't see a 1959 Cadillac DeVille.
7:58You don't see a 1962, you know, Ferrari, right? Like those are, that took a little bit of risk. So we're seeing this sort of homogenization of everything in the image space because now everyone can do it, especially with chat GPT. So it's how do we differentiate? How do we control? How do we stand out? I think that's really the sort of the way that we need to look at it. But, you know, in reality, if we're going to control these tools and get the most out of them, you just need baseline skills. You know, there's a million tools, but if you're good in a skillset versus an actual tool, you're going to be much better.
8:24So if you have a zone of genius, something you're great at, and some vision, you can really amplify this. Now, when I say baseline skills, this goes beyond images. I think I just like to throw this in here to make everyone maybe feel a little bit more comfortable with the state of everything. If you can use three tools well, you can use 100 tools well. So if you have skill with an LLM, most of them work exactly the same way. It's like the difference between an iPhone and a Samsung. They do the same underlying capabilities. They call, they text, they send emails, they have apps, but there's some little intricate details and you can learn those.
8:55But the baseline skill, you know, like I'm saying, if you understand chat GPT, you probably know how to use Claude. You probably know how to use perplexity. If you can understand mid journey as an image generator, you can use flux. You can use Leonardo. You can use chat GPT. You can use Korea. Same thing with video. If you know how to use cling, you know how to use VO3, you know how to use runway. They're all common. They're all built on a similar architecture and infrastructure. Right. And they're symbiotic. They work together. Like it's like a little team. You can use them for different tasks or to get to a common result.
9:20But, you know, typically, Basically, I think if it's something as if I want to describe an image or like build an image, I can start in ChatGPT. I can take that into MidJourney, create it in MidJourney. Then I can take that MidJourney image and go to Kling. Then Kling can animate it. I can extract a frame from Kling, describe it with ChatGPT. And this whole thing can be sort of cyclical and homogenous in nature. So I think it's just a good understanding. Just you don't need to have a million tools in your tool belt. You need to be good at certain skills. And then once you learn those skills, you're transferable.
9:45So when you start to understand how those things work, how it functions, and you can really start to push how it's leveraged and start to like really operationalize this stuff, because that's what we need to do to move this stuff up at a really, you know, scale this stuff at a quick pace. So let's transition this into the image side of things and what we're going to talk about here today. But, you know, really, I'm going to talk about like unique creative. Anyone can go onto an image generator right now in the chat GPT and say, do this, you know, okay, wait, change that. Okay, wait, do this. If you want to cut down that time and sort of get to a better starting point from there, from idea to first draft, I think that's where a lot of people are going to win right now.
10:17We're going to start with how this stuff works, as you can understand sort of how to get there quicker. So really, if you start to understand how the image generators work, what they do well and what they don't do well, and what works with them and what doesn't work with them, you'll save yourself hours and frustration. So like breaking down a lot of them, text prompting is still king. The same way you prompt chat GPT, you know, the art of the image is in the actual prompt. And I like to think about it like this, like a clear and direct prompt is going to equal a clear and direct output. An ambiguous prompt is going to equal an ambiguous output, meaning like the more direct you are in your prompt, the more you want to control, the more you will control.
10:49The more you leave up to the interpretation of the image generators, the more they'll take it and run with it. And it might not be sort of what you're looking for. Now, you'll see this a lot of times as a debate in the image space on should you go with short prompts or long prompts? I tend to think when you add more detail, you get better results, because that's just sort of the way that everything is in life. When it comes to anything like a client brief, or you're instructing someone to do something, if you just say, go write a social media content calendar, what are you expecting back? Versus if You say, I want a social media content calendar, three posts a week.
11:20We're going to focus on Instagram. We want to target this audience. The same way a chat GPT prompt would work, right? So something like this, you know, you might say, photo of a juicy hamburger. You'll get the hamburger, but it'll look like a stock image. But if you start to add more detail to it, you know, juicy beef, oozing cheese, bold and satisfying, you'll get something that looks like it's out of a magazine. So it's just a little bit more attention to detail on what you're saying. Now, when you understand how the image generators work, things start to become a little easier. um so like going back to this ambiguous nature if i asked everyone here to just picture a woman in the park if i you know i said like go you know put that image in your head picture a woman in the park and i gave everyone a second i can guarantee here that everyone has a different image of that woman in the park in their head because that statement's super ambiguous it's your brain's trying to fill in information that i didn't add what time of day is it what's she wearing where's the park and what does she look like what's going on in the rest of the scene all that has to be filled in because I only gave you like two pieces of information.
12:13Right now, when you start to add a little bit more detail, it makes sense. And it starts to shrink sort of that vision. So if I just said over here, woman in the park, so you can see we have different seasons, the woman's in different angles, you know, we have different lighting, it's like kind of hazy and gray, it's, you know, maybe it's a mid afternoon here or fall, this is spring. So you'll get this varying output because you didn't specify things. But if I said, you know, again, just like I'm talking to you, if I I say, you know, picture a woman sitting in a park at night during a cold night wearing cozy clothes and a single spotlight shines down on her.
12:44Everyone's vision collectively is much closer to what that image is than what just woman in a park is. Right. So the image generators work the same way, as you can see. So I do go over some some prompting basics. I think this stuff and some some basic techniques will get you to be good on any platform or at least have a fighting chance on any platform of not having to be frustrated to where you quit. So typically I start, and this is generally how I start to do things. These are not hard and fast rules, just suggestions. I don't use full sentences when I start prompt if I have a concept in my head.
13:15Each word is basically a token. And if you just write these long paragraphs, you might just have random words that have three or four meanings in them. And they might be sort of messing up the entirety of the prompt. So sometimes I just use cold or thick condensed style prompts. I like to call it, if anyone's familiar with concentrate or juice concentrate, If you have cranberry juice, you buy it frozen. So you have the concentrate, you add water to it over time. This is that concentrated version. You add water to it over time to sort of make the prompt. Everything is iterable. A lot of times it isn't just type in, get an image, great.
13:45It's a lot of iteration that goes into it. So I don't like to use full sentences, mostly just keywords and really building on top of thick, strong keywords. I like to separate this stuff by commas, not a hard and fast rule, just a way to separate ideas. More powerful language is going to help, just like in copywriting. Enormous versus big is going to be better. like vibrant versus like colorful is going to be, you know, going to tell a better story. So typically I like to focus on one thing to start. I wouldn't say like this person, this person, this person, this person, this person, like this person in this location doing this, like start there, build on top of it.
14:15You can, you know, you don't want to have so many variables that you can't go back and edit. So, you know, it's really like you're telling a story. It's just like you were talking to a friend, like you're setting a scene. That's how I like to think about prompting. But in my world, in the marketing sense, when we had to do this three years ago, we had to get photorealistic stuff, like the futuristic fantasy stuff was cool, we had to make things look real. You know, what we realized there was that, you know, maybe we have to start thinking more like photographers and start talking more like photographers.
14:39And when we put that into the image generators, they started, they started working very well. So, you know, the thing that we had to understand when starting this was like, what is a photo? Like what actually is a photo? What makes it up? What are its requisite parts? Because if we know that, then we can control it. So we essentially took a photo and then broke it down. Like if we're taking photography one-on-one, this is the sort of things you would learn about, you know, Subject in action. Who's in it? What are they doing? The environment, where it takes place. The composition and shot type, big part of it.
15:06A drone shot tells a radically different story than a close-up. The mood and emotion, what is the tone and the vibe around it, especially with branding. Specific cameras and lenses, this always depends, but I like to think cameras have different visual signatures. A big difference between a Polaroid and an iPhone look, so it's a good way to match a brand, a company, a style, aesthetic. Film stock can be another way to sort of differentiate there, but also lighting, lighting, color scheme, and then different details and modifiers. You can't have an image without lighting. If there was no lighting, you just have a black image.
15:34So it's like, these are things that have to be there. Like regardless of if you prompt for them or not, there's gonna be lighting in an image, right? There's gonna be a composition in an image. There's gonna be an environment in an image. You have to have it. It's a non-negotiable. So those are the things that we were like, oh, this is what we can control. Now, when we think about it that way, again, I like to think about these like visual building blocks. You have each portion of that image goes in, or each portion of that category goes into makeup and image. Now, again, if you want to just say you don't like the color scheme, now you can also then go swap this out.
16:04Great, I don't like that it's red and white. Let's make it blue and green. You can go identify it in your prompt very easily without having to go and redo the entire thing. Now you know that that's sort of where you have issues. So we took those visual building blocks and then we built it into a prompt formula. And this was sort of like our superpower in terms of operationalizing everything, right? Like this little piece, like breaking down the photo and turning it into this prompt formula. we worked off it a million different ways. So basically we just took all those elements, structured it into a prompt formula.
16:31So we have something like phototype, subject action, environment, color scheme, separated by commas, right? And you get something like this prompt below where it's just a collection of random words, which is what you might be seeing. But you have motorsport photography, Red Bull F1 car, driving on a racetrack, deep azure, blue, red, yellow colors with warm tones, 35 millimeter shallow depth of field, dramatic sunset backlighting, center framing, motion blur. Like I said, sounds like a bunch of random words thrown together. But, you know, when we run that, we get an image and sort of taking all those ideas and diffusing it together.
17:05But it's not as random as it seems because we're controlling the controllable elements here. You know, you'll see we have, you know, more motorsport photo. It's in that style. We're on a racetrack. We have the warm tones. You have that golden glow. We have Red Bull F1 racing car prominence driving. You can see that as the tires are spinning, you know, azure, blue, red and yellow. We have our blue, red, yellow. We have our 35 millimeter shallow depth of field, meaning things in the foreground are in sharp focus. Things in the background are blurred. We have center framing. It's perfectly centered right in the middle.
17:34We also have our motion blur. So you can see it looks like the car is moving to the right as the motion blur goes to the left. We have our sunset backlight, which is coming right from behind the car, which is exactly what we asked for. So when you think that you don't have control of this stuff or it's all random, you actually do. It just might not be controlling the right elements. So these are the things, again, you start here, not saying that's the only things you have to cover, now you can start to build it out. You can build it out further. Now you have the core idea, right? So solving the knowledge gap, though, was a problem for us as well.
18:00Like if we're going to tell our team how to do this and how, you know, to really get everyone on the same page, who knows all these photography terms, right? So that's where we brought in simple things, ChatGPT. Again, because we have that prompt structure, makes things a lot easier. Very simple formula for ChatGPT to follow. So, you know, just give it like, let's get five prompts for the Red Bull racing car on a track, right? We want to fill in this prompt formula. So it's basically just give it the prompt formula and use concise, visceral and powerful language, the same as our prompting basics.
18:28And then we just generate a prompt. So now we can create five iterations of this rapidly and sort of just like curate versus create all the time and go from there. And naturally we built this into custom GPTs and agents and things of that nature. It gets bigger once you understand sort of the core functionality of how to do that. Right now, the next thing was sort of solving consistent style. Once we could generate images and then we could generate whatever images we wanted to. Realistically, it was working in the agency world, how do we mimic the brand style? So we call this asset hacking. It's just for lack of a name at that point, which sort of made it up.
19:00So basically, it's reverse engineering. We had consistent imagery. I want to get consistent imagery, brand relevance. I try to tell people not to use this on our... Only use this on their own brands. But the goal here was to essentially break down an image into that style of prompt formula. So turn an image into text and then reconstruct it. Now, the analogy here would be like if I was thinking about a pizza. If I took a pizza and broke it down into its actual ingredients, what is it? You have dough, you have mozzarella, you have tomato sauce. Those are three ingredients. Once you've broken it down into those, you can rebuild it into a calzone, for lack of a better analogy.
19:32Take the ingredients, cook something else. It's just the same idea. Now, we would take brand assets, then we'd use that prompt formula again. It became very handy. We'd turn it into text, and then we could rebuild that image a million different ways. So instead of an F1 car, we could make it an F1 dune buggy, keep everything else the same, keep the look and feel the same. Right. So it'd be like, again, taking this F1 car, turning it into these text images or these text, uh, text prompts, and then rebuilding it. So everything looked consistent and in the style. Now that's a lot easier. There's things called style reference and, you know, you Laura models back then we were doing one, you know, one show, one trick pony, which was mid journey.
20:06That's the only thing that could do this. Now, why this is important. Those three skills, sort of understanding the image generators at a top line level, understanding how to prompt them, and then understanding how to reverse engineer, you realize everything's data. This can be translated across anything. You can do this with an email. You can do this with a video. Whatever it might be, you just need to understand the requisite parts so that you can go and rebuild it, right? Now, once you, again, once you understand this, it starts to become more, like, you start to get a little bit excited because you realize you can do just about anything.
20:34But currently, you know, there's just so many tools, there's so many uses, and there's so many decisions to be made on which tools to use. I think it's really about just limiting it again and coming back to this idea of if you can be good with one tool, you can start to understand how the other tool works. So a lot of the top image generators now, I'm leaving some off because I just had to. There used to be two, now there's 20. But these are the ones I probably use the most consistently and you have explored the most. So mid-journey, number one. ChatGPT has gotten very good. It's very, very complex in what it can do.
Read the full transcript
21:03Flux is a universal model. It's used in a lot of APIs, but it's also very trainable. We'll go through a lot of these. Reeve is a new one on the block. really good quality, sort of still under the radar, in my opinion. Imogen 4 in Gemini, it's actually a really good tool if you use it right. Runway, to me, has really stepped up their image game, and I use them for a lot of things as well. Ideogram is definitely pretty solid, coherent, does everything you want, not really super flashy, but good. Same thing with Leonardo. It's hard to go through each one of these here, but I'm going to try to give the best sort of understanding of what this looks like.
21:37Now, things have really changed. I used to have to go through all these crazy text prompts to get something good. Now I can say like, do this. And these image generators can understand, especially with the integration of ChatGPT into a lot of the tools as well. So three months ago, it wasn't possible for me to give a model three pieces of clothing and say, put it on this person. Now you can do that. So even down here, you can see you can do this with MidJourney, you can do this with Gemini, you can do this with Cria and or ChatGPT. And then you can do this also with, with Broadway my references and now flux context.
22:07So three months ago, super hard, super technical. You had to train lower models. There was a lot of editing involved. Now it's relatively simple. Please put these clothes on this model. And it's kind of just, again, comes down to your vision. But the problem that I'm seeing here is that all these tools are good now. They're all good at different things too. So it's sort of cherry picking which is to use for which sort of use case. Thing is they all speak different languages. Some like mid journey, like broad concept prompting, meaning like just give it short ideas and let it run with it. And some like Flux want like full, long, detailed sentences writing out every little thing.
22:40So, you know, it's like, how do we sort of bridge the gap here to be good at all of them? So let's, you know, get into that. But I also just want to show you sort of like how you can choose which model you need or how to sort of like test which model you might need. I kind of break them down into these categories. This is just my own sort of way of thinking about them. If you have your creative models, which are going to be like your playground, the ones that you can just type something in like, you know, a woman in red dress. Hit enter and you get something that's visually stunning. Those are your creative models.
23:05Coherent models is when you want to control a lot of different elements of the image. If you have maybe doing brand work or you're doing stuff for your business, that might be a little bit more useful in that sense. You have your trainable models, which I consider like your tailors. If you want to have what's called a Laura model, if I wanted to take this hat and train a model on it so I could reproduce this hat a million different ways with precise output, that's going to be more of a trainable model. Then you have your editing models, which is basically like you take this hat and put it on a dog or something like that, which is pretty simple, low lift, easy.
23:35But I'll give you some examples of this. So the tests I like to run to understand if these tools are creative models or coherent models. Again, this is my own terminology. I don't even know if this is really discussed that much out there. But I'll give it a very short, basic prompt like we did, Woman in the Park. If you give it something like that, let's see how it runs with it. So if you get more detail, you probably have a really creative model on your hands. If you have less detail, probably have a more coherent model, meaning it needs to be directed more, but probably means you also have a lot more control.
24:04So like this picture, this prompt, for example, just photo of a samurai. These are two examples from mid journey. Mid journey is by far and away the most creative model, just taking samurai and turning it into something like that. They're like, that's as easy as it gets for the quality. Now, when we run this on various models, you'll notice the difference. So you have mid journey, super detailed, super close up, vibrant color. It wants to be creative. It's sort of like your creative director that you need to reel in a little bit sometimes. They have all the ideas, but maybe not as much structure.
24:33Reeve, I think it's an all-around solid tool. Great visually, it does what I need it to do. Flux, as you can see here, as we start to go down the seam, I think the one thing you'll notice is a lot of the backgrounds. That's the one thing that's sort of not talked about a lot in the image generation space if you're looking for realism. Background plays a huge role because Because the further you go down this list, the less background is sort of prevalent because there's so much generating power is focused on the samurai. Like it needs to divert some of its assets. Mid-journey can do that typically very well.
25:03Foreground, mid-ground, background, asset allocation in that sense. But you see here Reeve, like decent background, still looks pretty AI to me. This one, you start to get the blurred background because Flux doesn't want to necessarily pull all that generation power into the background. It wants to focus on the samurai. Imogen, it sort of looks like a cinematic style view, but really not much going on in the background. ChatGVT just says, screw the background. And then Runway is like, let me put all this cinematic effort into just the Samurai. So they all treat it a little bit differently. Now, if I did the same test and just said, give me an editorial photo of a sports car, again, you get mid-journey with the crazy detail and just a bold image.
25:37Reeve does exactly what you tell it to do. It's got great lighting, great color, super saturated in that sense. Love it. flux can handle it but you can see like the road is very unrealistic the background to me is super very unrealistic but the car details are there imogen same thing and you notice like flux imogen chat gbt they all have the same position of the car it's all very like standard and generic they probably have the most data in a sense as to what they're doing and this is the most like likely outcome when you say something like a sports car so chat gbt to me really has the least amount of detail here because chat gbt is probably the most controllable in terms of like do this this this this, this, this, this, this, this.
26:14So it's like working with a blank canvas, if that makes sense. You're going to have to control every piece of it, which is why prompting becomes so important. And if you have a structure, you know how to do it. Now, again, if you look at these all side by side, right, you have your mid journey. Obviously, to me, I'm not trying not to be biased, but it's just the one that jumps off the page to me in terms of like visual stimulation. That's the one. Reeve to me looks like a really solid in the middle player. Like I want a little bit of creativity, but I also want you to listen to me. Flux, I can just tell it's going to do whatever I say.
26:41I just have that feeling by looking at it. You know, chat GPT, there is a certain look to it unless you control otherwise. It has a very yellow and greenish hue. You'll see it show up in a lot of images. You have to be able to understand that and control that because it wants to be muted and it wants to be dark. You'll notice in the next example we do also. But, you know, Imogen, Google's super wide ranging. Again, you know, it has a massive data set, so you're going to have to be precise with your problems. Now, like the creative models. I'll close it just for one second to add one small thing for people, which is to me, one of the other big differences between these tools, and maybe you're going to get to it, but it's something that I alluded to in the beginning, which is how you interact with them, right?
27:17The biggest benefit from my perspective of Imogen and ChatGPT is you can chat. You can literally go back and say, oh, I really like this aspect, but make it brighter. And then it will make it brighter, which is not something you can do with any of the other tools that are in here as well. So there are, the way you engage with these tools is also different. They have different user interfaces that provide pros and cons. Again, just like we were saying, there's no good or bad. The sliders and all the features in MidJourney are amazing, but you can't just tell it in English what you want to do. And there's pros and cons to each and every one of these sides.
27:49So that's the other aspect of this. Beyond the backend and how they work, the front end is also very, very different. So just to call that out now too, MidJourney added a feature, which I find that I'm using way more than I used to. It's called conversational mode. So it works very similar to ChatGPT now in terms of like, I have this idea, here's a rough structured idea, like make it. And then it makes it and you can edit the same way too. So there's a new function within there. Works very well, similar to their voice mode now, which is also great. Just like create this image. I'll add, I'll one up you on that as well, because I think this is awesome for people to understand in the context of this specific episode is the biggest difference though, is that the mid journey conversation will be around mid journey.
28:27And the conversation around the image and the conversation in ChatUPD could be, I'm creating this business presentation for the CEO of this company. And it knows all of that when it comes to create the image, right? Like the width of the conversation could be a link to the website and PDFs that I've uploaded and the deep research that like all of that goes into the context when I come to create the images, which gives me much more contextualized images to what I need versus me having to understand what I need to do in order to make it that. So again, not taking anything away from me, Journey. I still think like you're saying, from a creative perspective, it's the best tool out there.
29:01A hundred, definitely. And then there's no comparison of what ChatGPT can do with text and what Majority can do with text. ChatGPT is so far ahead in that regard. It's not even a game. So that's the one thing with these creative models. You can do more with less. I can say an artful skull hanging over a table and get an image that looks like this. Again, that might not be as realistic with some of the other models. Just needs a little bit more polish sometimes. They might be a little bit out there or a little bit too much. So the coherency side of things, too, I think is important because this is one of the biggest tips that I can advocate for.
29:34If you're familiar with the JSON style of anything, this is how I like to prompt when I want to prompt all models. They all respond to this, which is basically breaking down, maybe take the prompt formula that I used beforehand and just fill it in that way, utilizing where it says subject colon this, environment colon this. When you put it that way, it really limits the freelancing that it can do and it reads it better just like a computer would read, you know, like an operating system would read it. So a lot of times when people think there's not a lot of control in these tools, they need a certain prompting style for it to work.
30:09Something like this, and you can have ChatGPT do this for you. Here's my idea structure in a prompt just like this. And it'll spit it out. And when you do that, right, like I put this really long prompt with a lot of elements in there. And as you can see, it comes out pretty much exactly how I want it on every single image generator. So I think there's like, again, there's a common misnomer here that you don't control this stuff and the model really limits you in certain ways. And a lot of times it's just controlling the controllable pieces of it, putting it in a way where you want all of these things and then having it sort of structured in a way that all of them can read it.
30:41It's like a system prompt. When you use that, you can also start to build more things out, which we'll get into in a couple seconds here. But the coherent models, going back to this, Flux, definitely coherent, ChatGPT, Imogen. They listen, but it's like I said, it's more like a blank canvas, you're going to need to add the detail or explain the detail. It's not just going to do it for you. Right. Um, you know, so really when there's the malleable models and I'm trying to, trying to, because there's so much here, I'm trying to give you a little bit of a taste of everything. Flux is a malleable model.
31:07I think we have like two minutes, so we have time for the video stuff as well. So quick recap, uh, and, and then we'll be good, I think. Yeah. So again, you now have editable models in the way that you, uh, in the way that you can just say, Hey, do this, add a knife, put sunglasses on, add a hat. Those are more simple. Something like ChatGPT, Gemini, RunWake, and Flux Context can do this very simply. But really, taste is everything now. So last thing, I have a bunch of stuff here. I'll probably just send this to you so you can send it out if you want to. Really, once you start to... This is like, again, I was going to show some visual examples of everything here.
31:41This is, again, if you're not familiar with MidJourney, great at style, great at storyboarding, great at sort of character development, things like that. You have Reeve, which I think is really great in sort of an editorial context with really good text capabilities on the image, really good detail, texture, things like that. Flux, again, you can train custom models. So you can just take something like images of a car and just only replicate that car. You can also take sunglasses, put them on models, things like that, very simply. now again gemini we have a few things here really wanted to get to the chat gpt portion of this because it's very creative in that sense where you can utilize things like you know sketch to image i can draw this out get a great image don't have to always prompt you can say this is what i want i can give it a composition say fill in the composition with x very easy character consistency is now great with chat gpt just give me another like look at this character the sketch to image stuff you can take it and really push it in a far way or you can think about in terms of structure where, again, you have maybe an ad structure.
32:39Here's your headline, your sub-headline. I would create this, and then I would say, fill it in with X. And that's how we can start to build a template. And now I can do this template over and over again with different brands or different images, things like that. So if you understand sort of the template side of it. Where's one last... This was the one thing that I wanted to get to by the end of it. Regardless, this is this new tool. I'll wrap this thing up. Sorry for missing it. But this is Weeby. So Weeby is basically, once you know how to use all of them, you can take them into one space and start to build.
33:09You know, something like this, if I was to show you what the entire workflow looked like, this is like a virtual try-on generator where we have, we describe a model with a prompt, we use the reverse engineering to sort of develop, you know, to describe the close, then we can build a virtual try-on, then we create an entire product catalog using Gemini to then just iterate that into many different shots. We can take this and then animate the keyframes, then turn them into videos. So this is all an animated or automated systematized process. Once I put the close in and design my model, I can push run and all of this happens.
33:40I can do that on the next one. This is, you know, I tried to fit as much as possible into here. So my bad on running. No, this was, this was incredible. I think this was incredibly valuable to people who are just kind of like you're saying, oh, I can put on a prompt and I get an image and that's good enough for me. And this really takes it to a whole different level of understanding. and like you're saying, there's the big universe of automation layer on top of that, which a tool like Weavey can take an idea and turn it into 50 assets instead of one or four and do it at scale while having different checkpoints along the way.
34:13But overall, this was amazing. Again, for people just getting started, this was so much valuable information. I can't thank you enough. If people want to follow you, work with you, learn from you, what are the best ways to connect with you? Definitely on LinkedIn. I'm on X. I don't know if I'm, I feel like I'm lagging right now a little bit. A little bit. But also my, yeah, my buddy and I, Drew, do a podcast every week where we either one dive into tools and demo them and play around with them or just talk about what's going on. And a fun way for us to sort of hang out and engage a little bit.
34:42So if you guys are, you know, LinkedIn, YouTube, X, whatever, I'm around. Awesome. Perfect. Thank you. This concludes part three of the Ultimate AI Showdown. Part four is coming up next week in which we're going to cover and compare AI video generation tools with the amazing Tianyu Xu. And in the episode, he's going to show you the different AI video generation tools that exist today and what are the pros and cons of each one. And until then, have an amazing...
From the publisher
👉 Fill out the listener survey - https://services.multiplai.ai/lai-survey
👉 Learn more about the AI Business Transformation Course starting August 11 — spots are limited - http://multiplai.ai/ai-course/
Can you really trust AI to create brand-worthy images — without a designer?
If you're leading a business, you know how much time and budget visuals can drain. But what if you could generate high-quality, on-brand images in minutes — with tools you already have?
In this power-packed episode, AI image expert Rory Flynn joins Isar Meitis for part three of the Ultimate AI Showdown, diving deep into the tools, techniques, and secrets behind top-tier AI image generation.
Whether you're bootstrapping your visuals or leading a team of marketers, this episode is your practical guide to making AI your creative advantage — not your creative liability.
In this session, you’ll discover:
- Why AI-generated images are often indistinguishable from real photography — and how to use that to your advantage
- The key differences between top AI image tools like Midjourney, ChatGPT, Imagen, Flux, and more
- How to choose the right AI tool for your business use case — from social posts to product catalogs
- Rory’s "visual building blocks" framework to prompt like a pro (and never start from scratch again)
- What every leader needs to know about creative vs. coherent AI models
- A behind-the-scenes look at how Rory creates entire branded visuals — at scale — using automation tools like Wey
- The exact prompt structure Rory uses to control AI image outputs and speed up iteration
- How to reverse engineer your brand’s visual style for consistent, on-brand image generation
Rory Flynn is the founder of Systematic AI and one of LinkedIn’s go-to experts in AI image generation. From working with global brands to building full-scale visual automation systems, Rory’s insights are redefining what’s possible in modern content creation.
👉 Don’t miss Part 4 of the Ultimate AI Showdown coming next week, where we break down AI video generation tools with Nu Shoe.
Until then, go turn pixels into profits.
About Leveraging AI
- The Ultimate AI Course for Business People: https://multiplai.ai/ai-course/
- YouTube Full Episodes: https://www.youtube.com/@Multiplai_AI/
- Connect with Isar Meitis: https://www.linkedin.com/in/isarmeitis/
- Join our Live Sessions, AI Hangouts and newsletter: https://services.multiplai.ai/events
If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!



