Kling O1, Seedream 4.5, Z-image: AI's Biggest Week

4 Dec 2025 · 37 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown Denoised Podcast Summary

Episode Title

Kling O1, Seedream 4.5, Z-image: AI's Biggest Week

Hosts

  • Addy Ghani - Media Industry Analyst
  • Joey Daoud - Media Producer and Founder of VP Land

Episode Overview In this episode, Addy and Joey discuss significant advancements in AI models within the film and media industry, including:

  • Kling O1: A new multimodal model for video modification.
  • Z-image: An open-source alternative emerging in image processing.
  • Seedream 4.5: Notable for its unique batch generation capabilities.

They also compare updates from various models, including Runway Gen-4.5, LTX Retake, FLUX.2, and TwelveLabs’ Morengo 3.0.

---

Key Discussions

Kling O1

  • Described as a powerful video modification tool with capabilities akin to Runway Aleph.
  • Supports input of existing videos to modify elements (e.g., changing objects, scenes).
  • Outputs are sharper and more consistent compared to its competitors.
  • Can handle videos up to 200 MB and resolutions around 1080p.
  • Capable of compositing images into videos with high accuracy.

Z-image

  • Positioned as an open-source alternative with user-friendly modification capabilities.
  • Fully open-source, making it accessible for customization.
  • Three variants available:
  • Zimage Turbo: Distilled version for consumer-grade hardware.
  • Z Image Base: Foundation model.
  • Z Image Edit: Focused on image editing tasks.
  • Offers photorealistic quality and bilingual text rendering.

Seedream 4.5

  • Improved consistency and editing controls for images.
  • Unique batch generation feature allows multiple images in the same latent space, enhancing character generation consistency.
  • Stands out for its user-friendly interface in generating high-quality results.

Other AI Updates

  • Runway Gen-4.5: New flagship model aimed at video generation, but facing criticism for user output consistency.
  • LTX Retake: Introduces a feature for modifying specific sections of existing videos.
  • FLUX.2: Continues to evolve with promising image modification capabilities.
  • TwelveLabs’ Morengo 3.0: Enhances video indexing and understanding capabilities for easier searchability of video libraries.

---

Key Takeaways

  • Innovation Pace: The AI landscape is rapidly evolving, with new models emerging that significantly enhance video and image manipulation capabilities.
  • Market Implications: Open-source models like Z-image are gaining traction, emphasizing the importance of accessibility in AI tools for creators.
  • Quality vs. Usability: While many models produce stunning visuals, practical usability in professional filmmaking pipelines remains a challenge.
  • Future Considerations: The hosts speculate on the continued advancements in AI and the potential for future breakthroughs that will improve fidelity and integration into production workflows.

Audience Engagement

  • The hosts encourage listeners to share their experiences with the podcast and to leave reviews on platforms like Apple Podcasts and Spotify to help increase visibility.

---

Conclusion The episode highlights a significant week in AI advancements relevant to the film industry, with hosts providing insightful analysis on the latest tools and their potential impact on creative workflows. ```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28All right, welcome back to Denoised. So yeah, thank you for not being here, Joey. Thank you. Thank you for the consideration. All right. So busy week. I don't know. What do we say? Christmas came early from the AI companies. It just keeps coming. They really dropped them all after Thanksgiving. And every time we've kind of been delaying recording the episode. And then I'm like, ah, but like it's a hot topic. And then a new model comes out that morning. So you're so smart on your timing. I was like, let's go today. Joey's like one more day. Let's see if he comes out. Wait for it. All right. So I think we got a bunch, but I think the biggest one, especially the one I've been seeing on my feed a bunch, is Kling's new multimodal model 01.

1:0501. Omni-1? Yeah. Kind of being described as like the nada banana of video. This one, multimodal understanding, you can give it an existing video, tell it what you want to change. Similar to Runway Olive, but the outputs I've been seeing from this look a lot better, a lot more consistent, a lot sharper. Yeah. What is it? Is it native 1080p or 2K? What are we looking at resolution-wise? You can upload a video up to 200 megabytes, 2K. I think the output's going to be 1080, but I don't see a confirmation. But I think it'll be, I mean, it definitely would not be more than 1080. I mean, 1080p feels like the bare minimum in today's video generation world.

1:44Yeah, for sure. I mean, and if it's like a true 1080, because even with VO, it's sort of like, if you work in flow, it's like a 1080, but you have to up-res it. So it's really 720 and you're kind of up-resing. A true 1080 is good. Yeah, and I believe, I could be wrong, but I believe most of the diffusion models, image and video included, are square kernels. So it's either like 1024, 1024, or, you know, 1280, 1280, whatever. None of the video models natively run at 69 aspect. Okay, yeah, so they're kind of just faking it. Yeah, they're probably generating something bigger and then cropping the top and bottom to give you that cinematic aspect ratio.

2:20All right, so yeah, some of the examples I've seen. So like really you can give it a video similar to what we've seen with Olive, but much sharper. Remove objects, change people, change locations, change scenes. I've seen another interesting trick where you can give it a video input with an overlaid still image. So in this one, they kind of gave it a video. Just like Nano Banana. Yeah, but in this case, it's actual video. So you can give it this video input and then say composite this car and change the style of the landscape to look fuzzy. And so we get this Klingo 1 output with the car, composited, driving on the street, super accurate.

2:53And this one, they did a test against Aleph. And you can see Aleph sort of kind of just did some random, like sort of kept the scene, but did a bunch of random stuff. That's interesting that this is happening. As far as the entire ecosystem of all the models goes, O1 is being compared more to Aleph than it is to VO3.1 because it's essentially a video-to-video model. Yeah, and I think Olive is the only thing that we've really had so far where you can give it an existing video and through a text input or image input modify elements of that video. There hasn't really been anything else that can do that.

3:27There was sort of the VO3. You can modify the video, but that only worked on a video that you generated in VO and then you wanted to change later. You couldn't just upload a video and change it. Right. Cling, this one's the only one outside of Olive where you could. I'm sorry, I will also, in fairness, Luma Ray 3 Modify video. You can also do something similar, but I had similar outputs to Olive with that, where it would just kind of sometimes do it, sometimes make it a bit too mushy. Yeah, so it seems in our world of film and television, the most applicable models are obviously going to be Kling 01, Aleph, Luma Modify, and then VO 3.1 for any type of text-to-video, image-to-video.

4:11Yeah. And I mean, well, maybe we got another update from Kling that also dropped today in addition to Kling01. Yeah. I mean, well, this one is this is a demo video from Martin LeBlanc at FreePick of a Kling hype video that they did. And so you can kind of see this sharp detail with different character reference images, character performances. I would say you could also try this as a kind of, I'd be curious, like for character performances, how this compares to WAN 2.2. Because you can kind of do that same workflow of have a live actor human performance and then give it a reference image of another character.

4:43It sounds like a good comparison because I also haven't really seen, I've seen, obviously, I've seen a lot of Olive comparisons or Night of Banana comparisons. I haven't seen any WAN 2.2 anime comparisons. So it could be good for that. You're giving me ideas for my next tutorial video. Thank you for this. Yeah. I just want to, this is mind blowing how good even the demo videos are, because I don't know if you remember when we started this podcast about a year ago, having temporal video consistency was a big deal, right? Like as you have like a face and the camera rotates around you, just having the face be a face as the camera is moving was challenging enough.

5:20And now that problem is totally solved. And so we're on to the next part of the problem, which is fidelity and resolution and quality. This looks good. I would love to know, though, with all these hype videos, I would love them to do a video and have like a counter of like what number generation this was. Cling I'm a little maybe less suspicious of, but when we get to Runway, Runway's output sometimes, I find I have tough, tough time replicating the results that they demonstrate. Yeah, yeah, no, we could go off on Runway. I think we have the Gen 4.5 release on our AI roundup. Yeah, yeah, it's coming up later.

5:56Yeah, so we'll talk about Runway then. Yeah, I've heard Runway, the 4.5 has felt a bit like the meme of the pool and the runway was the hype for the day. And then now it's kind of at the bottom of the pool. Oh, it's so brutal being an AI company nowadays. I mean, your model hits the market within six hours. It's outdated or trumped by something else. It's crazy. It's good for us because, hey, we can do two of these episodes. We can not have enough time to cover everything. I know. Okay. Other cling thing. Cling 01 wasn't the only news this week. This morning that we're recording this, cling 2.6 dropped.

6:38And this is, let me look at the inputs. But the big thing with this is it could do audio, so it's kind of a bit more in line with VO3. VO3.1. Yeah, generating audio with the outputs. Let's see what else we got here. We've got some demos from Thal of character performances. I will say the voices sound pretty robotic and not that great in this astronaut clip. Tell my family I love them. Oh, the stars are beautiful from here. That's the uncanny valley of audio, Joey. Yes, not not that great. So making making a similar comparison, then would you say that one 2.2 animate and Omni one are more or less one to one and then one 2.5 ring cling 2.6 are one to one.

7:24oh that's a good that's a good comparison i think from the quick look that i've had yeah i think one 2.5 and cling 2.6 seem to be kind of on on par one 2.2 as well as v03.1 as well as v03.1 yeah uh one 2.2 animate i mean that can only do one specific thing whereas cling 01 you could do this character performance transfer but you could also just modify your video or tell it what you want to change and stuff so also you could throw a bunch of references into the video generation you could do references yeah so i'd say klingo one is like much more powerful than one 2.2 the other thing is one 2.2 is limited to like 720p resolution whereas we are not 100 sure but we think kling is 1080 it's it's interesting like late 2025 we're seeing the video models almost hit a fork in the road and go in two different directions so on one direction which is the main direction you have vo 3.1 cling 2.6 1.2.5 runway gen 4.5 which is the best image to video you can do the other direction is take an input video then modify it and then output a new video and that's luma modify aleph and now cling 01 and 1.2.2 animate because i mean the modified video i mean obviously that's has a lot of practical applications but it's still you know like you're basically taking whatever you shot your video at and then you run it through the model Now you're kind of compressing that video into whatever the model spits out, which we know since they have not hyped it, we know it's going to be like a 720, maybe 1080, 8-bit, super compressed video.

9:00So if you're a professional filmmaker or in a professional VFX pipeline, still not there yet. Obviously, as usual, this is the worst it will ever be. It's totally, totally right. Right. Yes. It's like, don't fret on that even for a second, because six months from now, we can completely overcome this hurdle. I think the building blocks are all there to do the video-to-video, quote-unquote, VFX. The last mile problem still persists, which is taking the output from the variational autoencoder, essentially what AI creates, and then making a bit depth resolution color space fidelity high enough to where you can feed it right back into a traditional production.

9:47And have the flexibility that you're used to because it's like you make the shot and then it's got to fit in the pipeline and then you got to color correct because you got to make it match every other shot. And you got to composite or blend or, you know, other traditional tools to bring it all together. Totally. And I think we're, I mean, I have not tested Ray 3 yet. I think Ray 3 is a good move in that direction, but I'm sure Google will have an answer for that, right? Sooner or later, Vio is going to have some type of HDR capability, maybe past 4K, and then perhaps having the ability to output in raw color spaces, right?

10:21So you can have way more color adjustability in the post. And also just like one final note with Kling, so it could do, it's a text-to-video or image-to-video input. So usual inputs you have there, but it's kind of like if you want to have more controller modifying video, then you would go Kling 01. What's next, Joey? Next update. Ooh, this website looks like it's from 1998. I don't know why we have this one to say. GigaZine. The next update is from Alibaba, ZImage. ZImage, yes. Image model. What do you got about this, Addy? ZImage is being perceived as the next SDXL. And what I mean by that is it's completely being embraced by the creator AI community as an open source, easy to adjust, easy to modify model.

11:03If you remember two, three years ago, SDXL was the darling of the creator community because you can easily build loras for it, publish those loras on Civit AI, or perhaps completely fine-tune, fully train the model. And it was a basis for a lot of products back then. So SDXL was the basis for MidJourney. Before MidJourney had their own model, a lot of the SDXL folks went over to Flux, and then Flux was born. And so Zimage gives you that same level of capability and tuning. It's fully open source, but with today's quality standards, right? So SDXL is no longer relevant in the quality standard of today's market.

11:45But Zimage is, you know, if you're going to go neck and neck with, let's say, a nano banana, then this is, I think, in that realm of quality. Yeah. And they've got three different variants. They've got Zimage Turbo, a distilled version that can run on more consumer grade hardware. And also all of this stuff, there's Comfy workflows. And you could run this on Comfy locally on your computer or Comfy Cloud or it's on every API. And that's I think the main appeal is that it's super lightweight, right? And it can run on lower end GPUs. It can run locally on Comfy. So you can go to town on modifying this on your own.

12:19Then they got Z Image Base, which is the non-distilled foundation model, and Z Image Edit, which is for image editing tasks and more in line with like a nano banana kind of editing. But yeah, it's really good to see kind of the showcase that they have on their launch page. The, you know, as you were talking about, the photo quality of the images are photorealistic, like much more up to date with other higher end locked off models. Proprietary models. Yeah. Stuff that you don't have the weight information. So yeah, the quality here is excellent. uh also accurate uh bilingual text rendering so english and chinese hopefully other languages as well yeah eventually so yeah i mean the demos we see here it's like the demos in itself it's like i've seen this is not a banana or other things but the power here is like this is a model that is open source you could download it you could run it you could build it into whatever workflow you want you modify it yeah so it's a very powerful model yeah and um folks like uh you know FAL, RunPod, Replicate, FreePig, like the big AI aggregators of the world have already ingested this into their ecosystem.

13:26So you can go run Zimage today if you like on most big platforms. I'm sure ComfyCloud is probably going to have it soon if they don't already. I think it's there already. Yeah. And I know I've been seeing some LORAs and stuff pop up for this already. Yeah. So yeah, this is a good step for open source. One of the more interesting workflows that I've seen in Z Image is not generating a base layer, but doing an image modification where you're adding detail to an existing generation. So we talked about AI sheen, plasticky skin, and that kind of stuff before. I've seen Z Image workflows where you can put an AI image through it, and then it comes out on the other end having much more realistic skin.

14:06Sort of like an upscaler-ish. I think much better than an upscaler, honestly. Okay. Yeah, because it's not just pixel resolution. it's actually changing the image. Okay, that's a good workflow. Yeah, I've seen similar workflows with like Nada Banana as well of like sort of upscaling and slight style changes. But yeah, to have that. I mean, also that sounds like you could find or build a LoRa that just specifically does that with ZImage. Yeah, that's exactly what it was. It was a LoRa that plugs into ZImage's main model for that. All right, next up, going back to what we talked about before, runway.

14:41mm-hmm gen 4.5 so this is a new update to their foundational model and then we've got the hype reel here which as we talked about i mean look the reel looks great you know super clear coherent shots some of the shots do fall apart but yeah most of the shots are fantastic i would just love to know what how many generations each one took to get 70 000 joey this isn't a coca-colle ad so now to clarify with this too this is their video model so this would be definitely text video i assume also image to video they haven't released this model yet you can't use it yet they just sort of i think it's coming out this week for on the platform but this isn't this is separate from olive so this doesn't have anything to do with being able to modify an existing video this is more just pure video generation yeah this is i would consider their flagship product in runway right now, probably their most state-of-the-art.

15:36I did see a couple of podcasts that Cristobal was on right when this launched. And I'm always fascinated by CEO of AI companies, as you know. We've interviewed a few of them. What is their internal clockwork? What is motivating them to build the things that they are building? And it's fascinating what Cristobal said. He essentially made a one-to-one comparison to what runway is building to a world simulation which is kind of like the cosmos nvidia model or the google genie model and it just happens to do media and entertainment now but eventually it'll do everything what are your thoughts all roads lead all roads lead to robots yeah we've talked about that before we're like all these world models are so the robots and cars can understand the world because film production applications are like that tiny little bit of the of the of the addressable market i i appreciate the ambition and the vision of these ai ceos like they certainly need it in a space that's very ambiguous having said that runway you guys still haven't cracked film and tv so don't just move on yet like the stuff is not quite usable would you agree yeah i mean unless you're doing some insert shots or something i've found more just from a user experience part like it takes me more generations to get something usable from what I'm looking for from runway than it would with like VO or C dance.

17:02And so I tend to just rely on VO or C dance first, because I know the outputs are, I'm more likely to get what I'm looking for out of like a few shots and runway don't use as much. Yeah. I mean, it's funny that we always talk about film and TV as like a small slice of the overall market, but it's arguably one of the most difficult of those slices to crack. Yeah. I mean, you know, like maybe you can't specifically put your finger on it but we have on candy valley you know when something feels kind of off i will say like i'm looking at some of their highlight their demo videos and you know the real world understanding and physics um i don't know if this is a shot at uh the genie 3 demo but they have a blue paint painting the wall with blue paint which was the uh the good demo in genie uh three from google right but this shot look the guy uh flipping a mirror back and forth and it's getting the reflection correct that's crazy yeah that there's definitely some physics magic going on within this engine for sure yeah so i mean i think the physics is definitely better you know i think i'm sure the quality is better you know it's just still this is i'd be curious to test it out because maybe maybe it does maybe you're getting better outputs on the first or second generation versus uh issues i had before like the success rate is probably much higher yeah like is it getting me what i'm looking for so i'm curious to try it out i got my runway subscription so yeah but again it's it's a shame that they dropped the same time as cling 01 which i think will take the bigger spotlight for the moment because video to video is just like in my opinion the more exciting of the two yeah not to sound jaded already but it's just like oh you do like text to video or image to video and it looks cool okay you sound totally jaded that bar is um yeah that that it's like a gradual input from from what we're doing uh we're so past the uh the point of like oh look what AI can do to the point like, all right, but can it really do it?

18:57You know, now we're more of a, like a realist in terms of what AI capabilities are. And also, I mean, it was like, we also know it's like, yes, AI is very good. It can make very good looking single shots, like, you know, but can we get the control and quality and stuff to make something more coherent and work in a more professional filmmaking pipeline? Yeah, for sure. I mean, you're stating the obvious from the studio world. I'm sure that's what everybody within those four walls are probably thinking about is like, how can we replace our existing pipelines with this stuff? Right. How can we modify it so that it works, that this can work into that workflow?

19:35Yeah. All right. Other update that dropped this morning. One of my favorite image models, Seadream, the Tencent. No, the ByteDance.

19:48yeah this is the same parent company of tiktok that makes cd by dance they have cdream cdream cdream is the image model cdance is the video model so cdream 4.5 uh dropped today sort of i mean you know also like it feels just it's not a banana everything else it feels like just getting up to parody with like what some of the other models could do you know more consistency see with reference images more fine-tuned editing controls um you know better text graphic layouts and design text yeah uh multi-image editing yeah nano banana pro i think came out maybe six months too early and these models are now catching up to it yeah i will say cdream is the only one still that does that thing that i like where it'll it can batch generate multiple images in the same latent space generation.

20:39There was a hack someone posted online and I was talking online with Matt Workman from Cometography Database. Yeah. Yeah. Great channel about this, where basically someone was using NadaBanana to create a grid of images, like a nine by nine grid. So kind of same hack, you're getting the image output in the same latent space. So you get like consistent characters and locations and shot variety. And then you would ask Natibana, like, hey, like up res, you know, frame three, but C dream does that in one shot. Wow. Like it just gives you the output images in the same generation. Yeah. Yeah. And I haven't seen anyone else, any other models do that yet.

21:18So I think that's still very cool. You're right. That is a big differentiator, especially in our world where we need that consistency from shot to shot. Yeah. I've used it for character, you know, generation and stuff. And instead of just making one character sheet, it's like, oh, just make me high-quality character images of different angles of the same character that I need. Damn, that reminds me of the next video we're about to drop. I should have used seed treatment for exactly what I was doing. Now you've got more video ideas for the holidays. There's no end of testing. Yeah, how much time I got, right?

21:52Yeah, stay tuned to our audience. uh we have a lot of exciting testing videos and yeah and podcasts that that are hopefully is going to drop in december and then uh yeah we'll get into the new year with a big recap this is a good demo from fall from cdream of uh i don't know if this is an input existing image and then they extracted it let me google map that address because that's here in la yeah let me see let me see but if that is true it's basically a picture of a tote bag with like a very small text of like its address in LA and phone number for Canyon Coffee. And then it was extracted and displayed on like a white background.

22:30But the text detail is preserved, which this is normally, this is an issue I've had where just small text and details get jarbled when you do AI transfers. That is a real place. Canyon Coffee's on Echo Park Avenue. So that's a real bag, probably. Yeah. So I think this is an input image of this real bag. And then they said, you know, place the bag on a white backdrop. So that's cool. Yeah, that's amazing. That's a huge industry, right? If you're talking about every single e-commerce that is selling clothing, right? All the Fashion Novas in the world and all the Abercrombies or whatnot. I mean, they could essentially replace traditional photography with this, which is scary.

23:15I mean, I think also once it gets fast enough and cheap enough where the user could do a virtual try-on. I know it's kind of tough now because the cost of the compute to generate that for every user on the website is expensive, but eventually it will probably get cheaper. Remember that Google app that we tested a few months ago, which was supposed to do? Which one? Remember the one where I wore a clown outfit? Which one? Oh, no. Did I do the toy? Yeah. I think it's – I'll Google it. I vaguely remember this, I think. Doppel. D-O-P-P-L. Google Doppel. I have not heard of Google Doppel since we did that episode.

23:51We did an episode on it. I vaguely, because I think we talked about that and we talked about they had like a kind of like a whiteboard tool area as well. This is part of like Google Labs, right? One of the lab products. But you're saying this reminds you about this virtual try on thing? Yeah, I mean, like they not only figured out the images, but also the video portion. And it was a video to video workflow. So actually, it was an image to video workflow. So that model would then walk and turn. And so you could see the backside. and everything. What else do we have? Camera controls, virtual try-on details, image quality enhancement, material preservation when editing.

24:26So yeah, it's a lot of stuff that seems on par with Nada Banana, but still pretty high quality and it's just good to have another option out there. Yeah, and the other thing is C-Dream is natively 4K output. Isn't that right? Yeah. Yeah, so you're getting... Yeah, it has really high quality. Yeah, the pixels are all there for you. Nada Banana Pro does do 4K now too as well. Right. Alright, next one. this is an update from LTX, a new feature for them called Retake. And this is interesting because this is in the realm of video modification, but you can give it an existing video and then just mark off a specific section of the video that you want to modify and redo just that section and then preserve the beginning and end.

25:05So I like this kind of specific frame modification, which the only other tool that I've seen that could do something similar-ish is Moon Valley, where they sort of have like frame specific editing. So this is a cool update. The motto, the retake demos that I've seen, the quality is a little iffy, but I like where they're going with it. Okay, so let me understand this correctly. So you give it an in frame and an out frame and then it'll generate before the in frame and after the out frame? No, you give it an existing video and then you can, I guess in their interface, mark a specific section of that.

25:41So, like, here it's marking, like, this actor performance is. Reminds me of seven or something. One video take. Yeah. And then you mark off this one section in the middle. Okay. And then give it direction and change the video. Oh, so it's, like, in-painting within the timeline. Yeah, giving it a specific range of the clip that you want to modify and then leaving everything outside of that range. That's amazing. I'm saying that's why the outputs maybe don't feel as natural or realistic, but interesting where it's going. Also, I'm putting aside all the issues of changing an actor performance with or without their permission.

26:18The quality is not there, but the concept and the execution is very interesting. yeah i like this concept of like a instead of i mean maybe you could do it with cling01 and if you give it in the prompt specific timings of like things you want to change but this like gives you that control of very narrowly like i like the rest of the shot i just in this very specific moment want to change this thing yeah and also it blends in with that overall clip as well because you're just taking a segment of it so then you don't have to do editing magic to blend all of that Yeah, yeah, that too. So yeah, interesting take on this.

26:53So I'm curious. And LTX, LTX is the is the one that's left that hasn't been acquired. Invoke and Weavie are already gone. They also have some of their own models. So like, they're not just a tool set like LTX has, I mean, this model, they have a, I think what's called LTX two, which we didn't talk about, but it's its own image generation model, it can do like 50 frames a second can do audio output. I think could do like one minute long generation so it seems like maybe they're not up to par with the quality of some of the other models but they've been doing other things that the other models aren't doing like super long generations and higher frame rates um so i think it's a cool i think it's a good strategy to like just do something different they're definitely unique in the in the ai ecosystem on what they're doing so yeah some of the use cases for this are rephrasing dialogue changing dialogue again putting aside the permission issues from performers just just can you do it focus focus on the technical ability of doing this the right stuff is you know and ethical stuff put that aside for right now for this conversation yeah i think uh consent and capability are two different buckets in ai at the moment refine tighten pacing adjust delivery you know this is interesting what did i just see i just saw some breakdown of some oh you know i think it was um what's his name todd the the the the vfx artist who worked on star wars and a bunch of films and he had he's big on twitter and he also just did like a variety video breakdown of like vfx shots and he it was dungeons and dragons or um one of the movies and it was basically like a long breakdown of how the pacing of like the shot of an intro of a character who like rides up on a horse and then the camera like 360s around them and then they like shoot a slingshot and the timing of that shot as they shot it like didn't work great and they wanted to like tighten it up and know it was like a very long extensive vfx process to like literally just like cut 10 frames out of this thing to like speed up the pacing you know potentially maybe not in the quality as it is right now but something like refine in this feature whether it's ltx or another model it's like those kind of use cases where it's just like vfx work that just needs to like tighten or clean yeah directors love to fiddle in post right like this is this is one of the things that delays movie productions is how many times you want to go back and shoot it and redo the whole thing so if you can prevent some of that and great that's water money in your pocket yeah way cheaper to mess around with stuff in post than on set all right next one yeah what you got this one's definitely been on the bottom of the pool flux oh come on flux two's great now i i mean in the world of well you know this came out this was announced on the 25th so you know we're like a week behind this and it is in the in the in the deep end of the pool so collecting dust black forest labs i find this company really fascinating they're in germany and um like i said uh some of them uh went from stable diffusion and formed this new company they are the david and the goliath right like they are still hand making models and competing with nano banana or cling or cdream all these big trillion dollar companies and they're doing it with like i don't know 50 people 100 people so anything that comes out of black forest labs i always do a double take because i love rooting for the small guys yeah i mean it's impressive what they've done with a small team yeah flux context uh if you remember when we covered nano banana one flux context was quite neck and neck with nano banana one and so um this is i believe their answer to nano banana pro which is Flux 2.

30:26I would say the only negative thing that happened here is their model is just a little too heavy. So I think it's 90 gigabytes, if I'm not mistaken, of VRAM needed. And most GPUs at home, if you're running comfy at home, you can't do that. So this is really a model that's designed to run on the cloud, on a 100 GPUs that are only available on the cloud. So having said that, the quality, the image modification capabilities, in my opinion, with the limited testing that I did, is maybe a good second place to Nano Banana Pro. Okay, that's good to know. And also you could potentially, you could work this more into custom pipelines as well, which Flux was good with.

31:07Whether you can run it locally or not, you could still run it on a cloud in Forensic Comfy and then add LoRa's and kind of customize this out a lot more than you could with a Nano Banana Pro. Yeah, the little bit of hilarity that ensued was that this dropped the same time Z image dropped and Z image is like, I don't know, I'm going to say under 20 gigabytes. Like it's a super lightweight model. And so the creator community, the comfy UI community, if you will, they embrace the image right away. And they're like, oh, Flex 2 is just too heavy, too big for us to really tweak and modify. Yeah, it's still good to know about.

31:42And yeah, I mean, a lot of the improvements are kind of similar on par with what we've seen in the other models. but multi-reference image support, image detail photorealism, better text rendering, better prompt following, world knowledge, and output or image editing on resolutions up to 4 megapixels. Also, it reminds me of Black Forest cake, which is delicious. That is good cake. Also German, I believe. Yeah, and there's a variety of models. as flex 2 pro is their state-of-the-art top model yeah flex 2 is their latest and greatest at the moment and again um so there you go um it's already probably built into firefly foul for sure i've seen it on there i'm sure comfy has it already yeah so yeah it's already proliferated into the ai ecosystem did this fold in context or is this is context going to be i think it's a it's a new product i think context is probably just a different product altogether so like eventually it would be like a flux to context yeah i think this is like a proper foundational model for them and a new architecture going forward that's my guess because yeah it doesn't seem like it's much about editing existing images but like generating new images from various inputs yeah and again um we're we're so spoiled and so jaded because in today's world in today's episode We have Z-Image, Flux2, SeedDream 4.5.

Read the full transcript

33:09And if you're going to compare the three side by side, I mean, the difference is minute. They're all really good. And we're swimming in like golden AI image models. Yeah, it's just like, which option do I pick for the most photorealistic? Exactly. And this is just after covering Nano Banana Pro like a couple weeks ago. So like there is no shortage of good technology for you to pick from. Yeah. I mean, that's a good question, too. Once you get to that point, what makes you decide? We've seen a lot of photorealistic stuff. I think some of the deciding factors, it's like, depending on what your project is and if it's not photorealistic, we've seen some models just work better with different styles or different types of images or different prompting techniques than others.

33:50And also, we don't have a cost comparison, but it would come down. Images are less concerning than video generation, but it would come down to cost per image. I know we're comparing literally pennies to pennies as far as like infras cost goes. If you compare that cost to like actual photography or actual videography, like it's like 1 % of an actual shoot, right? So we really shouldn't even be looking at pennies. I mean, whether it's on the cloud. I will say the question that you asked earlier, which model to go with for what? But so when you brought up the seed dance or rather the seed dream example with creating nine images in the same latent space, like that's a very unique thing to that model.

34:33And so when you need that, you use that for the testing that I did that we're going to show the episode for you guys. I love Nano Banana Pro strictly because of the image fidelity. Like for me, having pixel detail is was like number one. So I just went with Nano Banana Pro for that. All right. Last one. This one's a quick one, not generative, but an AI update. 12 Labs, not to be confused with 11 Labs. 12 Labs has the foundational models that analyze video. And so they're good for basically turning all of your video library into searchable indexes and doing all sorts of stuff that you want with that.

35:10They launched a new version of their model, Marengo 3.0. Basically, smaller embeddings, faster indexing, temporal and spatial reasoning. So not just looking at individual frames, but kind of understanding more of the context Is it a fast moving object or slow object? And then a better identity of like images and text inside the videos and stuff. So basically this is like for, you know, if you have your video library and you want to have AI understanding of your content to search or to, you know, find things or put things together. Or if you're building AI agents for production type work, you're going to need the eyes and the brain for it.

35:43Yeah. This is how AI can understand what videos you have that you shot so you can do things with them later. So it's cool to see a new improvement of their model. they've kind of been one of the biggest players in town to do video understanding uh on a on a level and um i i saw the 12 labs guys at a conference a few months ago um their presentations are amazing like the stuff that they show so big fan of these guys hopefully they continue to do good work yeah yeah so good update here that's pretty much their post-thanksgiving roundup yeah no that was a that was a great roundup and i'm glad we waited we got everything in one episode And guess what?

36:21The next one, we're going to have more because this industry just can't stop innovating. After we stop recording, there's going to be more news. I do want to give a special shout out. And I believe I've shouted him out before. Paul Trapani, thank you for your comments consistently on Spotify. That really helps. If you're listening on Apple Podcasts or Spotify, we would love a five-star review if you haven't done so. It really helps the algorithm more than you think. Also, Spotify Wrapped is coming out now. and if we have made here our top whatever 5 10 list of podcasts uh listen to please let us know tag us or send us over uh your spotify wrapped we'd love to uh see it and share it out we'd love to see that i'm curious at least everything talked about at denoise podcast.com thanks for watching everyone we'll catch you in the next episode

From the publisher

Addy and Joey analyze Kling O1's superior video modification capabilities versus Runway Aleph, Z-Image's emergence as an open-source alternative to proprietary image models, and Seedream 4.5's unique batch generation feature. Plus, they compare new releases from Runway Gen-4.5, LTX Retake, FLUX.2, and TwelveLabs' Morengo 3.0.

--

The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.

More from Denoised

All 101 episodes
Kling O1, Seedream 4.5, Z-image: AI's Biggest WeekDenoised · 37 min
Listen in VO