AI Roundup: Qwen Layers + Kling Animator + Wan 2.6 (Nobody's on Holiday)

22 Dec 2025 · 23 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Denoised - AI Roundup: Qwen Layers + Kling Animator + Wan 2.6

Episode Overview In this episode of Denoised, hosts Addy Ghani and Joey Daoud provide an in-depth analysis of recent advancements in AI technology relevant to the film and media industry. The hosts discuss several updates, including Wan 2.6, ChatGPT Image 1.5, and Seedance 1.5 Pro, while also touching on the implications of these advancements for filmmakers and content creators.

Key Topics Discussed

  1. Wan 2.6 Announcement
  2. Version Differences:
  3. 2.2: Open-source and downloadable for local use.
  4. 2.5: Commercial model requiring API/server access.
  5. 2.6: An upgraded commercial model with enhanced features.
  • New Features:
  • Casting characters from reference videos into new scenes.
  • Intelligent multi-shot narrative support.
  • Native audio-video sync with multi-speaker dialogue and lip sync capabilities.
  • Advanced image synthesis for cinematic photorealism.
  • Critique:
  • Discussion on the marketing language used, particularly regarding “studio quality” audio.
  • Mentioned the AI softness in visuals, suggesting a need for media professionals to create better demo content.
  1. ChatGPT Image 1.5
  2. Performance:
  3. Improved prompt adherence and editing capabilities.
  4. Competing closely with existing models like Nano Banana Pro but retaining some AI sheen in outputs.
  • Future Applications:
  • Consideration of using image generation models in tandem with detailer models for enhanced outputs.
  1. Seedance 1.5 Pro
  2. Updates:
  3. Introduction of native audio generation.
  4. Vague descriptions of film-grade cinematography and visual quality improvements.
  • Model Bias:
  • Discussion on inherent biases due to data sources, specifically relating to character and environment generation.
  1. Kling Animator Updates
  2. Enhanced Features:
  3. Actor performance tracking that allows reference videos to drive character animation.
  4. Notable improvements in facial performance capture, though concerns were raised about retaining micro-expressions.
  1. New Model: Qwen Image Layered
  2. Functionality:
  3. Segments existing images into layers to generate missing parts beneath.
  4. Potential for integration into workflows for animation and background manipulation.
  1. Luma AI Updates
  2. Ray 3 Modify:
  3. Enhancements to video modification capabilities using the latest models.
  1. General Observations
  2. AI Model Trends:
  3. Increasing shift towards API-only availability due to hardware limitations.
  4. Noted the underutilization of AI data centers and the importance of optimizing GPU use for large models.
  • Creative Application:
  • The hosts emphasize the importance of understanding one's prompting style and how it impacts the output quality.

Key Takeaways

  • Recent advancements in AI models are significantly improving video generation, character animation, and image editing capabilities.
  • Filmmakers and content creators need to stay aware of the nuances and biases in AI-generated content to effectively leverage these tools.
  • Collaboration between AI researchers and media professionals can enhance the quality and realism of outputs, leading to better integration of AI in creative workflows.

Conclusion The episode serves as an insightful roundup of the latest developments in AI technology affecting the media and entertainment industry. The hosts encourage listeners to engage with the advancements and share their experiences with different models and workflows.

---

For more details and updates, listen to the full episode on your preferred streaming platform or visit [Denoised Podcast](https://denoisepodcast.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28All right, welcome back to the noise. day break. Doesn't seem like it. Maybe this is the last week. We'll see. We'll see if things slow down for Christmas next week. All right. But one of the big ones, one 2.6 is announced. It kind of branches off into two streams because we've got 2.2, which is still open source. You get download that, run that locally. And then with 2.5, which was sort of more of the commercial VO3 kind of model, audio, they have this upgraded 2.6. It's commercial. You have to run it through their APIs or through their server. You can't download the models for this and run it locally.

0:58right some of the improvements cast characters from reference videos into new scenes nice uh intelligent multi-shot narrative so this is that similar thing with c dance or c dream same same latent space exactly and get edits so you get kind of more consistent outputs in your single generation yep native av sync generate multi-speaker dialogue with natural lip sync and studio quality audio so audio output okay that just sounds like marketing terms studio quality the audio yeah i don't really believe i don't yeah i take that out of the grain of salt but it does generate the audio with the video 15 second output to at 1080 that's huge pretty good 50 seconds that's huge uh one of the longer ones that i've seen yeah and then advanced image synthesis and editing deliver cinematic photorealism with precise control over lens and lighting support multi-image reference for commercial grade consistency and faithful aesthetic transfer i'm wondering if this is similar use case uh to something that i've been enjoying with Klingo 1 where you can just give it reference images and not necessarily like a first frame, last frame, and kind of make, be like, here are the reference images, like the character, the place, the object, make the shot with this stuff without giving it a specific first frame.

2:13Yeah, I agree with you. The AI video models are so image specific, and they're very picky on what you give it, and that has a significant sway into what it generates. Yeah. I've been doing some research on why more and more of the newer models are going into API only. Oh, okay. What do you see? We're hitting a limit fundamentally with how much VRAM a model is as far as size goes. So the way they do this, and this actually goes into why AI data centers are underutilized. So if you have a model like 1.2.6, which is, I don't know, I'm guessing here. Let's say it's 800 gigabytes, something insane, almost a terabyte.

2:56They're going to split the neural network into three layers. So it'll go from GPU one to GPU two to GPU three. Guess what? The two GPUs has to wait until GPU one is done. And then GPU three has to wait till GPU two is done. And this is why, you know, you see sometimes in the news that they're building massive AI data centers, but it's fully empty. It's underutilized. it's only got 33 utilization well yeah because these models are loaded across three gpus or two gpus and it has to wait has to wait so chat gpt certainly is in the like the terabyte i think they just launched 5.2 that's got to be huge yeah and so like there's no way that can run non-api at this point yeah they're just getting too big yeah too powerful yeah makes sense i mean unless that they also unless they do like a distilled version or something like i mean vo 3.1 fast is not a great example because that's still a commercial model but it could be something like that that's true yeah yeah there's also there's always room for new architecture you don't need things to be massive maybe it just has to be built differently yeah i will say in the demo video like i wouldn't the stuff doesn't look like it doesn't it doesn't look the quality of like vo3 it doesn't it's soft yeah it has this ai softness it doesn't look it doesn't look super sharp also you gotta remember a lot of times these are built by researchers and ai people if you want kick-ass demo videos you got to give it to media people yeah to like push it and make it something see how realistic they can get out of it yep all right so yeah one 2.6 i mean i don't think it's rolled out to free pick yet or foul but i imagine it would be soon where you can call it up and generate oh never mind i have the foul like next foul it is it is live on foul so you can spin it up yeah that's cool with the 15 second i think that's a big that's a huge leap probably what i'm most excited about yeah and uh i mean just us covering one 2.1 1 2.2 animate 1 2.5 like we're big fans of one on this show because a lot of it is open source and downloadable and comfy ui-able if that makes sense yeah yeah uh and in files no it does confirm reference to video use one to three reference videos maybe the met images run through reference images for character object consistency their demo looks sharper but this is more of like a cg demo yeah like look two years ago we saw that we would have mind's blown right now we're like oh let's do a little bit better alibaba i would be curious if this is one single generation with the multi-shot uh feature right uh well hey maybe i'll have time for another uh tutorial video yeah this break addy's been dropping the tutorial they got some good response i think they're they're really useful by the way i should be the i should be the monster right now i'm gonna use you for the next video if i may yeah if i can have the permission so um i remember when we said hey we want to talk about the ai bubble and open air and all that comment if you want us to do that somebody was like oh hell yeah yeah there were a few yeah yeah so we should do one of those yeah we just need to actually do a lot of research for that now we'll just be super opinionated and wrong it's a bubble okay next up chat cbt image 1.5 is now out we talked about this in the last episode of seeing some demos on lm arena with the what was it called chestnut or whatever some output demos from their site it's not bad it's pretty damn good no it's not bad i would say not well nothing here is like in the that one looks pretty it still has i feel like it still has the bit of the ai sheen yeah for the more realist photo realistic that that's the part like nano banana pro and vo31 is so good at getting rid of is that ai sheen looks like a photo yeah yeah i have heard that it's kind of prompt adherence or just giving it directions like remove this, move that is really good.

6:50And following instructions is just changing stuff in like first go around. Sure. So I've heard really good things about that. Everything else. I mean, I haven't tested it too much yet. And then looking at these demos, you know, I think it's the images. I can't pick one apart. That's like, oh my God, look at that. I mean, they look like 2025 level image generation quality, if that makes any sense. I feel like it has caught up to VO3. I mean, it's caught up to Nano Banana Pro. And maybe not in the output, but in the being able to follow instructions. Edit images. Edit images. Give it multiple inputs and create an image.

7:24I think, look, this is maybe not the first generation and be all. Perhaps you take this generation into an enhancer model or a detailer model. And maybe that is the way. So when we covered Z Image, I don't know, three episodes ago, when it first came out, that's one of the big use cases for Zimage was a detailer pass. And there seems to be a lot of workflows online just doing that. Yeah. Yeah, I'm curious if anyone's messed around with it, found it better, you know, at doing a task that Nada Banana isn't, like, let us know in the comments, because I'm, yeah, I'm curious. I also think with all these models, too, you know, it kind of depends on, like, what is your prompting style or what are you trying to do?

8:04Because I feel like there isn't a good blanket answer for, like, yeah, it's the best one. It's like, oh, maybe this is the best one for just how you like to prompt or work or generate. Right. Slash what type of stuff you're working on. So maybe this is great for animation. Yeah. Are you a prompter that doesn't include a lot of details? No, it's the shortest prompt possible to get what I want. Okay. I'm not writing an essay. Yeah. Well, me, I am very specific about camera language and lighting language, but the actual content itself, I let the AI decide. I've been trying to get a video to replicate a telephoto, like zoom lens, and it does not.

8:35It has been very hard. dude try doing anamorphic bokeh like what sorry what yeah to give it that really compressed look it no no no doesn't know it yeah oh speaking of which blazar just dropped some new lenses oh yeah so this one uh sorry folks full segue um so you know they're known for sort of budget anamorphics that fit onto dslrs and smaller cameras they dropped a new one where it has the look and feel of a two times stretch factor into a 1.3 stretch factor. So you're like really getting concentrated anamorphic look and feel without the crazy cropping of the super wide aspect ratio. Oh, okay.

9:21So it's still 16.9, but it feels like, you know, 21.9 or something. Oh, that's cool. So you get the anamorphic or anamorphic feel without having to like sacrifice your frame yeah yeah that's cool all right next update uh c dance one of our other favorite models has 1.5 pro is out ish they have announced it but i still haven't seen it kind of roll out into any into free pick or in a file or into any api so i have not been able to try it but updates native audio generation this seems to be kind of standard now like if you're doing the ai video model and generating the audio that goes with the video is kind of the baseline so they've added that film grade cinematography and visual quality yeah whatever that means yeah okay i have one nitpick thing about c dance yeah so because it's by by dance and by dance has obviously our tiktok and then the chinese tiktok yeah so a lot of the data comes from the chinese tiktok so it's like naturally inclined to generate asian people and chinese environments yeah you You almost have to forcefully prompt your way out of that bias to get, quote unquote, American results.

10:33I can see that. I mean, I'm thinking for C Dance, I've probably only given it images as input. So I haven't done text a video with C Dance. But I could see that happening if you just had a text image without specifying and just be like a woman walks across the frame that it would most likely be an Asian woman. yeah no bueno if you're like trying to make a show that's diverse americans just specify what you want in every prompt uh come on too much work i mean if you're in asia all of your prompts are going to be like defaulting american people and you'd be like yeah i don't want these american people all the time or it could be great for making a like a i don't know a short story that's based in asia right like a post-war vietnam or something wait say again so let's let's say your creative calls for you know you're building a post-war vietnam short film you see dance you're naturally going to get that bias oh like you use a model based on kind of where your story's from yeah it could work yeah i mean i don't know well they're down their homepages are all old white guys well yeah that's because we're on a u.s domain right yeah like it's tracking our location but the weird thing about cdance is i've uh or by dance is like i've tried to just connect to their api directly to like just use the service that way through versus through like fall or something and like it doesn't really let you it's weird it doesn't really like let you self-serve yeah so the audio thing and then i mean i'd have to test this out more i can't really say like oh they're very clear on what is better in this you know film great cinematography and visual quality is pretty vague.

12:08I've usually had pretty good outputs from it. Yeah. So just general improvement is probably what I'm sure they're going for, but nothing that's like super specific. Sure. I mean, I'm about to try it. Once I can get access to this, I'll try this model with a telephoto lens prompts and see if it can do any better. Yeah, maybe we'll do an episode of just trying all these models out with like one single creative, just side-by-side comparisons. Yeah. I will say after doing the last tutorial, which I ended up using Kling 2.6 over VO 3.1. Yeah. 2.6 is great dude both both oh on at 2.6 they're like i'm using those more and more now yeah yeah cling cling uh cling came out swinging oh and uh guess who makes them this company called kaisho yeah i could yeah i didn't even realize i was like wait i thought it was a big tech company but i guess not yeah unless they're an arm of another tech company right being a cling um there is a new upgrade feature to 2.6 actor performance tracking so you can take a reference video and use that to drive a character performance and apparently you're doing better than 2.2 animate or runway act 2 pretty good though demos that i've seen yes it does look much better than 2.2 animate from angry tom so we have uh yeah uh human performance to animation okay i'm looking at this one as well oh yeah that was the solid yeah solid look at the hand registration to the face and uh wrinkling on the on the shirt hand in front of face still goes but then again we're only going to see the very best thing right yeah i'm curious how it does on the facial performance because that's also where i see it kind of like neutralize a lot of nuances and face performances yeah a cling 01 not good at face oh okay yeah but uh this this one seems to be you know it all comes down to using a pose estimator or open pose or something underneath a hood which yeah I mean, this one is a figure skater spinning super fast.

14:00Did really well. That's solid too. But it's also a figure skater on a figure skater with just an outfit change. That's an easy. Right. We're not one to one. Making it a dinosaur or something. Exactly. This one is a little bit more difficult because it's a woman driving a man and a whole different ethnicity, clothing, everything. So, yeah. And this one was a demo of one single 30 second one shot. So you can do, you can do from three to 30 seconds. so that's also good 30 seconds wow duration yeah so one 2.2 i think is capped at either three or ten seconds and your computer specs yeah and i you know i wonder if you could do full hd in this one because also one 2.2 is kind of capped at like 720 720p yep this is good yeah this is a good step forward for us uh on top performance 2.6 has already been a really cool model just for like video generation and now like this bonus feature is added so that's cool yeah i I think our viewers know that I have a natural bias into using human performance over fully synthetic performances.

15:03Yeah, and I think that's... So, like, you need tooling for that. Yeah, and tooling to capture and, like, not lose it. Because I've seen still, you know, you have a human performance and maybe it gets big gestures, but, like, you're still losing the nuance in the facial performance. Yeah, like all the wrinkles around your eyes and things like that. Like, micro emotions. Exactly. All right, next one, another one from Quen, which was Alibaba's sort of open source video generation model. They have a new model called Quen Image Layered. And this is pretty cool because basically you give it, I think you can give it, you give it an existing image, and it will segment that image into layers and generate all the missing parts beneath it.

15:44And then you've got a multi-layer thing, an image that you can take into Photoshop or modify or kind of do whatever you want. okay that's it this reminds me of kubrick oh yeah kind of yeah except kubrick wouldn't paint in the other other layers areas where it just sort of segment that's amazing though like the fact that it's so easy to do now yeah and i feel like you could build this into a workflow where you could generate you know your first image or something and then what is it called quen what uh quen image layered you can make your first image and then you know if you want to animator kind of bring that into something like after effects where you have like fine-tuned control right then push it into this yeah wow that's really useful especially for all the basic little web animations that we're seeing everywhere i mean that also i mean look you could use it as a cubit kind of workflow thing if you have an image i mean it's open source like take it and build a product on top of it yeah they're kind of demoing it for like social media ads and stuff like that but you can take any image and have it automatically segment and now you've got your layers that you could run parallax in on a background environment or whatever you want.

16:50Sure. And maybe have more control over than World Labs, which I've been messing with. Oh, like building the entire world? Yeah, I mean, World Labs kind of does that because it's sort of like it generates this mesh 3D world. It's also low poly. Yeah. And it gives you, you know, it's not like a full 3D space. It's like it gives you faux 3D from where you're standing. You can move around a bit, but then as you keep moving around, it falls apart. Yeah, it breaks a little bit. Yeah. I mean, look, a lot of our listeners are from the virtual production world. Like this is one comfy workflow away from plugging into your wall.

17:22Yeah. I think, yeah. You know, take like a photoreal generation from Nana Banana or ChatGPT. Especially with all these newer ones that can do 4K outputs. Right. Put it through that. Segment. Segment. Move it around. Animated. Add your fluttering cloud, fluttering flag or whatever. Like do a image to video for a couple of the layers. Yeah. Keep the other ones still. Dude. Yeah. bada bing bada boom yeah so this one feels very practical yeah i'm excited about messing around with this one let's see some other smaller updates luma has a new update so now they have ray 3 modify so ray 3 was their model came out like a few months ago uh that could do 10 bit 10 bit i believe even more yeah it's it's quote unquote hdr it's hd dynamic range 4k video outputs right but now and then they separately had their modify video feature which is like kind of runway olive-ish video to video you can modify and transform things but that was on the ray 2 model the older model so now the newer model has modified video so you can give it an existing video give it a prompt or of what you want to change and it'll use its newer model to uh modify that dude luma and runway nick and nick yeah a lot of a lot of powerful very emily focused and yeah both of the offerings are yeah i think the only thing luma doesn't have correct me if i'm wrong is like a performance thing like a one 2.2 anime type thing no not that i know of i mean i think this might be the closest to it of like give you recording a performance like they kind of have it in this demo video and then changing the character but it was just them looking like i don't really know how well it would do if they start talking or right i don't think it's like geared for performance no i mean it's for character swapping but well that changed uh that that changed v sure yeah is it on free pick maybe we try it I'm sure if they have it in the API it will be but it just came out today on Luma you know one of the other things that I noticed interestingly enough running 1.2.2 animate across a bunch of different services so I ran it on free pick I ran it on my own desktop on comfy I ran it on foul and replicate So four places.

19:35Yeah. The best quality, bar none, was FreePick. Interesting. So I think it goes down to like how many steps are built into it and how much GPU it's going to use. Yeah, I'm curious with those models, like an open source model like that, because basically they got to set it up on their own cloud. Yeah. And how are they? They don't really tell you what settings they're using. And are they using an upscaler as it comes out? Like they're adding a bunch of stuff in. And the worst by far was FAL because it's like the budget one. well did you dig into because with file usually it gives you more options sure to change the settings i just left that default you left the default did you i think you can add more steps yeah there is a step slider i think it was like eight steps or whatever did you crank it up to see trying to save money man i'm curious how much the price changes based on how much like how many steps probably a lot you you you add i think that's an exponential scale like eight to 16 would be like, I don't know, four times.

20:27All right. Well, let's dig into that. I'm curious. Yeah. And this one, I don't know too much about this, but it's popped up on my radar from Deckart, Lucy Motion, image to video model with precise control over motion. And the demo is basically just dragging arrows on video of like where you want things to move and it moves. Yeah. By itself, not impressive. If it's a part of like a bigger solution to like use this, if I'm working in like, Like, let's say Kling. Would I come out of Kling, use this, and go back to Kling? I mean, I think this would be the video you get. I don't know. I mean, maybe it could work if you like are trying to do something very specific and you're like, it is not working.

21:07Yeah. The problem with, like, putting your video through a bunch of different models is it just degrades. Yeah. You just keep warping. Even if you're like, don't change anything, it still is a game of telephone. So, like, yeah, this needs to get sucked into, like, a bigger model or a bigger platform. But, yeah, I mean, it's promising in the way that we have control not too dissimilar from Moon Valley. Oh, yeah, they do have that feature. Yeah, that's right. Yeah. They have a bunch of features in their web thing that I forget about because you have to use it through the web. You can't call it up in the API.

21:33Yeah. But yeah, I mean, we're seeing more stuff like this of just like less prompting, more, hey, let me click and drag what I want, or other types of ways to control this. So if you do a dotted arrow, does it move slower? Or like if you could kind of do like a bezier kind of curve, like with speed ramps, that would be. Yeah, I think that's the next step. the AI researchers are like wait we just did this you're already giving us notes go back to your lab we want the key for the same keyframe motion that we had from before work through Christmas get this out January 1st January ship it yeah oh yeah that's kind of the roundup that I saw anything else you pop up on your radar I mean there's so much to sort of cover I think that's good for now but we should definitely do an end of year wrap yeah we'll do that uh we'll do that next week okay yeah so let us know what you thought thanks for everything we talked about as usual at denoisepodcast.com we are so thankful for all of you listening to us over the year uh we're gonna have a big episode coming up so if you have any ideas for what we should cover send some comments thanks everyone catch you next week

From the publisher

New AI model releases won't stop for the holidays! This week, Addy and Joey dive into Wan 2.6, ChatGPT Image 1.5, Seedance 1.5 Pro, and other video generation updates.

--

The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.

More from Denoised

All 101 episodes
AI Roundup: Qwen Layers + Kling Animator + Wan 2.6 (Nobody's on Holiday)Denoised · 23 min
Listen in VO