In short
Denoised Podcast Episode Summary
Episode Title
Wan 2.5 Brings Talking AI Video
Hosts
- Addy Ghani - Media Industry Analyst
- Joey Daoud - Media Producer and Founder of VP Land
Episode Overview
In this episode, Addy and Joey dive into the latest advancements in AI technology impacting the film and media industry, specifically focusing on Alibaba's newly unveiled AI model, Wan 2.5. They also explore various AI tools and models currently transforming media production.
---
Key Topics Discussed
- Alibaba's Wan 2.5 Model
- Overview: Wan 2.5 is an AI model that generates video with integrated audio and speech in a single output, positioning itself as a competitor to Veo 3.
- Accessibility: Currently not open-source; accessible only via API or specific integrations (e.g., Comfy).
- Multimodal Capabilities:
- Can process audio, video, and text simultaneously.
- Potential for creating lifelike digital avatars (e.g., a digital Robert De Niro).
- User Experience:
- Allows for high-resolution outputs with options for duration and audio generation.
- Initial tests show promise, though some quality issues remain.
- AI in Media Production Updates
- Qwen Image Edit 2509:
- New version of the open-source image generator allows for advanced text-based editing.
- Google's Mixboard:
- A tool that organizes generated outputs, facilitating ideation for creative projects.
- Updates to Topaz Models:
- Introduction of new video upscaling models (Starlight Sharp and NYXXL) focusing on improving video quality and detail preservation.
- Figma Integration with Claude
- MCP (Model Context Protocol):
- Integrates AI functionalities with Figma, allowing seamless transition from design to app building.
- Benefits: Bridges the gap between UI design and coding, making it easier for designers to implement their creations.
- Ethical Considerations in AI Music
- Suno V5: The latest version of the AI music platform, enhancing vocal generation quality.
- Concerns: Discussions around the ethical implications of AI-generated music and its impact on human musicians.
---
Key Takeaways
- Advancements in AI Models: Continuous evolution in AI technologies, with new models providing integrated solutions for video and audio generation.
- User Control and Customization: Emerging tools emphasize user-friendly interfaces and customization options, enhancing creative workflows.
- Industry Trends: A growing trend towards cloud-based solutions for AI applications, making powerful tools more accessible without the need for expensive hardware.
- Ethical Issues: Acknowledgment of the challenges posed by AI-generated content across various media forms, including potential impacts on traditional artists.
---
Conclusion
The episode emphasizes the rapid advancements in AI technology that are reshaping the media and entertainment landscape. As new tools emerge, professionals in the industry must navigate the balance between harnessing innovation and addressing ethical considerations, ensuring a sustainable future for creativity in a technology-driven world.
For further details, visit [denoisedpodcast.com](http://denoisedpodcast.com) or check the show notes on YouTube.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00All right, welcome back to the noise. it is time for our weekly roundup of all the ai news stories and filmmaking this is the ai roundup let's get into it
0:17all right we're back good to see you again remotely this time addy yeah good to see you too uh we're not in person again but uh we're still in the same yeah you know traffic it's later in the day so you might as well be in two different states all right uh so big week in ai for the open source model world sort of it was like there's a sort of a caveat to that uh but probably the biggest story this week is wan 2.5 i texted you i'm like hey did you hear about 2.5 and then you're like cling i was like no wan another model they dropped another one but this one's well so that's why i said this is sort of where my quotes and open source were this one is not public or open weights yet.
1:00I'm assuming they will eventually release it. So you have to use it through their API or through their website or Comfy has integration with it. But it's basically the next level of one. I don't know why they skipped 0.3 versions, but it is the next level that is sort of a VO3 competitor. It can generate video with audio, with text, with speech, all in one platform or in one output. Yeah, yeah. And it's really powerful. Well, Well, first of all, because it's completely open source and free. Like anybody can just grab it and use it in their company. At this point, you can't right now. You have to use it through the API.
1:34And hey, yeah, that's what I'm saying. I mean, I'm assuming eventually there will be an open weight model you could run locally. But right now, because they're calling it 2.5 preview, and you have to use it off of through their API or through the WAN's website or through any of the partner integrations. Oh, interesting. I wonder why they're doing that. Maybe because they want to just see how it's being used and feed that into further improving the product. But the test that I've been seeing online is a lot of character-to-character video transformations. So think of it as puppeteering for animation rigs and things like that.
2:12I think you're mixing it up with another model again. There is one 2.2 Animate, which is a...
2:26got him again are you gonna fire me no but i think now we need to like make a like make something where it's like got out of you again nailed him oh okay god i'm so embarrassed and i'm on camera so well i mean it also goes to show that there's so many new updates i think i think we covered okay we got one 2.2 animate which is open source and download it is basically like an open source runway olive or a runway um act two where you can give it a driving video for performance and a still image of a character and you can animate drive the performance of that character that's free open source you could load it in comfy they got workflows and stuff that's like a week old so it's a million years old already in ai land this week was WAN 2.5, the newest model, preview model.
3:22You can't download it quite yet. You got to use it on the API. This thing is either text or image input. But the big lift of this is it generates audio. It could generate speech inside the video. It's one unified model. The closest thing we've seen to VO3 out of any other model. Absolutely. Absolutely. And yeah, they're making a big fuss about it being multimodal, which means you can train it on audio, video, and text at the same time. I'm guessing it's going to be very fine-tunable just the way 1.2.2 is. And what that means is if you have the likeness of a person, let's say, I'm just pulling an example here.
4:00You take Robert De Niro's physical likeness, his vocal likeness, and then you could train the model to have a digital Robert De Niro that is outputting audio and video. Interesting, yeah, having that fine-tuned control. Yeah, so that's just like a digital human example. Other examples could be taking entire styles of story. So let's say you are building a world, and then that world has its own noises and sounds in it. We talked about Star Wars having the lightsaber sound and all the signature sounds of Star Wars. So let's say you designed a world like that, where it's not just a visual design, but a sonic design as well.
4:41You could train the model to have those two things output together and work together. Yeah, that's a good point. Yeah, I ran a test, one quick shot inside Comfy using the model. And the nice thing about it is you do have a lot of resolution and framing options. So your usual one-to-one, but you also do full horizontal or for full vertical, starting from 480p all the way up to full 1080p outputs. Duration, you can either do five seconds or 10 seconds. And then you have the option if you want it to generate audio or not. And then by default, it came on with a watermark turned on, which is a bit weird.
5:16It had this little Chinese watermark in the bottom corner. I was like, why is it there? Oh, it's still in preview mode. And then I found that you could just turn it off. I will say, I mean, look, it was maybe a dollar, I think, 1080 for a 10 second generation. So a lot less than VO3. And once, assuming they will eventually release this as a open source, then you just run it on your hardware, which depending on your hardware could be take a long time or not. But I will say the quality, this was like a quick test for some action shot prompt. Oh, yeah. That looks great. Yeah, it still has like that AI-ness feel.
5:52Yeah, it definitely feels synthetic. I ran the same thing in VO3. I'm sure you prompted for that camera rotation thing, right? I did. I asked it for a Michael Bay 180 orbit kind of thing, so that one is a bit more of a push-in. Oh, it didn't get Michael Bay at all then, because Michael Bay does it around the character. It didn't get that. VO3 didn't get that, but VO3 definitely has more detail. Oh, well, Moon Valley would probably nail that camera move. Moon Valley, you could tell with the camera move. With the spline, I'm guessing. Yes, blind camera mover thing in it. Yeah, that's my one quick test.
6:25So this is definitely obviously not definitive, but it's good to have another VO3 competitor. Maybe one, yeah. Who makes one? Is it Alibaba? Yeah, Alibaba. So maybe Alibaba's strategy here is to get market adoption with early versions of one, 2.1 and 2.2, and then start charging people when they get to 2.5. Yeah, I don't possibly, because yeah, we have wondered what is the business model for all the other models being free, essentially. I mean, it's probably to a point, too, where eventually you want to process powerful enough stuff than you want to use their service to not have to wait an hour or two hours for one shot to generate.
7:05I'd be curious. I mean, but their whole playbook has been open models, so I would be surprised if this one, they're just like, no, we're just going to, you have to use it through an API and pay for it. I think the benefit to one over VO3, if I were to pick one versus the other, is the only way you can customize something in VO3 is if you give it reference images and you sort of just trust it to build those references into the generation. With one, at least with one 2.2, you have really high level of control in that you can build a LoRa on the side and then attach it to the generation. So I'm guessing if we have the same level of LoRa and fine tuning available for 2.5, that in itself could be a major game changer versus comparing apples to apples.
7:50Yeah, to be able to put these into pipelines and do a lot more customization, that's a huge lift. And then you can start to combine multiple Lauras at the same time. So if, for example, you have your, I hate this example, but you have your Star Wars world, like you have your galactic world, Laura, which is building the backgrounds and sort of just putting you into that environment. And then you can combine it with the actor's likeness, Laura, so then you have to control over who the people are and obviously with audio attached to them they would even sound like the actor yeah i'd be curious too if there's more control over the audio too if this releases as an open model because right now it's sort of like the same bucket as as vo3 where you can just describe the audio that you want and the prompt but if there would be a way where it's like you can give it a voice sample or you can give it audio inputs or give it something else to have more control over what the audio part of the the multimodal uh video generation would be yeah and also can you generate like completely new um audio like not just dialogue but foley or soundtrack and things like that yeah it did the similar i mean it did a similar thing in that test clip my audio is not hooked up right now but it added like some explosion sounds and some music sounds so like it did you know like similar with vo3 where it's like oh yeah that sounds like it kind of fits but it also sounds like it was ripped from a movie sort of garbage i had uh i did another test where i had like a talking dog podcast and i said make the dog say something and it did do it but the like lip sync wasn't as good as vo3 in this one test i want to caveat these were just like one test output so it was not very it was not a very thorough test the vo3 audio at least the examples that i'm thinking of uh i mean that audio is not usable at all for anything professional.
9:41No, unless you want it to be weird and funny. It makes for a great LinkedIn post. I mean, the audio sounds weird, but the lip sync is usually pretty close to what it would look like. I just noticed, too, here in the WAN update with the new model, that they specifically called out image capabilities with advanced image generation and image editing control. And this is interesting because this was sort of like WAN 2.2 came out and it was like intended to be a video generation model. But then a lot of people realize like if you just do one frame, it's actually also a really good image generation model.
10:19And so it seems like they're also leaning into that, too, where it's like, oh, you use the same model. You could just do the same, use the same model and same control for just doing image outputs and editing. Yeah, I've used WAN for image generation in the past. It's equally as good as any image generator, if not better. I think the one thing video generation models have over image generation models is that when you go into the latent space, you have not just a notion of the things that need to be generated, right, like the tree or the sky or the person, but you also have a notion of time. And that really helps with image generation, because if you're generating somebody, you know, that's in motion or they're doing something there, the AI model is aware of how they're moving in time.
11:08And it's almost like a spatial understanding and animated understanding of that person, which really helps with achieving like a slightly higher level of quality. Yeah. All right. Yeah. So, I mean, 1, 2.5, excited to see where that goes. All right. Next up, also from Alibaba, Quen Open Source Image Generator got an update. So Quen Image Edit has existed, but now there is a new version with a terrible name called Quen Image Edit 2509, assuming that is for September of the September update of 2025. But they basically said they rewrote the entire model. And it is with text-based editing on the level of like a Nano Banana or a Seed Dance 4.
11:50But it's open source free. There are comfy workflows you could download and run it, or you could use it on open platforms that have integrated it. Yeah, I'm on the Quinn blog right now. And the main AI researcher, he's all over it. He puts his likeness on a bunch of the tests. On the blog post? Yeah, if you go to the Quinn blog post. So the first one's really funny. They just take two arbitrary people and put them into a wedding photo. It's kind of a bleak showcase of the future. It's not just synthetic people, but synthetic marriages. You see the guy who has all of these fake relationship photos of him and Taylor Swift that he posts on Facebook.
12:34Yeah, this is right up that alley if you're into that kind of thing. Look, Quen actually is quite well known in the open source community. It's been around for a while, and along with Stable Diffusion and Flux and some of the other open source image generation models, it's like a building block for a lot of custom workflows. And this is a cool, too, and this also kind of ties into why it's powerful to have Comfy in the workflow, because you could use the Pose node to draw out exactly the exact pose you want and build that into the workflow. And so then it poses your image generation with exactly what you want.
13:14So it really gets into that fine-tuned control of being able to maneuver and tweak things to get the outputs that you want. Yeah. And I wonder if there's any overlap at Alibaba between the QN technology and the WAN technology because those are at the same time. I mean, yeah, they're two powerful products with different names. I've been curious, are these completely separate teams, completely separate models? or is there a lot of stuff pulled from the same training or models? Yeah. I mean, to train models is so expensive, as you know, and I wouldn't be surprised if the one foundational model is just under the hood of the Quinn model and the Quinn is just running like an additional couple of layers of network on top of that to achieve all of the customization and adjustability.
14:01Yeah. And then also showcasing high text adherence, both in english and are we ever gonna have a week where we don't talk about a new model because it hasn't happened yet i don't know what we're gonna talk about then this photo restoration one too that's pretty wild i know that's crazy that's uh actually it's funny uh some of the early examples of nano banana was photo restoration yeah yeah i mean it's great it's a great use case yeah i mean have been rivals for sorry another week more models no models no problems there are more problems like keep you track of them no what i was saying was um you know google and alibaba have been tech rivals for decade right or even more and uh you know i think when vo came out nano banana came out they couldn't wait to drop their version of those two things it's just like an arm for various between those two companies but i am still just so curious of like where the drop in the free models where that comes into play.
15:03I mean, I guess maybe the play is because, you know, they're coming from Alibaba and they're basically kind of like the AWS of China, right? I'm not off base there in saying that. That like, oh, you got these models are cool. You mess around with them. You build them into your apps. And then when you want to deploy and scale it up, you can't run that on your local computer. You got to power up some servers. And if, and Alibaba would be an option so that you might use their servers to run these models. that's my yeah exactly my guess at the business model all right next one another one from google labs this one is a mix board yeah interesting product so this feels like somewhere between miro and nano banana if miro and nano banana had a baby this would be i mean it is literally it is literally nano banana and it's a bit bit miro yeah i mean it's a very lightweight yeah I kind of like, well, yeah, I guess it's too low lift of a comparison to call it, uh, to compare it to like Flora or the other AI, not mood boards, but the AI.
16:02Yeah. Like node based, this is definitely not a node based workflow. What I would call this is maybe just a visualizer for multiple outputs, uh, an organizer of ideas. Yeah. It's very nice, simple interface. You've got your board features. You could drop images in, uh, down here on the bottom. You could just start generating stuff and it'll just throw it in there. You don't have any model selections. It's just everything's running through Nano Banana. You know what would be nice is if it did have model selections. Or if FreePick or somebody like that released this product within that platform.
16:32Well, we'll have news about that in a second. Yeah, it's simple. And then it's like if you just click on an image, it adds it as a source image. If you highlight your images, it adds it as multiple source images. So you can just tell it what you want. Again, if you're ideating and you're going through a lot of generations, let's say you're you know you're ideating on a movie and you're going through a hundred different generations you have to organize them at a shot level at a sequence level i mean this is one way to do it you know it's yeah maybe you have a batch of your characters and a batch of some locations you click on the character click on the location and then say you know make a medium shot of from you know a high angle of this yeah it's just it's like a digital white yeah and it's just convenient because it's like yes you it isn't really anything new that you can't do in any of other photo editing tools but uh it's nice where it's like okay because you don't have to download the files and re-upload them it's like here you want to do something different you just highlight them you tell it what you want it makes a new image and then you uh you just repeat i'm not quite clear how as far as like plans or if you have to be on a paid plan to use this or if there's any usage limits um it's like you know in their labs feature so that's usually not i don't know as restricted.
17:45But I'm sure there's some usage limits. I'm on an enterprise Google Workspace account, and I was able to use it just fine. Like, it's part of Gemini and Nanobinette. Yeah, I mean, I have a Workspace account, too. So it's like, I know I'm logged in, and so, like, I have Google AI plan. But I don't know, like, if I was just on my regular Gmail account with, like, no Google AI plan, I don't know if there's a different usage limit, because there's no indicator here of like credits or usage or anything like that. So I don't know. I mean, I assume - I'm guessing not. Yeah, I mean, look, if you want to free now to banana hack, this might be the way to use it.
18:21But I'm sure if you churn out a bunch of stuff, we'll eventually yell at you. But as of right now, can't quite tell if there actually is a usage limit. Yeah, the other thing, they've got some other handy tools here too, I will say to either make more images like this one or remove background, which is also a handy feature to have. Absolutely. So yeah, new cool tool from Google. I'm surprised at how good AI is at removing backgrounds just in general across all the models. Yeah, it's like funny how like that became like rotoscoping or just delete removing background stuff. It's like pretty, pretty, pretty soft.
18:52Yeah. All right. Next one is Google Flow. Yeah, it's a grab bag of updates in Flow. So Flow is Google's web product for their video creation tool. So it's where if you want to use VO3, that's sort of the home base. If you're using Google's built in product. They added a couple new features. One is prompt expanders. And so this is actually kind of cool. Basically, you can assign sort of presets of like your scene or your look inside flow. So when you do your scene generations, it can have some persistent memory to keep some elements the same, like keep the style the same or the location the same.
19:30So you can kind of build out has like more prompting options to keep some elements consistent in your prompting while you try to change or modify other elements. so i think that's like a nice useful like kind of filmmaking ai filmmaking it's like a natural evolution of of a platform or i guess enough user feedback was gathered where they're like yeah we obviously need a notion of what we're working on memory of what we're working on yeah i mean this is something you could have built yourself when you're building on a prompt thing like in a gemini gem or in your clod project where you can kind of build settings and have it incorporated it to the prompt every time you generate a prompt.
20:07But this just, it's like if it can save you a step of having to like round trip between one tool to make your prompts and one tool to make your videos, like awesome, great. Yeah. I mean, to be honest, there's just, the Google tool set is still a little bit unorganized in my opinion. Like they're, they're still just a little bit all over the place. I think, I mean, yeah, I think a lot of their stuff is like, they kind of just throw a bunch of products out there and see what sticks and then potentially double down on it. I mean, that's also why I'm always like a little bit hesitant of fully adapting a Google tool into my workflow because I think we talked about this a while ago, but like I have a long history of using Google tools over the last 15, 20 years that they end up killing, whether it's like Google Notebook, Google Reader.
20:53I was not a big user of Google Wave, but I don't know if anyone remembers Google Wave when they tried to build an alternative email. Yeah, you talked about this before. I'm guessing you must have lost something significant in one of those product discounts. I mean, they're really good products that I used a lot. And then there's just, you know, they just left a void in my life of not having Google Notebook around. Yes. Which you're still trying to fill to this day. You should just go to Alibaba. Yeah. So, I mean, I don't think Flow's going anywhere. But, you know, like if you're like, oh, Google Mixboard, I'm going to like, that's going to be my new Mixboard.
21:28it's in the labs feature and it's like i wouldn't be surprised if one day they're just like well you know we'll just toss it out bye bye yeah yeah other update in flow too and this is also ties into just like making the platform easier to use now when you have your images your frames inside flow you can instantly edit those images with nano banana so nano banana is like integrated inside flow so you could modify your images again saving you like that round trip step of like pulling your images into different into different platforms which is always good to see and then also speaking in now to banana now in photoshop beta now to banana is integrated inside photoshop so you can that's amazing prompt you change it as it comes into the layer you can also use it with uh the harmonize feature which blends elements together which we talked about a few weeks ago so yeah that's i think there was that guy that built that plugin that we talked about like a few episodes ago and then we're like well it's only a matter of time before adobe just does it natively yeah yeah and it totally makes sense it's it's solving that last mile problem right it's what you just talked about it's like eliminating all the round tripping uh when you're you know doing a giant project imagine like the number of times if it's like uh three operations per image times 100 images that's a lot of time you're just wasting yeah and like to be in your main tool set that you're using and just have everything there and so you just call up what you need but stay in your tool set and have that control and send layers like just the layers you need and save all those export steps there's yeah there's an awesome setup for doing that and adobe opening up adobe opening up their ecosystem to third-party models is probably the smartest thing to have done in a long time because stuff is just so powerful now way more powerful than firefly could have ever been yeah i will say um because i've been messing around with firefly boards which is actually like the best firefly product that adobe has made um it's just it's a good board generator kind of It's sort of like the mixed board thing we just talked about.
23:22I think we were both out of town when Firefly boards came out. But it's basically the same idea. It's a mood board, but you can call up a bunch of image models. You can call up a bunch of video models. But it also tracks the prompts that you used, and it tracks the model that you used in that image. So if you move that image around to other apps, it'll track where it came from, which is just nice for kind of AI authenticity and tracking where data came from. But it has Firefly 4, the newer image model, built in. And I messed with it. And it was actually, like, decent. It was, yeah, no, Firefly was sort of the punching bag of AI models in their early version.
23:56But I was pleasantly surprised. You're also a really nice guy. Yeah, well, I try to be optimistic. You're very easy on things. Well, no, I mean, we've talked about how the first Fireflies were pretty useless. But the newer ones were pretty good. Good for Firefly. I mean, it's not a... Good for Adobe, good for Firefly. But the Firefly boards is actually a pretty handy product that they've built. Well, it's funny you bring that up. That's like the number one dilemma with any image generation model is like when you go a commercially safe route, at least the notion of what commercially safe is today, you're looking at hundreds of millions of images to train on.
24:40But if you go with publicly available data, then you're looking at billions and billions of images to train on. And so you can't have maybe both licensed commercially set of training data and output quality at the same time. Like, these are some of the things that are still up in the air. Yeah. Yeah. To be figured out. All right. And then speaking of boards, another update from FreePick. So they sort of been teasing this. They're calling it FreePick Spaces. It's not out yet. I think they have a wait list or some beta testing. but it's basically a board node-based version of FreePick. So similar to Weebly or similar to Flora, but inside FreePick.
25:23This is cool to see. Yeah, shout out to Joaquin. He was on our show a few episodes back. I mean, look, you and I use FreePick. We're not sponsored by FreePick. We just use it. We paid for it, but I like their model. Yeah, I paid for it too. and it's just such a well-designed aggregator of everything that's available today more or less yeah they've got the models in there right there and i mean like the unlimited image generation it just removes that mental barrier where it's just like yeah cool like i feel free to free to make picks exactly yeah i i'm also hoping that this does because the one weakness and we talked about this in our interview with joaquin is uh is um their collaboration tools like are kind of non-existent so like if you know we're working on multiple projects and we're trying to do images in the same feed they kind of their organization is kind of rough it's so like stuff gets mixed in together all the time it's kind of hard to find elements from other projects it's hard to collaborate with other people so i'm hoping the board features also the board feature also adds more kind of collaboration tools and features into free pick as a platform yeah i mean it keeps history of everything you generate, but it'd be nice to like put it into project buckets or something.
26:34Yeah. I was like, Oh, we're working on this project. Everything we generate should just stay in this project. And then if you flag it, it's like a select or something and you can organize it there, but just to keep better project separation. You know who does that really well is frame IO frame IO frame IO just as the way it sorts out projects and keeps like who should access what and so on. And so you mean just more from like a, like an asset management perspective. Yeah. Oh, yeah. Yeah. I like Frame.io for that. I was thinking, like, I was just chugging my brain, like, you could generate AI in Frame.io.
27:06You didn't know, Joey? Excuse me. Frame.io, my favorite model. Yeah. They dropped last night. I mean, they did finally add transcription to videos, like, a few months ago. So that was a cool feature. Yeah. All right. Next up, some other updates from Topaz. they dropped a couple new um upscaling models so uh one is starlight sharp so it's their newest model in the starlight family and it basically can upscale with an extra level of definition not present in any other offering yeah it's interesting that it's a diffusion based video restoration model so it's doing a lot of infilling with generative ai is what i'm guessing yeah i I mean, that's what their initial Starlight model did.
27:55So this one just seems like extra sharp. But yeah, the Starlight model is good for both synthetic data, like AI-generated stuff that you want to upscale and just restoring video footage that you want it to look sharper. Like I've run the original Starlight on some archival video and I do struggle just my brain of trying to decipher like, it looks good, but it's like, does it look really good and sharp and I'm just not used to it? or is it too in the uncanny valley weirdness phase and i've had a tough time with my brain figuring out like if it just like looks like it should have looked you know from upscaling like a 360 video to to hd or um uncanny valley weird but it's good yeah it's a really good it's a really good upscaler i upscaled some old vhs tapes of my childhood and um how to look the shots that it did really well in is like close-ups and medium shots but then when it's like super wide like if you have the beach and the palm trees but then you have the tiny people like you just can't figure out what yeah when there's just no detail there and the wide shots with like crowd stuff or details in the background yeah it doesn't it doesn't still doesn't do a great job because it just doesn't know what was supposed to be there well that's the thing i think if you integrate uh a vlm into the solve then it could look at the frame and detect that it's a beach and the beach those dots on the beach are people and then tell i think that's what they're improving on yeah so like it knows that that should be a crowd that should be sand that should be building skyline in the far distance that's like kind of blurry maybe they're working on it already i'm sure they are and the other model is um we are i guess we know we're we're aware of this yeah The other model from them is NYXXL.
29:42NYXXL. Yeah, that's one specifically focusing on denoising videos while preserving detail. So this is a cool, like, you've got some older footage or you've got, like, some super zoomed in iPhone footage that you're trying to sharpen up. You know, I like these kind of models where they're very, like, specific for, like, you've got a very specific problem with your video. Here's how you can fix it. This is probably in Topaz Video. I know they've been doing some rebranding because initially it was Topaz Video AI, which was a software you ran on your computer, and these models would run locally. I think they've been shifting a lot more stuff where you can – it's just a subscription package and stuff runs on the cloud, so you don't have to have a beefy computer to run Topaz anymore.
30:22Yeah, I mean, video restoration must be so GPU intensive. First of all, how many people out there have a computer good enough to do this? So you're significantly limiting your customer base if you just have a local version. And also the opportunity to mark up a cloud access and usage on the cloud. I mean, that's just like recurring revenue at its best. Yeah. I mean, now that you mentioned that, it's kind of funny that they sort of seem to have found this new use case, too, with so much AI-generated video that needs upscaling. Because before Topaz was sort of like this thing you knew about if you were like maybe in like the documentary space and you're like, we got this like archival footage.
Read the full transcript
31:02How can we improve the quality of it? Or you need to like fix it in video with issues. You know, now it's like has so many other use cases and like such a bigger market of users who just like, I mean, AI video, I need to like make it sharper, make it look better. Yeah, Topaz is kind of like a name that is saturating into the mainstream a little bit. I was talking to somebody outside of our industry the other day, and they knew what it was. And I was like, I was really shocked. Because most AI tools are stuff that you and I talk about, and nobody else has any clue about. Yeah, outside of, like, ChatGPT.
31:34Right. Exactly. Like, even, yeah, even Claude. I was like, I don't think most people know what Claude is. Speak to Claude. The painter? So there's this new integration with Claude Code and Figma. Nice. So through MCP, the Model Context Protocol, which we talked about a couple months ago, which is sort of like Cloud's protocol where AI can talk to other apps. This integrates with Figma. So the idea is you could build out your interface in Figma and just point it to Cloud Code and it will automatically look at your interface and build out your app. But honor the interface that you built. So it kind of solves front end design.
32:13You build out the front end. It programs everything. And give our viewers what Figma is for those of them who don't know what it is. I mean, yeah, Figma is basically a web-based version. It's a web-based design tool. Right. Or user interfaces and stuff like that, I would say. Yeah, I mean, it's sort of like Illustrator. 90 % of the user interfaces. I was saying it's sort of like a web version of Illustrator, but I don't know if Illustrator is a good comparison because people might not know what that is either. Oh, yeah, no, it's like Illustrator on steroids. So you can do page transitions when you click a button, what that button opens up.
32:52And so you can actually wrap up quite a bit of your user interface into Figma without ever writing a line of code. And it's really easy to use. There's like no technical hurdle to overcome. But then the challenge is how do you take this like really sleek user interface that you design in Figma and then actually implement it back into an application that you're building? so i think this is a game changer because most of the times ui designers aren't going to write code and then the people that write code don't understand ui very well right so this is bridging those two worlds really well i mean yeah and there's yeah there's a lot of jokes too where you like built out something that looks really cool in figma or your like design looks awesome but then when it actually gets built and turns into a functioning app or website that the design completely falls apart.
33:41Yeah, if you go to any modern website now that has the fancy buttons and as you slide up the website, things happen. Or if you're running any app outside of your Android app or iOS apps, the native ones, chances are it was designed on Figma. It's the number one tool for UX designers across the world. So yeah, this is cool. I'm excited to see how people use this. And also, I'm thinking, this is an example where it's like, oh, it's integrating with a popular app. What happens when that also integrates, like you have Cloud integrating with something that's more photo-related or video-related, and it can do powerful, automated, either rough kind of assemblies and sort of cloud-based energy features?
34:26Yeah, and it's funny you touched on MCP, Model Context Protocol. Like, the more and more I hear about it, the more it's sort of just replacing traditional API connections and going into that route. It's like a more sophisticated version of how two applications talk to each other. Yeah, I found pros and cons with them because I've been messing around with cloud code and I was trying to connect it to like Airtable. And I initially tried the MCP connection and it would just sometimes take a long time to find the thing because it's using AI. And so the advantage of MCP is you don't have to give it such defined data like you do with API APIs.
35:05They need a very defined data set and fields and all sorts of information. MCP could be like, yeah, find me this task on this thing or this record on this thing. And it can understand the abstract language. It just took a lot longer. And the thing I was trying to do was just automate a bunch of fields fast. And so then I found it was actually quicker and easier to just have it use API. And it built a script. I don't know how to do any of this stuff. I just tell cloud code, like, do this for me. And it's like, okay. And it built a Python script to connect to Airtable. Well, first of all, it's pretty cool that you took on such a technical challenge.
35:38And it's crazy that you almost got it figured out. I mean, I figured it out with the API. That worked. That worked faster. So now when it needs to update stuff, it just, like, runs its little Python script. And it's like, okay, and it's updated. So basically, I'm saying there's, like, pros and cons to the API and the MCP. But the MCP is definitely for, I think, stuff like this Figma integration. like you wouldn't be able to do that any other way. Yeah. And we're going to see more MCP. We've had that blender integration demo, like a couple months ago, which I'm, I haven't really seen anything new come out of that, but I'm sure we'll see more of that as well for like, yeah, you know, start, build out a basic animation of this character and just kind of builds out the founding, the groundwork.
36:17So you can jump in and take care and modify it, but at least you have a starting point. Yeah. I wonder why not like some of the more complex enterprise level tools like Houdini, Maya, Unreal, Come to Mind Unity. It's just ripe for the taking. Somebody just has to come in and build a really kick-ass user assistant feature. Oh, that integrates with these tools? Yeah. If you train a custom model on all things Maya, like learning Maya takes years, decades. What if you could shortcut that to like two weeks? Yeah. I don't have any disagreement to that. That would be very powerful. yeah i guess uh the question is like where's the money in it and why would you do it and when there's ai i mean if you could expand your user base but then it's like what do you do like are people just still going to get maya or is it like is it just better to build a more simplified 3d program yeah exactly that yeah like start from start from scratch i i don't know i was at the i went to the the comfy ui meetup uh yesterday at uh that was that uh fabric in uh in Venice.
37:23And there were there's definitely like a growing number of people in the crowd who are traditional M &E background, and are AI curious and are growing and like wanting to figure out like how to adopt these tools. And there was a good group of people who were like, traditional video editors. And then there's also someone who was like, VFX 20 years, you know, he's like, I was around, you know, during the formation of Nuke and all of these other tools and saw the shift to node-based VFX work. And then he's like, I see Compu UI. And he's like, this is the next VFX thing. This is the next nuke for the AI VFX age.
38:01So that was a very interesting comment from someone who's been around in VFX for a long time. That's very encouraging to see some of the tides shifting. And I think we talked about it before. It's like the folks that have expertise in VFX, I think those are the game changers of our world because they have the eye, they have the discipline, they put 12 hours in front of a computer, no problem, right? All they have to do is just learn this new stuff. Yeah, yeah, 100%. All right, and the last thing on our list of the AI roundup is Suno. The controversial music platform is back with a new version called V5.
38:42Now, it's kind of cryptic. We're having a hard time finding what's new with it. But it's better. It's better. And how is it better? With most audio models, as you can imagine, it's all about clarity, control, the ability to separate out the tracks, so have stems, the ability to have more control over vocals. so perhaps they have um a better vocal generation you know a cleaner higher frequency sound and things like that yeah um yeah i'm discussing yeah better better vocals better quality yeah i mean i found suno out of the box is pretty good at making i mean i only use it for making trying to make some background beats or something yeah i don't i mean the lyrics are funny i i i wouldn't That's one thing that Suno does actually well is it'll do, I think it'll do voice to voice.
39:36So you can give it, use like scratch reading or scratch singing something and it'll turn that into actual vocals. That's cool. Yeah, I can see that's like good if you're a musician and trying to kind of brainstorm ideas and or test things out more quickly. yeah but uh yeah i was people talking about the future but i saw spotify just came out with some new guidelines around clarifying some when songs were ai generated to sort of fight back on the flood of ai generated music without any human input so yeah you got this kind of battle of of sueno 5 coming out where it's probably gonna sound even better and more realistic and more uh higher quality and spotify yeah everyone else fighting against the flood of ethically i'm so torn ethically when i listen to ai music as well as looking at some of the ai quote-unquote movies but the music hits harder for me because um it's so much closer to real music i mean the music thing is also like it is a one button and you get something that you could listen to and sounds kind of decent where there's like no human input at all really and then if you Think about how long people need to learn an instrument proficiently enough to play it at a level where the output is usable for music.
40:55I mean, you're shortcutting all of that with just a prompt. It just doesn't seem fair. Well, the AI Suno band can't go on tour and play for you live. So if you're a real musician, you got the live show buffer. I mean, I feel like a bleak, I don't know, I'm feeling bleak today for some reason. So the dystopian view of the future is the AI band is just projected up on an LED wall with digital avatars. And you're just jam-packed physically with a concert grower next to you. And you guys are all looking at the screen with digital performers. No, we're all looking at the AI band inside the metaverse in our meta headsets.
41:42That's right. That's right. Yes. The meta Ray-Bans. Yeah. Yeah. Meta Ray-Ban fives. And then even worse, it could be like the version of that social media app where I forgot what it's called, but it's basically like a, an app. It was like a fake newsfeed where like you log on and it's like, you know, you post stuff and then a bunch of like AI bots reply to you and like kind of give you that, like that dopamine boost of like, you know, kind of going viral, but it's all, it's all fake. so it's you at your concert in the metaverse surrounded by the whole crowd but the whole crowd is um ai is bots yeah yeah it's it's you ai bots in the crowd watching ai music performance oh is that is that a bleak enough future it's way more bleak than my i don't know what you've been doing all morning but yeah okay you beat me to it i was just i was just going there it's just ripping that's great uh you know i'm kind of glad you didn't come over the house today.
42:38Log off. All right. Good place to wrap it up. Yeah. On that note. Thanks for everything we talked about at denoisedpodcast.com or down in the show notes on YouTube. If you're listening to this on Apple podcast, do us a tiny little favor, leave us a five star review. If you scroll all the way to the bottom of your iOS app, it's down there. Most people can't find it. But yeah, it's all the way at the bottom. All right. Thanks, everyone. Catch you in the next episode.
From the publisher
Alibaba unveils Wan 2.5, an AI model competing with Veo 3 that generates video with integrated audio and speech in a single output. In this week's AI Roundup, Joey and Addy analyze the flurry of new AI models and tools transforming media production, including Qwen-Image-Edit-2509, Google's Mixboard, Topaz's new upscaling models, Nano Banana's integration with Photoshop, and more.
--
The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.




