In short
Denoised Podcast Episode Summary
Episode Title
Is Nano Banana Google's New Image Model? Plus a World Model You Can Edit in Real-Time!
Podcast Description Denoised is a bi-weekly podcast hosted by Addy Ghani and Joey Daoud that discusses the intersection of AI, film production, and creative technology. This episode dives into new AI tools that are transforming creative workflows.
Episode Overview The hosts discuss several developments in AI tools for film production, including:
- Nano Banana: A potential new image model from Google.
- New World Models: Introduced by Tencent and Skywork, allowing real-time interaction and editing.
- FantasyPortrait: A tool for multi-character animation.
- Updates from SIGGRAPH: Covering innovations such as Meta's hyper-realistic VR headset prototype and practical workflow tools like Autodesk's free Flow Studio tier.
Key Discussions
- Nano Banana
- Overview: Suspected Google image model showcased on LM Arena.
- Performance:
- High realism and strong understanding of world context.
- Capable of generating consistent character profiles from text prompts.
- Comparison to previous models suggests an evolution in AI intelligence and image generation techniques.
- World Models
- Tencent's Foundational Interactive Video Generation:
- Composed of three modules: YON Sim, YON Gen, and YON Edit.
- Capable of generating interactive video environments that remember past interactions.
- Skywork's Matrix Game 2.0:
- Open-source, interactive world model that supports continuous video generation.
- Targets the gaming industry with real-time video rendering capabilities.
- FantasyPortrait
- Functionality: Allows animation of multiple characters simultaneously.
- Significance: Addresses a common limitation in animation tools where only one character can be animated at a time.
- SIGGRAPH Updates
- NVIDIA's Cosmos: Focused on automation and physical AI applications.
- Gaia: Generative interactive avatars with impressive animation capabilities.
- Meta's Tiramisu: A prototype for hyper-realistic VR headsets, emphasizing the evolution of display technology.
- Other Notable Mentions
- FAL Workflow Tools: Update to improve user experience in calling AI models.
- Autodesk's Flow Studio: Introduction of a free tier to attract users, while enhancing workflow efficiency.
- Color Grading Game: A fun tool designed to help users improve their color grading skills through a competitive format.
Key Takeaways
- The rapid pace of AI advancements is reshaping film production tools and workflows.
- New models are focusing on interactivity and user engagement, particularly in gaming and animation.
- The industry is adapting to integrate AI tools more seamlessly into existing workflows, enhancing creative possibilities.
- The hosts underscore the importance of engaging with these new technologies as they become integral to media and entertainment.
Conclusion The episode captures the dynamic landscape of AI in the creative industries, highlighting key tools and advancements that are set to impact filmmaking and media production. The hosts provide a thoughtful analysis of these developments, encouraging listeners to stay informed and engaged with the evolving technology landscape.
Links and Resources: For more detailed discussions, visit [denoisedpodcast.com](http://denoisedpodcast.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00All right, welcome back to the noise. We're going to do our Friday weekly roundup of all the new AI stories for film production. There's a lot? There's a lot. Yeah, it wasn't, at first I was like, I don't know, was there anything, there wasn't anything like big, groundbreaking this week. But as I was going through the links, I was like, actually, there was like quite a bit of interesting stuff and some new stuff under the radar that I think we'll point out that might be big in a few weeks.
0:24AI is leveling up incrementally every week and we're covering it. It's a lot of fun. I think it was a comment where it was just like, even like last week, it was just like, you know, like GPT-5, like not that big or something. Somebody had AI fatigue? Something like that. Yeah, I know. It's like, it's funny because it's like, we're so used to so many new stories where it's like any one of these stories a year ago would have been like groundbreaking new news. And now we just keep getting all these crazy updates. It's like, oh, okay. Yeah. If you think about like all the VC money and investor money, it's all funneled into this.
0:54And so there's people working on this constantly. All right. First one, Nano Banana. I almost spat out my coffee. All right. So I've seen Nano Banana pile up a little bit on X. And I was like, what is this? So this is a new image model that appeared on LM Arena, which is the main website where people, companies upload their new models. Yeah. And they get tested and ranked. And that's where the ranking stuff went. It's like this model ranked blah, blah, blah on this thing that's from LM Arena. It reminds me of the Coliseum. It's like an actual arena. Are you not entertained yet? Are you not generating?
1:30So there's a new model that showed up called Nano Banana. and a lot of people think this is a new model from google so it's not confirmed yet that it's from google but it seems most likely this is a new google model to either a separate model or something to replace uh imogen their current model and the examples that i've seen are very good world understanding and very realistic yeah this looking at this one that's yeah we're a side shot of a woman and it says hey have her look at camera and it looks like i mean everything about the person's face i don't know what she really looks like but it looks like an accurate rendition of like what this person would look like yeah the fact that you could do this prompt based with just text and not actually like you know put a spline or anything else in the image that's pretty impressive and then this one too where i think they gave the same image but then had it like hey just generate a complete character profile okay like a passport photo from this person and the consistency looks wow accurate that's impressive so yeah very similar to kind of flux context uh or chat gpt image this really impressive world understanding and just simple text prompting to modify yeah images yeah i mean to me it feels like it's so much more than just having a larger sample set and like more billions of parameters yeah it's beyond that it's starting to take on like a level of intelligence yeah and just understanding i there was i forgot the specific name but i was digging into because i was like the i've talked about how i would chat chat gpt image is like so good at just modifications for specific things and it does a special because when it generates you can see it generate like blocks like yeah like a like a regional matrix printing thing yes a printer matrix the yes if yeah i was yeah yeah it's doing a line scan line yeah there's a different method it was like a like pixel diffusion or it's a different diffusion method than the diffusion method we've talked about the most image generators use where it kind kind of takes the whole image and then diffuses it.
3:21It was a different method. And I wonder if this is adapting that as well. Interesting. So yeah, that would be a whole new architecture under the hood. Yeah. Yeah. I also want to point out, because I was looking up, trying to find out more info about Nano Banana. And it was just to show how crazy fast things are moving in AI and just more of the like vibe website and just opportunist market. Someone, I can't look at the exact date, but someone bought Nano Banana.ai and then spun up this website. called Nano Banana that is like, oh, hey, you could use our AI image editor. And I think it's just running on a Flux Context API.
3:56But the fact that they saw Nano Banana trending, spun up this whole website with all these features and all these things, which are most likely vibe-coded in literally less than a day is crazy. The power of AI to exploit AI. I love it. And also the color scheme and it looks like the logo. Yeah, it's on point. Google should just hire this guy. who knows how to vibe code. All right, other updates. So kind of a group full of updates. We talked about Genie 3 last week, and that was sort of the big. That was huge. Yeah, that was impressive to watch. I feel like now it is Revenge of the World Models, where there was a new slew of other companies being like, hey, we got stuff too.
4:37So this first one is Jan from Tencent. And it's a world model. They're calling it foundational, interactive video generation. And the interesting thing, I mean, it does a lot of similar things to Genie 3. You can generate from a text or an image prompt and start having this wall that you can navigate. It is persistent. It remembers things that it generates. If you move back to that thing, it remembers that it's still there. But it's composed of three separate models. Also, this one, it says it could run at 1080p at 60 frames per second. For games. Which is cool for games. Yeah. I think Genie was 24 frames per second.
5:11But the interesting thing about this is it comprises of three core modules. Yon Sim, Yon Gen, and Yon Edit. it so yansim enables high quality simulation of interactive video environments jen uses text and image prompts to generate the video but then edit supports multi-granularity real-time editing of the interactive video so you could be in the world and edit it which also reminded me going back to crystal ball of that demo of speaking of generating the world and speaking to it and modifying it in real time yes this is like the first actual practical implementation did they beat crystal ball to it i mean if this is out yet i don't know if it's a research paper Is that what I can say?
5:51Is that any way of saying come on to the podcast? Yeah. I mean, you could come on the podcast and we could talk about it. So with the editing, they have some demos of doing a whole style transfer in the world and they just change the complete style of it or structure editing. So like a character's moving and then you generate a wall in the world. So yeah, the editing thing, I think this is like the big. That was the promise of Gaussian Splats. You remember a few years back, it's like, no, we have editing capabilities. coming to Gaussian Splats. No, I don't remember that. Yeah, so like a lot of the Gaussian Splats startups were working on the way to get in while you're in the Gaussian Splats, you can delete this building or you can, yeah.
6:30But I guess it never really fruitioned. A really interesting thing here that I'm noticing if you look at the headlines, it says AAA level simulation, obviously super targeted at games. And again, it goes back to what I was saying during the Genie 3 coverage. Everybody's coming after the Unreal Unity space. it's like move over real-time renderer here comes ai renderer right it's essentially it's going to be and i think it's going to be like that scene in indiana jones where he puts the treasure in the sandbag like just does a swap that's what it's going to be like the consumer is not even going to notice that this game is generating ai it's just going to be what you're talking about i mean because there's two elements there they're talking about like the the world creation of like mapping it out and planning it and like the experience and then the rendering yeah and that the AI rendering part, that's DLSS.
7:16Like that's part of that where it, with any. Yeah, DLSS is not able to fully generate. It's just able to make something better. Right, but that's like, it's using AI to speed up the real, the quality of the rendering. Right. But it's like, that's based off, you have an existing world, you built something, and then you just want to look better on hardware and faster. Right. This thing, so the generating the world part, that's where, I guess, I mean, this i feel like from a game developer this ties in the same question of like the the filmmaking space where it's like okay well is it going to be a new game generated every time for the user right is it more like a prototyping method or like a initial world creation method and then you're saving that and you know storing that elsewhere yeah yeah i'm curious like i mean i don't know as much development yeah again we haven't taken any of this for a spin sorry folks um but my guess is like the guys that are developing this know that this is for triple a so of course persistent world generation is at the top of the list.
8:14Yeah, this speeds up generating a massive map. What's really interesting is that on the AAA level simulation headline, they're using something called 3D variational autoencoder. Okay, what is that? So remember the VAE block in comfy UI? This is a 3D version of it. So it's essentially going from a latent space. So the world is being generated in latent space. And now the VAE is not outputting a 2D thing. It's outputting a 3D thing. So I think that piece alone is magic because now you can attach that to other open models, like let's say Genie 3's latent space out to a 3D variational autoencoder, out to a 3D object.
8:53The problem with Gaussian splats was that you're still in a Gaussian splat world. It wasn't really meshing with a 3D world. Like it was just fundamentally a different type of asset, right? But here it feels like you're outputting something that could blend right back into Unreal or Unity or Maya. Yeah. Right. Yeah. As long as you get that mesh with something that you could keep using a persistent space. Yeah. I wonder where like inworld.ai is. There were a startup that was very early into the game to do exactly this AAA gaming with AI. And I feel like if I haven't heard from them, if we're not covering them on our Fridays, then they're kind of behind on all of this stuff.
9:32mm-hmm yeah and yeah i mean you know i think of like prime example is gta 6 which has taken forever to develop yes you know it's like well what if you could speed up making the map even something like gta where it's like sort of based on real world maps it's like the google map data and yeah speed up the foundational level of of building out these yeah and you can take something like ccm right which is uh so ccm ties into google maps and builds a 3d representation of the planet okay so then if you have like a latitude longitude coordinate uh of your game like you could be on this exact street in google maps and it's a 3d representation this house and everything is accurate so it's like a digital twin of the world yeah yeah yeah yeah if anyone else is more of in the game development world i would love to know what you think about this because like i'm curious how this not knowing as much about like triple a game development how this would fit in yeah i bet the triple a game folks are rolling their eyes right now nothing's gonna same way same thing with like when we see the r.i.p hollywood yeah exactly yeah there's like yeah this is not gonna this is not gonna make it i made a oscar nominated short film for two hundred dollars and then other one from sky work is matrix game 2.0 this is another world model oh wow but they're saying well they said this is the first open source real-time long sequence interactive world model long sequence meaning is just uh like a longer video maybe that it will we have that um the limit with the genie 3 where it could because it is how much memory you can save to keep the world persistent so genie 3 can only run for a few minutes um this one doesn't give us exact number but it says minutes of continuous video uh it's running at 25 frames per second yeah interactive move rotate explore yeah yeah so another world model this one's open source yeah I think these are great steps into this world.
11:20And, you know, over like the next year or two, we're going to forget about these early steps and just wonder over the achievements that we have made in this stride. You know, think of like this. I mean, think of how far, you know, image generation has come over a year or two. And this is like the, the, this is like the dolly to image generation. Do you, do you even remember a 5.25 inch floppy disks? You don't. Well, because the 3.5s took over and that became the standard but that was the gateway into that floppy disk in the same way like these beginning steps we're not even going to really remember the 5.2 were actually floppy the plastic ones were not floppy oh you mean like floppy floppy you could bend them they were floppy yeah okay what was the uh it what was 700 kilobytes 800 kilobytes per disc the black ones the the the plastic ones were 1.25 megabytes yes i think but the actually the black ones yeah they were like i don't think you're getting megs dude no no it was kilobytes like yeah 600 kilobytes or 700 it was enough to put doom on one disc you could yeah i remember i had this tool where i wanted to back up a file or something and it would split up the zips yes into one megabyte chunk so i could like do it was like put it in your first disc and it's like okay put up your second disc like okay all right we can talk about this forever i'll do one thing for you aol used to send out those free discs and then you just reformat it and then just have have another extra disc i don't know you can do that yeah i just use them as coasters they had a cool lord of the rings one lord of the rings promo oh really yeah when the movies came out oh that's awesome all right enough to stop it and then last well there's actually another world model update but this one um this one's not new we've talked about this i think like a month or two ago how do you pronounce it joey i don't know why my like american brain keeps not messing up uh hun yan yes i need to see order at a chinese restaurant
13:24hun yan's gamecraft which we covered like a month or so ago and was a similar-ish world model i don't think it was as persistent memory maybe it was anyways they open sourced it um so not new in the sense that this has existed but now it's open source you can go mess around with it yeah the chinese manufacturers tencent alibaba you know these guys are just quick to yeah they just crank and then release crank release you think there is like some bigger thing at play here like i don't know if it's just like the way to compete with u.s models or companies it's like well we'll make something similar but just push it out there so like yeah maybe more people adopt it because it's free and then yeah like what is their strategy like i don't know obviously they're not making money from any of this and this costs millions and millions of dollars to make i mean something like tencent this stuff's getting rolled into tiktok or will be you mean see dense by dense uh yeah yeah there's other models yeah yeah i don't know i mean yeah that's the thing i i mean maybe i don't know the chinese market enough uh to just know what their media landscape is but it feels like to us we're just getting stuff for free here yeah like a non-stop i mean i just don't know but just a free to adopt play to like get everyone using it.
14:33And then, right. I don't know. They charge later and that would make sense, but I don't know. Okay. I'm curious. Yeah. And from the, some of the AI researchers that I've talked to, like they're all impressed by Chinese models. Like everything 10 cent drops, even on the image video generation side is solid. Look, I talked about a few weeks ago, like C dance. 1.2.2. 1.2.2, which is, yeah, completely run locally. That one is like the top comfy run locally on your computer. computer uh video model yeah c dance is great i've used that for a bunch of stuff and uh i think it's getting more and more like people are picking up on that one it's been like oh it's like actually like a really good reasonably priced video generation model that can handle a lot of different things yeah i think uh caleb at curious if he's dropped that video where he compares c dense to vo3 okay supposedly that's better in a lot of ways yeah it depends on the thing i mean look i think vo3 is still sort of king in one shot generation because it's the only one i think now that could still generate audio with the video yeah so if you're just like trying to make rough if you're trying to make people uh-huh maybe there's no better tool because of the lip sync situation yeah for like a one shot don't have to mess around with it too much afterwards right vo3 is probably still can't get that but cdance for like giving a different inputs and getting outputs and also not spending three bucks a generation yeah really good okay other let's see other one multi-one yes okay fantasy portrait so this one's interesting uh this is from alibaba i believe it's open source but uh this is a character animation but with multiple people so that's sort of been a limitation with a lot of you know like runway act two or hedra or a lot of the uh character animator tools where if you give it some driving audio and like a still image and you want it to you know have the image start speaking the audio it's usually limited to your hey gen it's usually limited to like only one person in frame right at a time if you wanted to do two you'd have to kind of run them separate and composite it using traditional methods later fantasy portrait says that it can do multiple character portrait animations so this yeah it looks like it's at a research paper level and they have provided github code so all somebody has to do is grab this turn it into a product yeah yeah you can run and test it yeah so i got some videos of like the driving performance with a video with multiple people and then then image with multiple people and they're respectively getting the correct performance even this one with the presidential debate
17:03from the predict from the debate oh that's great they do have a sense of humor in those researchers and then it works in different styles um yeah i'm gonna try this one out because we're working on some uh speech detects projects yeah so i mean cool to see development in here because i feel like character animation and driving works on animals too talking dog talking pug yeah i mean this is all in like the deep fake territory all right like the big fear with ai remember a couple years ago is like what if somebody just makes a video of some politician saying something horrible yeah and that starts world war three yeah like i think we know now that that we could see sense ai versus not in a lot of these situations for now for now for now oh god yeah it's like that meme when it's like i'm seeing less ai generated images on the internet and then it realizes like oh it's because you can't tell the difference.
17:54I guess AI went away. Yeah, I think it goes back to what we've talked about before where it's like we're going to have to have the camera manufacturers for like the news gathering source is going to have to have some sort of verification in the files or something to verify legitimacy of actual meeting files. Sam Altman had an interesting take. Sam Altman was interviewed by Cleo Abrams. I think I covered this in last podcast. Yeah, so Cleo had asked like, yeah, what is going to happen when we can't differentiate from AI anymore? And his answer, more or less paraphrasing here, is like, we're just going to have to live with it.
18:27I mean, I think so. Yeah. Yeah. I mean, his argument was really on shaky ground. He said something to the effect of like, the iPhone is lying to you. Because when you take a picture, the iPhone is doing so much processing. And then by the time you see it, it's not really a reality anymore. In the same way you're accepting that that is reality, you're going to have to accept some of the AI stuff as reality. Yeah, I don't know. it is again he's out there dude i mean yeah he's got the whole world coin thing and the ubi stuff and the yeah okay uh sigguraphs yeah sigguraph happened this week the big i believe it's still happening i think it just oh no no you're correct because one of these papers is presenting today i've been getting uh pictures at midnight from parties people that i know are there having drinks they're like hey we're here i was like thanks yeah so it's like um uh the graph what's the official title of it it's a graph computer graphics yeah it is uh probably the most respected and oldest running uh conference of nerds of computer vision researchers traditionally it was computer vision and machine learning and now then in the middle there was a big rendering focus uh computer graphics focus and now it seems like it's a lot of ai pivot yeah yeah but focused on like the scientific and research part it's all the white papers yeah so well the graphics image generation ai image generation yes like very like very nerdy science part like siggraph moves uh media and entertainment forward every year so uh when those papers drop it takes a few months for those papers to turn into products and those products move the world I mean, you have people like Disney Research, who's like a big arm of Disney in Zurich.
20:11They drop their annual papers there. You have people like NVIDIA. They also do the same. Yeah. So yeah, there are a couple. I didn't see a ton coming out of it. Like, oh, it's super new, interesting things. We got a few things. But there was a opening talk from Ed Catmull, co-founder of Pixar. And someone asked him about AI and paraphrased the answer. AI will create lots of good things and lots of bad things. We don't know where it will go, but it's not going away. so artists need to engage with the technology. Don't we say that on Denoist? I feel like we kind of say that. A lot of the time, yeah.
20:41But yeah, wanted to recount that. But yeah, the two things that I saw coming out of it that were interesting, actually three things. One, in NVIDIA's opening note, they talked about Cosmos, which is their new world model. Yeah, they've had it for a while. Yeah, I think it was an update. But this one definitely seemed geared more towards automation, infrastructure, autonomous vehicles. Autonomous vehicles, physical AI with robots. That's really what it's built for. Yeah. The other one was Gaia. And this just popped up on my radar. Yeah, Gaia is pretty awesome, man. Yeah. And I think they're literally giving their talk right now as we're recording this.
21:16So maybe we'll have more info or the video next week. But it is a generative, animatable, interactive avatars. Very cool. Expression, condition, gaussians. Basically, it is a human looking person that's generated with gaussian splats and a bunch of sliders. and you can adjust them and manipulate the look of a person, but it's all generated with Gaussian splats, which is unique. Yeah, and I think we're going to be looking at the video as we're talking about it. So you could see that the quality there is pretty impressive, and as this woman is adjusting her expression, she's also shifting in parallax and perspective, and it's all holding up pretty decently.
21:52What did you say that reminded you of on music video? The Michael Jackson video. Black or white, was it? oh yeah i was gonna say i don't remember the name of the song i know the video but i forgot the name people have done a lot of recreations of that video with ai recently it goes just always harkens me back to like how did they actually do it in the 80s yeah yeah yeah they didn't have ai back then was it an early version of computer morphine i believe so yeah yeah but it looked great yeah i'm gonna say it was by pdi who did the term the terminator t1000 liquid going through the prison great like that was a very early version of digital human yeah okay yeah yeah okay so here uh my prediction is that this is potentially gonna change the way we do video conferencing and remote work okay how do you see that because uh if you look at the google beam product which is their really fancy conference system previously called project star line it's essentially two spaces that are connected by the internet and the reason it makes it so immersive is the other person is i i think potentially represented by gaussian splat so the google beam system has multiple cameras so it's almost like a volumetric capture of you and then it's somehow compressing and sending it over the network and then the other side you're being recreated with gaussian splats so if you like you know shift your head around a little bit like that you actually see like the left side of that other person's face right side of that person's face.
23:23So if NVIDIA pulls this off and releases this as a technology other people can adapt, let's say somebody like Zoom, then I think we could see potentially a real big level up in video conferencing. You think any M &E applications, like a gauze and splat work, or what are the limitations of doing actual visual effects with that? Very little. To be honest, digital humans are the most usable when they're rigged, when they're skinned. And then some of the Houdini stuff we talked about, you put muscle sins on. Meta-humans. Meta-humans, yeah. And I think that's where you get repeatability and high level of control.
24:03So I don't see this stuff really useful for primary character work. However, I think for crowd work, like background characters, stuff like that. Or something where it's like you need something more lighter weight real time. So that's streaming stuff. Yeah. Maybe the real-time live stream avatars. yeah or metaverse applications yeah yeah you know um remember when uh mark zuckerberg was interviewed by lex friedman and they were in the in the meta metaverse and like their torso and their face was represented by gauzy and splat or some form of volume however meta is doing it yeah similar with like uh the vision pro how it scans you to represent you on voice calls and stuff so both of those technologies although they were impressive at the time they weren't quite you know lifelike and i think this is a step forward in quality so imagine this in the metaverse yeah yeah i mean looking at the demo the details is great all right and then uh there one other thing i saw out of sigraph is from meta meta speaking of meta and it is a prototype next-gen version of their glasses, the display glasses, headsets called tiramisu, hyper-realistic VR.
25:15So high contrast, roughly three times the contrast of MetaQuest 3 with a wider angle of view and brighter. Those are all good things. Yes. And so some of the demos I saw on X, people trying it out were saying like, uh, things like, you know, excellent, wild. I mean, it's definitely a prototype because you look at the image right now. It is like, it looks like the cartoon with the coyote shooting its eyeballs out like it sticks out a lot yeah so it does currently look like this so it's a prototype but um yeah i mean wild to see you know where this technology is going yeah so it's interesting that they uh put the retina thing on there so do you know what retina is like what apple uses as retina is actually not retina i mean i know it's just a higher pixel density right yeah So in optical science or color science, when a display is quote unquote retina, that means you can't tell the pixels apart.
26:08Right. Like visually it's as lifelike as what you and I see perceive the real world. So if there is a display this close to your face, obviously you need far more pixels than a movie theater, which is like hundreds of feet away. so for them to achieve retina on a headset i mean you're talking 8k plus perhaps uh 16k in resolution do you know if that's more dense than the apple vision pro i don't know how the apple vision pro i don't think is retina oh you don't think yeah they may use that term for marketing but in in the sort of truest sense of that word it hasn't for the truest sense of the word is it a combo of the pixel density but also the like display distance yeah both because like it's an equation right it's not like a set number because it depends how far away you're would normally be looking at the display yeah so here they're saying 60 60 pixels per degree so you can imagine one degree of your eye at that close resolution is not that much real estate right so you're looking at like a maybe a fraction of a millimeter and within that millimeter you need to have 60 pixels so you're talking about an extremely pixel rich environment yeah i know there's some other headsets that were like flatter too but i don't know the specifics of them i'm trying to scan but uh they're called boba one or boba three but it's basically more of a flatter version um of the headset yeah yeah yeah i mean nobody knows that the biggest barrier to headsets is the fact that you have to wear headsets more than meta yeah yeah i mean it looks like the boba has the same style as the the sony adapted with their professional headsets were sort of like a flip down flip up flip down somewhere in a professional environment if you're like going back and forth you could just keep it on your head and just flip the screen up right which yeah you look cool for a hot second i mean it works if you're in like an industrial environment and you're like i need to like look at something or you know and flip it up and down it's like a welder's glass exactly and the coolness factor isn't as big just don't do it on a plane we talked about this don't be a glass hole But yeah, did we talk about how like for some reason that guy like had no reason to be putting his arms like everywhere?
28:19You could just like sit there and probably pinch and operate. I didn't know you could do that. But yeah, he was grabbing stuff, you know, his hands in the aisle. We were all trying to get in the plane. Yeah. All right. What else we got? Skyreels. Skyreels. So yeah, Skyreels popped up. This is similar in other audio to character animation. At first, when I watched their demo reel, I was like, I don't know. It doesn't look that good. but then i saw some of their other tests and examples and some of it looked did look the human performance on top of an image look pretty good it's i think the company is skyreels it's another ai company i haven't used them um i don't know if you use them no yeah so i mean i might give it a shot because like i mentioned before we're kind of project with some uh audio to to character animation but yeah i'm starting to put on people's radar uh because it looks like i mean you know like all these things it probably works well with some stuff not so well with other stuff Yeah, challenges with video generation is length.
29:12The longer you go, the more you start to hallucinate, the more memory you use, and there's all types of problems. Yeah, this one did say, I think, that it can run on longer videos. Oh, you know, actually, also, this reminds me, I forgot to save it, but one other update this week that was in the character realm was Pika. So Pika also rolled out a new update. I think it's just coming to their app, though. We haven't talked about Pika in a while. No, we haven't talked about it ever since when we had a little octopus run around the table. yeah it's uh yeah their own version of a character animator uh where you could give it a audio or a performance video sort of like their own runway act two kind of feature right but it sounded like it was just rolling out on their app like i think they're focusing more on mobile social space and mobile and creator space than yeah the desktop space yeah they're probably looking at the user data and maybe most of the generation is happening on a mobile yeah you know and i mean that's probably just you know talking to other people about this stuff it's like a bigger addressable market for for sure these companies for sell ten dollar subscriptions than m &e space where they're going to still pay the same subscription but demand more when when i think desktop i think you and me like yeah people that are trying to do professional work versus mobile everybody oh right right yeah or the or the generation that grew up native on like cap cut or stuff for video editing Exactly.
30:30Video creation. What's Premiere? Oh, you mean Adobe Rush that I never use? I use CapCut. So FAL, who I think I've talked about them before. I'm a fan. I'm a fan of FAL. Yeah, they basically every like pretty quickly between FAL and Replicate, they'll take all of the AI models, put it on their own servers, and then you can just quickly call them up with an API and you just pay per the usage. So if you don't have a powerful computer or some of these models, like you just you can't run them on a computer anyways. They're like a good one stop shop. You get like one API key, one central billing, and you could like run and use relatively cheap to AI model.
31:04Yeah, they're usually either just give you the rate that the API from the company is charging or if they're hosting it on their servers, it's like a pretty reasonable compute rate. Yeah. One of the I'll talk about in the future. But one of my pet projects has been trying to bring in file APIs in a comfy UI so that you could use. That's a good one. You can call them up. Yeah. So you're doing everything on the cloud, but still have a workflow. Exactly. You build your workflow and comfy, but call up file APIs. Because there are some nodes that already exist, but they haven't been updated. And so file updates stuff all the time.
31:34And I was like, I just want a way to bring in my own APIs. So that's a pet project I'm working on. I'll talk about it later. But file did just release a big update to their own canvas based editor. Much needed. Much needed. Yeah. So this looks cool. It's called workflows 2.0. And it is basically, you can drag and drop all the file nodes and apis and build out your own workflow and then just run it and file so they completely stomped on your work a little bit yeah i mean i still think there's advantages to comfy because uh you could the file saved your computer and if you're doing a bunch of batch stuff you can just have them all saved locally but this also looks like a great solution and i don't think there's any additional cost i think you're just paying for whatever for the generation apis you're running but it's just a really nice workflow and now they're getting into the invoke uh ltx flora territory exactly yeah where it's like a very comfy ask ui without the technical barrier of comfy if that makes sense exactly yeah there's very a the canvassy click and drag thing yeah it's just nice that they have this built in so you could an easier way to use their api models yeah i mean we talked about this when we did our comfy ui episode it's like i don't understand why comfy doesn't do this uh as a cloud service or as a generation service what do you mean like they You can easily build a Comfy on the cloud that connects to a bunch of GPUs.
32:56Oh, why they don't have, right. Because there are third parties that have Comfy as a service on the cloud. Right. Why Comfy doesn't do it. I don't know. Yeah. I mean, Comfy has the API nodes. They're the most ready for it than anybody else because they're the universal standard in the world of AI development. Yeah. They've sort of been around the longest. They've cornered that market. Established. Yeah. I don't know why they don't have their own cloud hosting service. Yeah. I would imagine like services like Fall, Invoke, or anybody in this competitive space, I wonder if they take the JSON files that the Comfy workflow is in and then start to build that in their own environment.
33:32Oh, I see. You know what I mean? Do you even need that though? Because I mean, most of these APIs are pretty self-contained. It's not like how Comfy works with the local model with the clip and the VAE. Yeah, you don't see the clip or the VAE here, right? Yeah, when they call up the API, the API just already does it. Yeah, true. So yeah, I don't know. Okay. Yeah, I mean, I think it's a cool tool if you mess around with it. I know we turned someone on to, not invoke, what's the other one? That's not Flora, not invoke. There's a... The other workflow chart. Yes. Canvas tool. Yeah, actually the one that we're thinking of is I think the most powerful one because it's layer-based and node-based.
34:08Yeah, with the canvas-based tool. So I think it was invoke. Someone reached out and said that we actually turned them on to invoke and that they've been able to build out a bunch of like cool concept commercial videos using these all-in-one platforms. Invoke seems powerful and purpose-built for our industry versus like for social media content or anything like that. Yeah, yeah. To give you a lot more control and like a one-stop shop for building out these flows. And especially if a node-based workflow kind of works better with how your brain operates. And also Invoke has layer-based stuff on top of the node-based stuff.
34:37So like layers, like you could layer the images? You can have masks and stuff. Oh, okay. Regional in-painting stuff. Oh, and then run different functions on nodes to like that specific section. Yeah, that's cool. Yeah, that's cool. All right. Other updates. Actually, this did sort of come out of SIGGRAPH, but I don't know. This is I just realized on this press release thing, it says SIGGRAPH 2024. And I'm like, is this the right? This is an article. No, this article did come out this year. Basically, Autodesk has released a free version of Flow Studio. So Flow Studio was Wonder Dynamics. Wonder Studio that they acquired.
35:07They rebranded it to Flow Studio. So Wunder Dynamics is a great tool where you could give it a video and then it could automatically map and replace and rig the video to a character, actual character. And then you bring that into your software, professional 3D software, and use that as like a first pass to start animating. I could be wrong here, but I don't think Wunder Dynamics or Flow integrates very well with Maya. From what I heard, since the acquisition, Wunder Dynamics is still not super well integrated into other Autodesk products like MotionBuilder Maya. Okay, yeah. You would run it in, like it's a separate product and then export.
35:46Export and import. Yeah, I don't think there's a direct path, which is a shame because Autodesk is certainly a company that's capable of doing something that's really, you know, well executed, elegant and available for M &E. Yeah. so they have a new free tier actually i'm looking at this pricing tier because i haven't checked it out since i used wonder dynamics last year and it was wonder dynamics was like either you could try it or it was 100 bucks a month there was no in between yeah which kind of pricey um so it does that they do have a bunch more tiers uh the pro tier is 95 a month so now they have a free tier where you can run it you get 300 monthly credits i forgot each generation is like whatever handful credits and some like duration caps but oh no it is point of marked so you're gonna say it's not watermarked it is watermarked but you can let's mess around with it you can export stuff yeah you can export to other scenes in the free tier i'm guessing if it gives you the files the files aren't going to be watermarked yeah i'm also guessing that this will help drive adoption my guess is they're having a hard time with you know recurring usage usage and daily active users so this is something that a freemium model if you will yeah it's good to be able to mess around try it and they have two other plans in between the pro hundred dollar month plan they have a ten dollar light plan and a 45 dollar standard plan so yeah it's good too that they have more in between options to to get more usage it also goes to show that no matter how brilliant your ai technology is end of the day it's still a business and a business has to generate revenue offset you know margins and all that stuff you need money to keep developing yeah to run this stuff so because wonder dynamics was one of the early ones right you remember they got acquired almost two years ago now.
37:23Was it two years? Yeah. Yeah. I'm going to say, yeah, it was pretty early on. And so their life cycle is actually much further ahead than Runway or Luma's. And they've actually gotten acquired by a big company versus Runway or Luma. And you could see the interesting sort of curve where how AI company matures. And now they're at a maturity point where they need to generate real revenue. Yeah. Yeah. And so that's where they're at. The company, Autodesk Bottom. um it's like well no now we acquired you yeah we need the roi yeah we need this acquisition to to make money yep all right another quick one uh gemini is now also rolling out memory i think we talked about claude added memory last week memory is one of the best functions like that's the main reason i go back to chat to et is because i remember stuff yeah but now claude remembers stuff and now gemini remembers stuff so that makes these way more useful you think they're using floppy disks there's someone just in the back just like we need more swapping the discount yeah that's how it remembers sorry about that joke folks all right the last one this one's a fun one a colorist sort of vibe coded i don't know if they actually vibe coded but i'm gonna say that i'm gonna assume they vibe coded by a montanari louhi sorry if i said the name he built a color grading game and it's called match the grade and he basically shows you an image of like some shapes and colors uh in one side and then you have a couple sliders and you're trying to match source to match the grade and then you check it and you try to and it gives you a score based on how well you matched it it's fun yeah i mean you know we all could use more color exercise to sharpen our eyes oh for sure yeah just even to learn or start about this yep so yeah i think that's just a really fun uh example use to uh this is the best score i've ever gotten should we make an episode oh 95 is great it's literally just like the image was way brighter i messed around with this before for like it has a 60 second time limit and i messed around with it before and um i was getting like 60 70 and now i'm gonna beat you doing this on the show i get 95 this was a very easy image to do so yeah anyway shout out to dubai this is a fun game and a really good use of um yeah just i'm assuming vibe coded but of building something fun and practical we should do a episode of you and me playing this competitively for an hour most boring episode whatever we just live live stream yeah on twitch comment if you want to see that yeah let us know because i'm thinking no but if you're into it we'll do it we'll do it if you guys want us to yeah all right that's pretty much a wrap up for this episode uh links for everything we're talking about as usual at denoisedpodcast.com thank you for your support on youtube hope you got to see free pick episode with uh ceo joaquin yeah check out the interview give that one a watch yeah all All right, thanks, everyone.
Read the full transcript
40:07We'll catch you next week.
From the publisher
AI tools for film production are advancing rapidly, with multiple new releases this week that could transform creative workflows. Hosts Joey and Addy dive into Nano Banana (a suspected Google image model), several new world models from Tencent and Skywork, FantasyPortrait's multi-character animation capabilities, and updates from SIGGRAPH including Meta's new hyperrealistic VR headset prototype. Plus, practical updates to fal's workflow tools, Autodesk's free Flow Studio tier, and a fun color grading game.
--
The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.




