Comfy Cloud, AI Camera Control, Nano Banana 2 and more AI updates

16 Nov 2025 · 40 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Denoised Podcast Episode Summary Episode Title: Comfy Cloud, AI Camera Control, Nano Banana 2 and More AI Updates Hosts: Addy Ghani, Joey Daoud Release Date: [Insert Release Date] Podcast Description: A deep dive into the latest trends shaping media, entertainment, and creative technology.

Episode Overview In this episode, Addy and Joey discuss several recent developments in AI and their implications for the media and entertainment industry. Key topics include the rumored Nano Banana 2, Comfy Cloud's new workflow, and various updates in AI camera technology and video manipulation.

---

Key Topics Discussed

  1. Nano Banana 2
  2. Rumors and Leaks: Discussion around potential leaks regarding the Nano Banana 2, an AI tool with implications for media authenticity.
  3. Implications: If true, the updates could enhance the tool’s capabilities significantly compared to its predecessor, which already allowed for remarkable functionalities.

Public Perception of AI-Generated Media

  • Deepfakes: Concerns about public recognition and skepticism towards AI-generated content.
  • Political Manipulation: The potential misuse of advanced AI tools in political contexts was highlighted, with speculation on the public's ability to discern genuine media from manipulated content.
  1. Comfy Cloud
  2. Overview: A new browser-based service that offers an affordable option for creators, priced at $20 per month with up to 8 GPU hours per day.
  3. Features:
  4. User-friendly interface with preset workflows akin to the desktop version.
  5. Ability to upload and modify model workflows.

User Experience

  • Initial Feedback: Positive impressions from users regarding its affordability and functionality compared to handling local hardware setups.
  1. Veo 3.1 Camera Controls
  2. New Features: Enhanced camera positioning options allowing for dynamic camera movements within generated shots.
  3. Technological Impact: This feature could significantly change how creators approach video production using AI tools, providing more control over the final output.
  1. Industry Updates
  2. Acquisitions:
  3. Figma Acquires Weavy: Figma's move to enhance its AI capabilities in the design realm by acquiring a competing technology.
  4. Adobe's Acquisition of Invoke AI: Reflects the ongoing trend among tech companies to bolster AI tools for creative professionals.
  • ByteDance Video Upscaler: An affordable tool for upscaling videos, priced at around $0.0072 per second for HD, with implications for post-production workflows.
  1. Research and Development
  2. Cam Clone Master: A new tool that allows for precise emulation of camera movements in AI-generated videos.
  3. Infinity Star Framework: A proposed model for unifying spatial and temporal data in AI video generation, potentially enhancing video synthesis capabilities.

---

Key Takeaways

  • AI's Role in Media: The hosts emphasized the dual-edged nature of AI in media, where it can be used for both creative storytelling and political manipulation.
  • Tools and Accessibility: Innovations like Comfy Cloud and affordable upscaling technologies are making advanced tools more accessible for creators, heralding a new era in content production.
  • Future of AI in Production: The conversation hinted at a trend towards real-time control and user-friendly interfaces, which could revolutionize how filmmakers and content creators operate.

Final Thoughts The hosts concluded the episode with a sense of optimism regarding the advancements in AI tools that empower creators while warning of the potential pitfalls these technologies bring. The discussion was lively, engaging, and filled with insights pertinent to industry professionals.

---

Additional Resources

  • [Denoised Podcast Website](https://denoisepodcast.com)
  • [VP Land Newsletter](https://ntm.link/l45xWQ) to stay updated on the latest news in creative technology.

*Note: The views and opinions expressed in this podcast are personal and do not necessarily reflect those of the hosts' employers.*

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Alright, welcome back to Denoise, a roundup of new AI tools. Ready Addy? ai roundup let's get into it

0:11all right let's talk about nano banana again there are no these are all rumors and i would put half obviously there's not a banana 2 in the works there are supposedly not a banana 2 leaks online i don't know how true these are but we'll talk about it what uh what did you see adi no just um A couple of the YouTubers that I follow dearly on YouTube that cover a lot of AI news, there was glimpses of Nano Banana 2. Not sure if it's true or not. You know, when we first covered Nano Banana, when it came out, there was like a fake website around it. Do you remember? Yeah. I mean, it was called nanobanana.io.

0:49Or something. Yeah. Or something. I don't know what to trust anymore, but if the news is true and Google is working on an update to Nada Banana, oh man, the implications for that. You can already do so much that you couldn't do six months ago. So if there's an update and you could do even more, I'm just curious to know what that is. Yeah, I mean, in this tweet from Roberto Nixon that may or may not be BS, it is apparently leaked images from an uncensored model. And so it's extremely photorealistic images of like Trump taking a selfie. Yeah. Snapshot looking thing of P. Diddy and Jeffrey Epstein.

1:34CNN screen grab with extremely detailed graphics. Wolf Blitzer. with will blister and donald trump getting elected for a third term all of the chyrons and everything are like accurate the date is accurate like it looks like a screen grab i mean um if this is true yeah somebody could just easily photoshop this and build hype around and get some clicks yeah the leak apparently came from media.io like one of the ai aggregator tool aggregators that I've never heard of yet. Again, it's always like, this happened again, OVO 3.1. It was leaked from some aggregator sites that I was like, I have not heard of you.

2:17You think the examples that they show, like the Epstein-Diddy thing and the Trump winning a third time thing, why is it always political? I don't know. And I think the public - Is it the worst case scenario of the worst deep fake looking thing that you do not want on the internet to like confuse even more than things are confusing? Do you think the average public by now knows that stuff is deep faked and AI generated and that they should be kind of cautious about what they see and consume? Define average. Like the most average Joe Schmo out there, you know, lives in. Average? No. We're looking at the bell curve medium?

3:01No. I mean, there are those crazy sore videos on Facebook that boomers 100 % think are true. Okay. So this could have potentially a big political impact if it's used correctly. All of these things could. Yeah. It's just getting crazy. And then reality is crazy too. I mean, there are real images, like the guy fainting in the White House. That image looks fake, but it's a real image. we just dude we just need nicole kidman back in here telling us what's real and what's not sorry yeah so no i i i yeah i just i have a tough time questioning reality as well too yeah there was a what did i just see i just saw some tiktok video of some there's a crazy over-the-top parade somewhere i don't even know michigan or somewhere north and they do these really elaborate floats and then I saw a video of floats that were scenes from the Titanic sinking.

4:02And I thought it was another Sora video because I kept seeing Sora videos of the Titanic amusement park that were fake. But I think this one is real. Okay. So I think, I don't know, I think eventually we're just going to have to give up on believing anything we see online. Do you think the, follow up to my question, Do you think the content protection stuff and all of the AI checkmark stuff is going to have any impact at all? I hope so, but it's up to every platform adhering to that. And I've seen even some tests of, yeah, you could, if you build Adobe products, and Adobe's pretty good, it has this tracking and content protection or authenticity and giving you the history of the image.

4:50But if you take that image and put it somewhere else, it just strips all that metadata. and then it's just a regular image. So that could help and work if all of the platforms sort of agree to keep the metadata and display stuff around the metadata. Maybe that'll happen, maybe not. I have no idea. On the flip side, it could also help to have some sort of authenticating what was real images, what is real images. but you know again it's like when when does that line come off of depending how much you modify the image or edit it and then if you shoot something with a real camera but then you edit it in premiere and then you export it is like it's still going to have authenticity metadata in it i don't know yeah yeah i mean i have an optimistic take on this what's your yeah please so we're going to definitely see a lot of like political mishap and all this stuff that could potentially harm us and the planet, right, and put us in the wrong direction.

5:54Having said that, I think the people that are the best at this tool at the moment, the experts with AI generation, image and video, are our people, are film and TV people, media and entertainment people. And they're probably not going to use the tool for anything like that they're they're rather going to use it for creating stories making something that wasn't possible before and it's everything that xavier talked about brin talked about it's like enabling the single creator to just do something amplifying something that they couldn't do before and once you have enough of those examples where you're like how that film was actually really touching and moving and it was used you know it was built by ai using two people you have enough of those And then in time, the public will sort of label generative AI as a storytelling tool, as a filmmaker's tool and not a political manipulation tool.

6:57I think the latter will go out of fashion very soon and it'll be quote unquote cringe to use it in that way. I think maybe, I think you're thinking of like the kind of meme-y hobbyist use case. I'm thinking of like the bad actor use case, like foreign states to like mess up the election. Bots, scammers, you know, like that, that, that crazy scam story with the worst photoshopped Brad Pitt images, uh, where he's in the hospital and like some woman was scammed out of like half a million dollars. Cause she thought she was like chatting with Brad Pitt and he was like in the hospital and needed her help.

7:34And the images were like his face cut out on like a hospital bed. So the images were terrible, but imagine those images actually looking realistic. So, yeah, I'm thinking of a bad actor case and how easy these tools are to use. Yeah, we don't know. I'm glad we're able to talk about it. Viewers, if you have comments on any of this, we'd love to hear it from you on the comments. Yeah, let us know. I mean, nobody's right or wrong here. I think it's just TBD. It's just too soon to be set in stone. All right, let's talk about something more positive. Yeah, let's. Comfy Cloud. Comfy Cloud. I just used it this morning, actually.

8:10I am very positive on this. It's exactly what you'd expect Comfy to be. It runs on a browser. It has all the nodes, and even the keyboard shortcuts are the same. It's quite affordable. Joey, how much are we looking at? Yeah, according to their business, so it's a public beta, so it's still obviously they're working out the kinks. But right now, Comfy Cloud is available for$20 a month, which includes$10 credit every month to use their partner nodes or the nodes that are the API nodes we talked about before where they just charge you the base rate of whatever they charge. And then up to eight GPU hours per day.

8:45That's pretty dang cheap, if you ask me. It's pretty good. Yeah, I mean, just$20, and basically it's sort of like, yeah, up to running eight GPUs. Yeah, I don't have eight GPUs at home. No, yeah, and it's powered by A100, 40 gigabyte GPUs. So, yeah, I think that's really fair pricing. Absolutely. The nice thing about it also is once you launch ComfyCloud, like your browser kind of launches itself and then uh all of the workflows are presets so if you have like a for example image text to image with nano banana that's a preset you click on that it opens up the entire node tree for that workflow or you could do a video one and then you can go in and just like regular comfy you can upload a laura you can delete nodes add nodes change the resolution like you have full control over i don't know if you got that far down the rabbit but how does that work?

9:36Does it already have the model? If you load up a default comfy workflow, does it already have all those models pre-installed? And then if it didn't have the models, can you download other models now? I haven't done the latter yet, but the most popular models, I'm doing some stuff with 1.2.2 Animate, which we're going to share with you guys. It's such an intricate workflow that's already built up there, and it actually combines a YouTube video with that workflow. So you can go watch the video, figure out how they built the nodes and then come back to the nodes and modify it. It's pretty cool. Yeah, and that's a great way to start because they built out these custom workflows that are ready to go and you can play with them and use them and they have great documentation of like put your image here, put this here.

10:19But then if you want to start taking it apart and modify it and rebuilding it, that's there for you too. So they're a great way to start and learn. Yeah, actually I just answered my question on the blog post. Coming next, upload and use your own models of Laura's. So I think right now it's just sort of the preset. It's enough. I mean, for 90 % of things, they're going to use those presets. I wonder, yeah, I'm curious how that would work. Because some of these models are 10, 15 gigabytes. Like, is it just, how do they store it? I'm just going to, I'm curious how they're going to figure out if, like, you have to pay for additional storage for models or how that works.

10:50I'm guessing they're just going to link to, like, the GitHub repositories of the models. So you don't actually have to download and upload. You just tie the GitHub to... But you still, the model still needs to download together. Yeah, I think they'll handle that part. Yeah, that's my guess. Oh, you think they'll just pull it in, use it, and then toss it? Yeah, if they have a direct hugging phase integration or something. Right. Maybe. Yeah, so I mean, yeah, that's great, and it just makes it super accessible. Do you know, can you use it on, because I remember when they first announced it, they were like, oh, you can still use local files and run it on desktop.

11:27Do you know if you could still use the comfy desktop version, and And just instead of running the workflow on your computer, you're like, send that workflow to the cloud and process it. I don't think it's hybrid like that. The comfy cloud is like a whole different thing than the desktop application. Yeah. One feedback and note I do have is right now I'm trying to generate some videos. And then the workflow compiles correctly. Everything is good. It goes into the case sampler. It starts to churn those GPUs. and then it crashes, which is fine because desktop comfy crashes all the time, but it doesn't tell you what the errors are.

12:07Whereas the desktop one will tell you, this node, da-da-da-da-da, and then it's a troubleshoot. Yeah, I'm sure that's coming. And yeah, this future pricing model after beta, all plans will include a monthly pool of GPU hours that only counts active workflow runtime. You'll never be charged while idle or editing. That's cool, that's a good model. That's good. Yeah, I mean, pricing seems super fair. I'm a big fan. And if you're doing comfy on desktop, I highly recommend you switch to comfy cloud. Or maybe just wait until the beta is over. Because, look, you're never going to have eight H100 GPUs.

12:40I mean, the power you have over the cloud compute is insane. I was a big believer in local GPU in the past. I have a pretty solid NVIDIA RTX that I invested in. And now it feels super dated. Like, I don't really need it. Yeah, I'm kind of like, I really want to wait five minutes, ten minutes for something to generate. I used to wait. I just hit it in the morning and then come back in the evening. You don't have to do that anymore. The novelty of being able to do something locally. I'm like, all right, whatever. I think the local stuff is still cool for image generation. Image generation is still in the few minutes range, but for video generation, forget about it.

13:18All right, another update. Another VO 3.1 flow update. This one's cool. This is one of the stuff I've really always wanted. camera control, camera position. So there's a new update in Flow, which is Google's web service to use VO3. And you can give it new specific camera control. There's a drop-down menu option to reposition how the camera behaves inside a generated shot. Greg Frazier's probably rolling his eyes looking at this like, that's not right. Yeah, some of these names, I'm like, I would not have called it that. So yeah, there's camera position, move up, move left, move closer, orbit, dolly, and then there's camera motion.

14:04But this is great. This is technologically really impressive. The whole spatial awareness and consistency as the camera moves up and down without any distortion. Again, the example we're seeing here could just be a really polished demo. But if this is actually at this quality level, that's a big game changer. I will, yeah, I will test it out because I think it's only on the ultra plans right now. But also the other thing to keep in mind, and I found this out because there was that other feature that came out with 3.1 that was you could take an existing, or you could take a video and you could like mark an object and be like remove or like replace.

14:40These features only work in clips already generated inside VO3. So you have to generate a clip first, and then these tools are only to modify already generated clips. You can't upload a video clip from another source and then be like, I want to change the camera position or I want to modify an object. This only works on stuff you already generated inside VO3, inside Flow, to further refine that video clip. So keep that in mind. I did not know that at first. I thought it was like Runway Olive where you could upload a clip and then be like, yeah, change this, change that. But you have to generate it in VO3 first.

15:16It's probably looking at all of the metadata and the intrinsic assets in that generation in order for them to do the camera move. So it's probably looking at the prompt, the seed, and the CFG, and all that stuff in order to then take it to the next step. Yeah. I was wondering if it was content safety stuff, too. If you're able to generate it in the first place, it already went through its content moderation. That stuff is pretty easy to block. You know, I would say it's not ready for user uploads yet because it's just not robust enough to do so. That's my guess. Okay. Interesting. Yeah, maybe.

15:49Either way, I mean, VO3 is already a pretty really good tool. So, yeah, that they added more of this stuff in there, that's super cool. I'm going to mess around with that. All right. More Node updates, sort of, on the business side. Oh, yeah. I'm loving this. Exciting news on the creative tools front. So Adobe, who we covered last episode, was the big Adobe Max, the fact that they acquired Invoke AI. Well, guess what? Adobe has two big competitors on either side of the business. I would say on the consumer side, it's more Canva and sort of replacing Photoshop and things like that. And then on the business side, it's Figma.

16:26Figma is right now the standard go-to tool for any user interface design, any user experience design. So Figma just acquired Weavey. Weavey is a competitor to Invoke. So Adobe picked up Invoke, Figma picked up Weavey. And now they're more one-to-one and straight-up competing with each other on AI node-based workflows for image and video generation. Yeah, it's like all of the established companies started to build up their AI toolkit or either building or acquiring existing node-based workflow tools. which is sort of been one of, like, sort of once you kind of level up and need to do more complicated things, nodes.

17:12Yeah, and Figma has historically been, like, their whole ethos is simplicity, right? Photoshop, if you've ever used it, or any Adobe tool, like, it's really difficult to master because it just starts to get really complex really fast. With Figma, you can actually still retain that level of simplicity and do a lot more with it. There's less clicks, less sort of, there's less effort to get to the level of complexity that you need. So them acquiring Weevy, I think, is an indication that they're going to bring a whole new type of node-based workflow that's far simpler than Comfy UI or perhaps even Invoke.

17:55Yeah, I'm curious to see where these go and how they kind of all integrate it because it's like they're already buying established tools that already have their own workflow and use cases. Like, how are they going to just blend it in? Are they going to kind of keep them separate? I'm kind of curious to see how these, well, we know Invoke, they sort of shut the website down. So that'll be somehow probably, I'm going to guess, built into Firefly. But yeah, we'll see how that plays out. But yeah, Weevy, I'm curious if they're going to keep Weevy separate or roll it in. If you're a big tech company, there's still a few other players left for sale.

18:23So I think we're looking at LTX Studio. FloraFauna, LTX Studio. and Joey Comfy, Comfy UI. It's true. I mean, that is the ultimate tool at the moment. Yeah, that is. I'll be really sad to see Comfy go the way of big corporate takeover. I know. I think Comfy is so attractive, too, because it's just open source. Yeah. It feels very like early days Linux, you know? Yeah, a bit, sort of. Maybe slightly easier to use. Yeah. All right. And then a couple kind of papers slash, you know, new under the radar and models popped up. One is called Cam Clone Master. And this was at SIGGRAPH Asia. And this sort of addresses one of the things where I'm always like, I just want to be able to have a camera and shoot a shot.

19:12And like that is the camera movement for the AI video. This sort of does it. You can basically give a camera reference, a camera movement as your reference input. and then it mimics that exact camera movement in your ai output on top of your starting frame so we see one example of this sort of like barrel roll camera reference and then it does the same exact barrel roll on a variety of different images okay i already have a crazy workflow that i'm thinking in my head so you you build a beautiful set uh with human actors and instead of bringing your aria alexa you just bring in your canon 5d like a dslr and you take really amazing reference photography and then you take an iphone and you do your cinematography with it you're never renting the aria alexa right because that's super expensive and then you use this to get your aria alexa level quality out of your iphone footage wait so where does your aria alexa come in again if you were to do a really cinematic really high quality shot let's say two people sitting in a bar talking you set decorate the bar you bring the two people in with costumes you light it and the only thing that you're lacking is a really high-end cinema camera like a red v-raptor ari alexa sony venice those things are generally like i don't know twenty thousand dollars a week to rent plus additional crew to man them and it's cumbersome so what you do is you get a high-end still camera like a dslr you do reference photography in all types so you retain a lot of pixel quality dynamic range quality and then you take your iphone and you're just doing the cinematography with your iphone right and you bring that into this model theoretically you should be able to convert your camera language into the quality of the dslrs without the need for a sony venice my response to that is if i was already going through the entire hassle of getting a location and casting and lighting and bringing a crew there and my one hiccup was like i couldn't afford an aria alexa uh i still think we'd get really good results with like the black magic ursa or even a pixis or any other you won this round joey you won this round

21:38hey we're all here we're all here everyone hold up let me just take a photo right here i totally forgot how cheap okay all right we're good cheap professional cameras have gotten and how good they are i mean also that goes to the argument where it's like you know really good lighting can can make a you know whatever quote mediocre camera look pretty good uh with some decent lighting i've been thinking more for this like in the um kind of light craft jet set world where it's like, what if we're filming the movements, but we're just in a very basic bare-bones studio with very minimal props and sets and stuff, but we want to have that camera motion.

22:13And so we record that, and we use that as the driving for something that's a bit more of an AI workflow. Okay. And we save all the money on Lightning and the rest of the crew for that scene. I mean, if things go the AI success story way, hey, you're talking about elimination of most costuming, most gaffing, lighting, grip. I mean, so much of it you won't need. Just people in a... First, I'm saying, yeah, I think it's going to end up being both. Like, you know, I don't think there's going to be some, you know, just AI takes over everything. I think it comes down to, you know, what do you need for this?

22:53What do you need for that? Filming in a smaller location, a lot of, you know, there's already been invisible VFX for decades. like yeah that kind of stuff true and in some cases maybe you know maybe for depending on the project yeah there are synthetic characters driven by humans human performances but i think it's gonna be a whole gamut i don't think it's gonna be like every like gaffers and lighting crew and everything just no i don't i don't think that as well i just think that you know in in some cases you would need 20 people to light a set now that number comes down to five people and like half the amount of lights or whatever just because you can cover a lot of relighting needs in post that's true i would agree with that yeah i mean also i don't even know if that is entirely on ai or just like light light technology and um a lot of the tech behind that has gotten better and you can do more with a smaller crew or there's just not even a budget There's not enough, but yeah, that's probably the biggest, um, biggest motivator, biggest enabler for sure.

23:58Other paper motion stream. And so this was pretty cool. It is real time generation click and drag. And the motion inside the video is being modified in real time through the mouse movement. Play the video. Yeah. Can you see it? Basically the mouse is, let me try to get better. Let me see if it's a better example. Oh, you're doing this live. Oh, it's a video. No, it's a video playing. But the video is doing it live. The video is clicking on this thing, and now it is dragging this dog head around with the mouse, and the dog head is moving. Dang, that's like a control rig from animation. That's insane.

24:36Yeah, and it's happening. Yeah. And it's near real time or absolutely real time? What are we talking? 0.4 second latency. Near real time. 0.4 seconds at 29. 29 frames per second, 0.4 second latency. I would say that's real time. Yeah, look at this thing. This elephant's moving based on the direction of the mouse in real time. Holy shit. This is crazy. Our model runs casually in real time on a single NVIDIA H100. All video results shown here are raw screen captured without any post-processing. Yeah, this is a character control break. Between these two papers, the camera control and this, this is where I keep seeing the AI video generation going.

25:15More control, you have the camera control, you have real-time control. When we click and drag and move things around, that's where I see this going. For sure. I think if you give this technology raw in its raw form to an animation studio like a Pixar, right, the character riggers and the people that build control rigs for characters will far better build the tooling around it, the user interface around it, to make it usable than the people that write the white paper. Because for them, it's not, you know, they're just getting this idea out to the world. Yeah, they're doing research. But I think the tooling is where it's going to make a big difference.

25:57Yeah, this is like the seed of just like, oh, look, we could do this. How do we integrate this technology into our existing pipelines and bring artists on board to make this work? You're right. Yeah, between this and the camera control tool, I agree with you that we just saw. Like, this is the AI 2.0 tool sets that we're going to see over the next couple of years. Like, what happens after generation? How do you get? Yeah, get rid of the text prompting. How do we, like, how do we interact with these tools in real time, sort of how we normally would with a real camera? Or, you know, if you're doing miniatures or models, like, and I'm moving this cup around and it moves in real time.

26:38How does that, how do we bring that to the computer? All right, a couple other quick updates. This one is maybe a little niche, but I was in a specific bind and looking for something. And it sounds kind of technical, but it's cool what they built. So Gemini, they said they added the file search tool inside the Gemini API. What does this actually mean in practice? Basically, if you use any of the AI chat tools and you want to give it a bunch of information as a knowledge base, the AI chat, they have the context window. And so pretty soon it's going to fill out that context window and you're going to max it out.

27:16So like if you're working on a film and you want to give it the script or the script maybe not – it's not so long. But in our case, we're working on documentaries and we want to give it like transcripts from all the interviews so it knows that. But if you give it the whole stack of transcripts every time, it's going to max out the context window and you're going to like hit your limits super fast. And so the solution to that was using a RAG. Retrieval augmentation model? Retrieval augmentation. Yeah, it's a database retrieval implementation. Yeah, you store all your files in this database, and it sort of indexes them, and then your AI chatbot can go and find the information that it's looking for if you ask it, hey, where does someone talk about this line and this transcript?

28:00It's kind of clunky to set up if you're not a technical person. You have to kind of connect a couple tools together. Gemini now just announced a new product, but basically it's still an API, so it's still a little bit technical, but basically they merge the RAG and the Gemini chatbot into one tool. So you can just give it all the files and then ask Gemini about your files, and you can give it thousands and thousands of files, and it'll be able to index. It automatically does the indexing. It automatically does all of that stuff that used to take a lot more technical know-how. It just does it automatically.

28:34And so this is useful. I mean, it's definitely useful for for our documentary projects, but it makes me think of just having that sort of central knowledge base of any project you're working on and be able to have a chat interface to talk to it. Sort of what Bryn talked about when we interviewed him, where he sees each film having its own sort of custom AI model or kind of custom knowledge base tied to it. This is a step in that direction to make that easier for any small team to set up themselves and have a central chatbot database about anything about their project. For sure. I think you're absolutely right.

Read the full transcript

29:10These are like the quality of life improvements that make or break a production, right? And this is going to go under the radar, you know, sort of like, yeah, we want to see image and video, but no, this is the stuff that actually makes the difference in a massive production that would scale across multiple teams, multiple shows, and you name it. Yeah. what version is this file on this thing where is that file what is this quote on uh yeah so i think that's a good quality life improvement okay by dance yeah this one's another quick update um but by dance now has their own video upscaler okay it's on fall it can upscale videos to hd 2k and 4k 30 60 frames per second the pricing i mean it's 0.0072 cents a second pretty cheap 7 tenth of a penny so yeah this is you know I mean I love Topaz use Topaz a lot for everything but it's good to have a couple options there too because like all the AI tools depending on what you're trying to upscale some of the generative upscalers are good at some types of shots and some are better at other types of shots absolutely and I will welcome 7 tenth of a penny price thank you by dance yeah is it subsidized by China?

30:24yeah good update there do you think that's why it's so cheap? well let me see it does say that was for the 1080 upscaling I wonder what the 4k upscaling is it's probably a slightly different pricing here. If we do 4K, 60 frames per second. I ran out of my foul credits in time to re-up and try all these new things. It is. If you go to 4K, it's 2 cents, almost 3 cents a second. So a little bit higher. And then if you go 60 frames per second, that doubles the cost. So if you're going 60 frames, 4K upscale, that's about 6 cents a second. Yeah, the examples are always generated video upscaled. Why can't they just take real footage and upscale it?

30:59it would look so much better. Yeah. With like weird color space issues and lighting issues. Yeah. It's having, he's having a bad hair day. Yeah. That's why I'm so angry. It's like, my hair's a mess. Yes. Joey, you had to do some dad jokes on this episode.

31:20And then speaking of fall one, 2.2 animate has an up upgrade, a four times faster inference, It's sharper, cleaner visuals,$0.08 a second at$7.20. And this is a good way to use WAN 2.2 Animate because also as we were DMing before, but there's no, it's not, as of now, it's not really built into many tools. It's not built into FreePick and stuff to use. So if you don't have ComfyCloud or kind of want to get into the whole comfy thing, you could just mess around with Fall and just pay for your usage. And that's a good way to dabble with WAN 2.2 Animate. It's a solid tool. Yeah. It's super slept on, underrated.

32:00I think it would solve most people's animation needs, honestly, if used correctly. Yeah, the performance in it has been great. Yeah. I told you that I didn't... If you use the 1.2.2 Animate workflow in Comfy, the default option for whatever weird reason is not to take an image and drive the image performance with the video, which is the common use case. it is to maintain the person in the video and then sort of blend them with the input image. So I did it with a test performance of myself. I guess if you wanted to kind of just put yourself, animate yourself into a new environment, I don't know.

32:45I think there could be use cases there. I got to mess with it. It wasn't what I was expecting. But yeah, I gave it a test performance video of myself and then it turned me into like a half looking elf because I was trying to drive an elf video performance. And then it like was video me with like elf ears in like a Santa's workshop. I'll upscale that. I'll actually, I'm going to do some post-processing on that video once you send it to me. I mean, you know, actually like what we're literally talking about, it could be for this use case that we're seeing right here where the woman wants to maintain her look and she's just re, she's changing her environment, but her, she's not changing herself.

33:22She's changing her costume. She's changing her environment. So I think it would be more for that kind of use cases. Whereas if she wanted to use her motion to drive a completely different character, that's where you would not want anything from the original video. So in this example, how do you determine the woman's head doesn't change, but the body changes? Is there controls for that? In the comfy workflow, there is an extra step where you mark off what you want to keep and what you want to delete. And it does a very rough mask. I don't know how that works on Thal. Yeah, I would imagine it's all API driven, so it's probably in there.

33:57Okay. Oh, and I was going to ask you, so if you're running Comfee locally, you could bring in a file node and have that node point to 1.2.2? How would you do it? There's not an easy way. I use Cloud Code to make a node. You're advanced, Joey. Maybe we'll do that video about that, because that was one of my cloud code use cases was, yeah, I wanted to turn... Oh, I did a demo of it at the Comfy UI meetup in LA. That was where I talked about this. That was where I talked about this. But that was my dream case, was I wanted to bring in fall API nodes into Comfy. So there's a couple of reasons why you don't want to do that.

34:42But yeah, one is if you want to use models that aren't available with cloud. Yeah, you're always at the latest and greatest model, which are generally not available for a while onto like a free pick or whatnot. And then also fall is quite cheap, quite efficient, you know, and like a dollar goes a long way. Yeah, false pricing is great. Comfy UI doesn't have every single model as a native API node. And then also if you're doing client work and you want to just maintain separate billing, there's no, in Comfy, everything's built under your account. There's no way to segment how you build things. So there's a couple reasons why you would want to bring FAL API nodes into Comfy.

35:21But there's no easy way to do it. But I have a cloud code where I just give it the URL, and I'm like, turn this into a Comfy node. You're up against the clock here, because Comfy cloud is quite good at 1.2.2 animate. It's just the bugs need to be worked out. Yeah, I mean, that's one specific thing. But there's tons of models on FAL that are just not on Comfy. If you wanted to really cover the whole gamut of what is out there and try out a bunch of models, fall and replicate. Yeah, like, for example, the ByteDance upscaler that we just talked about. That's probably not going to be anywhere else.

35:53Exactly right. I think the only upscaler on Comfy is Topaz. All right, Joey. So another research paper, another day, another paper. This time we have Infinity, Infinity Star, quite the name. And guess what? The description is going to have you talk about Interstellar because it is a unified space-time autoregressive framework for high-resolution image and dynamic video synthesis. I feel like there's no way to talk about that without Hans Zimmer. So what it is, and so I could be wrong here, and viewers, if you know better than me on this paper, please comment. it. My interpretation is that generally for video model generation, you have latent space that is separated into two spaces.

36:43So you have your temporal latent space, which has the notion of motion. And then you have the spatial latent space, which is more consistent with image generation. It has an awareness of what the beach, what the palm tree, the rocks look like and then it'll build the aspects of the image and then the temporal latent space knows motion knows camera movements and some of the time-based things so it'll then take that image and build the video out of it so what this infinity star paper seems like is it mashes those two things together into a single latent space so that you have the awareness of the temporal aspects as well as the spatial aspects all in the same part and that gives you robust generation abilities when it comes to extended video right like beyond five seconds beyond 10 seconds or if you want to generate an image that has time like applications to it so if you're doing motion blur for example on a still image things like that what are your thoughts on this So you're saying it's kind of like instead of a video generation, we're trying to kind of predict what the next few frames should be that makes sense.

38:00It also has an understanding of like what the world is like to kind of make it more consistent and basically like follow the – have a more consistent feel. It says it's a space-time unifier. So I think we tend to separate those two things out mathematically in the AI neural network world. where this is just treating all of it as the same type of thing. So time is just another asset. You know, any time-dependent things like motion is just another object in the latent space, if that makes sense. And look at the demo clips, and it's like the demo clips don't really... Yeah, it's a little bit cryptic for sure.

38:42And, you know, again, at this point, it's just basically a white paper. No, there is a demo because they have a Discord where you can prompt and generate some stuff. But, yeah, I would say there's nothing in the outputs that makes it look like, you know, it's not like I'm thinking of ByteDance's, you know, other model like C-Dream where you can create cuts in the output and it's like in the same latent space generation. So you have that consistency. So, yeah, I think it's I think combo. It is a paper that has some demos, but it's not really showcasing what. OK, well, TBD and I'm sure we're going to do another retake on this when we get more information.

39:18Yeah. Yeah. And we see some. Yeah. interesting outputs. I think it also might be an open model as well where you could download the weights and run it yourself, but it might be incorrect about that. But yeah, I mean, I think it's cool to just kind of see how different models are rethinking the way they generate. Well, a lot of good progress today, man. Yeah, you know, for a week we've got a lot to catch up on. All right, yeah, links for everything we talked about at denoisepodcast.com or over down in the show notes over on YouTube or wherever you're listening to this on Apple podcast or Spotify.

39:52We just got our 21st review on Spotify podcast. So thank you, whoever you are. We appreciate it. And also on YouTube, we have a new feature called Hype Points, which we're going to show you in this recording here. So if you are on the app, you scroll right underneath our video, there is a new button called Hype. Give us some Hype Points. Let's see how it goes. Sure. I still want to know what the high points look like. So yeah, let's see that. All right, thanks everyone. We'll catch you in the next episode.

From the publisher

This week, Addy and Joey dive into the Nano Banana 2 leaks and their potential implications for media authenticity. Then we explore Comfy Cloud's powerful browser-based workflow ($20/month for 8 GPU hours daily), Veo 3.1's new camera controls, and real-time AI video manipulation that's bringing unprecedented control to creators. Plus: Figma acquires Weavy, ByteDance launches affordable video upscaling, and practical AI tools that are transforming post-production workflows.

--

The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.

More from Denoised

All 101 episodes
Comfy Cloud, AI Camera Control, Nano Banana 2 and more AI updatesDenoised · 40 min
Listen in VO