“Nano Banana for Video”: The Simplest Way to Understand Gemini Omni

28 May 2026 · 40 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode is a post–Google I/O recap focused on Google Gemini Omni, “Nano Banana for Video,” and how it differs from prior video models.

Guests

Addy and Joey (podcast hosts). Addy recently attended Google I/O and discusses Omni demos; Joey compares Omni to other tools (e.g., Runway Aleph 2, Sora, Clay/SeaDance) and shares hands-on tests.

Key claims

Omni is a “world model” for video-to-video modification (not diffusion like VO4), with strong physics/world understanding and best-in-class editing. The biggest quality gain is the avatar feature: a likeness-calibration process (90 seconds/1–2 minutes) tied to the account, producing better face motion than Sora, though not 100% likeness.

Notable examples

turning a Venice dog into a robot dog; restyling a LACMA “Metropolis” city while keeping geography/camera; making a Google building launch like a spaceship; “vintage objects” spelling “DENOISED” with correct letter placement; and video-to-video robot/chimp swaps. They also cover Flow (agentic batch video edits, character system) and Google Pics (Canva-like layered editing).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring Google's New Omni Model

0:46 to 1:59

Discussion of Google's Omni model's unique capabilities and comparisons.

“Yeah, the Husky shirts or whatever, the cheapest ones you could mass print.”

Avatar Feature Insights

2:00 to 4:16

In-depth exploration of the avatar feature and its performance.

“And basically, VO is more of a diffusion model that was just really good.”

Video Modification and Creativity

4:17 to 7:01

Hosts analyze the video modification capabilities and user prompts.

“as well as the Miami Vice gray T-shirt shot.”

Gemini Model Announcements

7:02 to 12:39

Discussion about the latest Gemini models and their implications for AI.

“Although Omni does have really good reasoning in the same way that Sora does.”

AI in Science and Future Prospects

12:40 to 14:02

Exploration of AI's role in scientific advancement and future discussions.

“Yeah, no, you said you saw Sundar, the CEO, as well as Demis Hassabis, speak on stage.”

AI's Impact on Science and Health

14:02 to 16:49

Explore the potential benefits of AI in scientific research and health improvements.

“you talk and frame uh when around ai and the risks it brings, but also the benefits that it could bring, which is also why they kind of did do a big push on like AI for science.”

The Future of AGI and ASI

16:50 to 19:24

Discuss the timelines and implications of Artificial General Intelligence and Artificial Superintelligence.

“The side of non-AI jobs, there's an oversupply, And so you're fighting and wading through like thousands of applicants and so on.”

Waymo's Training and Real-World Scenarios

19:25 to 24:16

Analyze the training processes of Waymo vehicles and their responses to unexpected scenarios.

“The same way you and I, like, hey, do you remember a time, Joey, before cameras and before editing and before color science?”

Google Flow and Video Creation Innovations

24:17 to 28:04

Learn about the features of Google Flow and how it's transforming video creation.

“which sometimes is good and sometimes makes figuring out which Google product to use for your problem.”

Exploring Agentic Creativity Tools

28:04 to 29:16

Learn how AI tools are enhancing the creative process through iterative workflows.

“It's probably not, but it's more like assisting you in the creativity process.”
Show all 14 chapters

The Rise of Vibe Coding

29:16 to 31:23

Discover the concept of vibe coding and its implications for app development.

“Yeah, if you're just like, oh, we copy and paste this prompt and redo it.”

Innovations with Google Pics

31:23 to 33:54

Learn about Google Pics and its capabilities for editing generated images.

“But this one, I think, was the most impressive or coolest surprise thing that I saw.”

Runway Aleph 2: Video Editing Innovations

33:54 to 36:50

Explore the features and improvements of Runway Aleph 2 in video editing.

“Separate topic, but I mean, how are they going to make their money?”

Comparing AI Video Models

36:50 to 39:43

Compare the performance of Runway Aleph 2 and Google's Omni model in video generation.

“Yeah, it's so much softer and lacking so much detail than what we saw with Omni.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Joey:I think the best way to think about it and the way that they were explaining it at the event was like, this is not VO4 and it's a completely different model. This is the video version of Nano Banana Pro.

0:14Joey:All right, welcome back to Denoised. Addy, good to see you. Nice to see you, Joey. Welcome back to town. I am back in town. I only just spent the last few days at Google I.O., which was a lot of fun. I am decked out in swag. I got a Gemini hoodie and Vibe blank Vibe code hat. Maybe I'm just getting old, but when did tech gear, free stuff, start to look so good? Because remember, you used to go to CES like 15 years ago, you'd get these free t-shirts, you're like, oh, I can't wear that. Yeah, the Husky shirts or whatever, the cheapest ones you could mass print. I didn't give out. No, that stuff's very nice.

0:54Joey:Okay, so yeah, there's a lot of Google updates, obviously, and a lot of I.O. updates, or updates from I.O., so let's talk about it. Yeah. Biggest one, Google Omni, their new video world model. Yeah, big update. I know we get video models that are coming out every couple of weeks, and Google also drops, they're also in the fall, but what I noticed, I mean, please go through the Omni specs with me, and then I'll tell you what some of the things that I thought was really amazing. I've been sitting on this for about a week of like, where does this fit in or what is this good at? My initial, when I first had access and messed around with it, obviously my first jump to like, okay, how does this compare to claying and sea dance?

1:39Joey:And on the surface, I would say for like, if I'm looking to do cinematic visuals, it still is not at the level of four. And it's a completely different model. This is the video version of Nano Banana Pro. It's a world model. It has world understanding. It understands physics. And basically, VO is more of a diffusion model that was just really good. And so this is completely different. This is like V1 of this new Omni model. And the thing that it is really good at is editing videos and modifying videos. I feel like we're going to get stuck with calling things editing videos. I still think editing videos is like snip snip.

2:19Joey:You're cutting videos together. but that lingo has yeah i call them video to video video to video or just video video modification you give it an input video and you change something in painting yeah in painting video um really good at that really good at just giving it world understanding prompts and really good or pretty good um avatar feature which we can talk about in a second i thought the the like the biggest sort of um you know the biggest gain in quality was definitely building a human from the likeness of like, I don't know, like a one minute or 90 second calibration, very similar to what you did with the Sora app.

2:59Joey:Yeah. So the, okay. So yeah, the avatar feature, let me say if you give it an input image and I tested it with like a couple of like photos of me or my wife and like image to video. And it has, if you go in the Gemini app, there are a lot of like templates set up of just like fun things you could do with your stuff. Yeah. It really, the likeness fell apart. And that was my first impression when I kind of gave it some images of me from like a trip and it tried to make a video. It completely changed our faces. And my initial impression was like, what the hell is this model? What's the fuzzle? Yeah.

3:32Joey:And then actually at the event itself, I was speaking with Justine Moore from A16Z. And she was like, oh, you got to try the avatar feature. And I was like, what avatar feature? and it's like buried in the UI is this avatar feature that it's assigned to your account. You basically use your phone. It's very similar to Sora. It has you like turn your head left, turn your head right, and then you say a bunch of numbers. Yeah. And that is your avatar. It's tied to your account. And then you can make videos of yourself with your avatar. And the quality there is way better. Oh, so good. I thought this is way better than Sora's avatar was.

4:13Joey:It's way better than stories, Avatar. And even, like, I think you used Avatar for the champagne shot here, as well as the Miami Vice gray T-shirt shot. Like, to me, that's 98%, Joey. Like, I've seen you in person enough to know, like, maybe you're not. If you see this, it's like. You're not this much of a stud. No offense. But, like, it's pretty fucking close. I like Omni that it definitely, you know, kept making me more jacked than I am. Yeah, it just gave you a sharper jawline and just fuller hair. It just amplifies. The common denominator is a really handsome man or a woman, and it just kind of pushes your avatar into that direction.

4:53It doesn't retain 100 % of your likeness. But I think that's a creative choice on Google's part. I think if they really wanted to, they can dial that aesthetic thing down and just keep it true to this case.

5:05Joey:Yeah, there's probably a big balance of just like, what is it actually capable of and how much are they letting people do. There were times where I tried to test out prompts and stuff and it was just like, either it was overloaded or couldn't do it or wouldn't do it. And sometimes it was like, what's the reason behind that? The other really impressive thing about the avatar feature, there you go. Yeah. You're pulling it up right now, is the vocals. Like the AI voices miss a lot of the inflections and some of the, like the, hey, you know, and then know who like the range is very limited but here i feel like they're they're getting a little bit better with that range it's still not 100 as natural as how we sound every day but it's getting there and it's i think omni really made kind of like the issue with the avatar and so you can see like this was the source image like i trained this at night in my hotel room on my phone so from the visual quality uh really good of what i was able to pull out from my phone the audio of my voice, let me try to phone microphone, because it was.

6:07Joey:And there's no like, like no Adobe podcast audio. There's no like audio enhancement that this model is doing. Google just dropped Gemini Omni. It's nano banana for video. Yeah, that doesn't sound like you at all. It's funny because some of the other examples that I've heard is pretty spot on to the people that really, maybe like, you know, if I plugged in a DJI mic and trained it with a better microphone, then maybe it would give me better outputs i don't know if it's the mic does it make you just say the same thing for no it's in the prompt it's whatever oh oh oh you mean like the calibration process the calibration is literally you're saying numbers okay that's it okay yeah but i'm just you know i'm just holding up like so you don't say like mary's sales she see no you're literally the same numbers yeah okay but you're it's whatever you know your mic is here so yeah if i had a better mic plugged in that maybe I would get the outputs of this.

7:01Yeah, but I was quite impressed by the avatar feature more than anything else on Omni. Although Omni does have really good reasoning in the same way that Sora does. Whereas if you give it just an idea, it'll expand on that idea and really go to town on creating the entire editorial. Well, yeah, let me show you some of the outputs.

7:21Joey:First up, an honor view. Oh, yes. 1970s New York City walking down the street. Although it is technically, I mean, these cuts are incredible. Yeah, and this is a very basic prompt. I didn't give it any info. It just kind of reminds me of Saturday Night Live, like the Travolta scene where he gets the pizza. Two slices, two, two. And then he does the walk. World understanding stuff. This was a video-to-video test. And so the original video was just like my wife walking the dog. on venice so it's like okay dog's danny panning around and then i said okay turn the dog into a robot and like the dog turned it to the robot the movement looks good everything else yeah no no distort the only thing i saw was her left hand and the leash was a little funky when she grabs it again right there but i'm being super picky obviously the the sheen on the and the finish on the metal seems very plasticky and not as reflective and responsive to the environment.

8:31Joey:Yeah, I think those robots are matte plasticky. It even has the same kind of movement of it, but it also kept his arms on it. I'm being super picky. This is incredible. Like robot dogs here. Yeah. Wait a minute. This is the best. Well, because also, once we get to our later updates of Aleph 2, I will reprocess the same shot. Looking forward to that. I haven't tested. I assume you could give it, I could give it an input image, like a restyled image to give it more direction and be like, change the dog to this. This is my prompt. It's literally right. Change the dog. It goes and picks the robot that it wants.

9:03Most basic prompts possible.

9:05Joey:Yeah. I picked the robot, but also like, you know, literally understood the dog, swap the dog out, everything else, like the transparency of the background. But the fact that it changed it to a robot dog and not like a robot, the size of a dog that's a biped that's reasoning like it's going through some sort of intelligence yeah it kept the same right it kept the same dog thing yeah i mean i also did this one with same same source video i said change him to a chimpanzee that looks better see like when you don't do metals and reflections to me that feels more real these look really good out of the box i'm just kind of these video to be once it was like okay don't look at it as like a video generation model but a video modification model.

9:47Joey:This one, this was a shot from one of my favorite installations at LACMA. I forgot the name of it. The Metropolis City with the bunch of cars. Really? It's like a permanent installation. It's like a massive permanent installation. Those cars are running full speed. Yeah, it's like a bunch of like Matchbox cars in this like crazy futuristic city. And I said restyle it to a futuristic city but keep the same geography. Oh yeah, they kept that shot right. Yeah. And kept the cars, kept the stuff. A little bit of warping there, but like this was one of the best outputs where it still kept the - The shape of the tracks.

10:22Joey:Structure. The camera movement, and it just changed everything else. One more, this one was actually fun. This one I just gave it an image of one of the Google buildings, cause they had a bunch of these and they just looked really awesome. And then I said, make the building take off like a spaceship. Yes. Yeah, it did it. It's like, it kept the building. It was like, it added this extra shot I wouldn't have done, but like it kept the building and had it shake and bust out of the ground and have rockets underneath crazy um okay the world understanding this i said create single shots of vintage objects the first letter of each object spells out the word denoised the letter should also be on the object in a natural way so basically i wanted the first letter if it was like you know whatever like I said vintage objects, but it was like dog.

11:10Joey:Yeah, it's a good reasoning test. E, N, O, I, S, E, D. Didn't get the last one. Don't know what the last one was. But it got basically dial phone. Yeah, the filament on the Edison. Because I asked him to explain. I love that. Yeah. Yeah, I said explain what the list was. and it also the list objects here are slightly different than what I put in the video so the understanding between the two is a little mismatched but and newspaper oh oil can the ink thing with old-timey pens yeah ink holder i yeah um s stopwatch e i think that was edison fan eyeglasses okay uh all right well look i don't know and then the key i don't know did you generate this in one shot off the rails but that's crazy this is literally this one the fact that it is able to generate so many different things and then cut it together into one clip yeah that's impressive that's where the model part comes in of like it knows objects letters uh you know much sharper text rendering it just it just trained on the internet man like it just knows every single object known to man through the through the years through the decades yeah yeah like it'll know all types of bulbs not just this bulb right it'll know fluorescent bulb led bulb and so on so omni probably biggest announcement super cool they also had gemini flash 3.5 the names always confuse me tell me nano banana yeah names are confusing and it's still basically they're like it's not the pro model but it's beaten benchmarks for them that their 3.1 pro has and that they're like kind of pushing it as the new video model for now until there's like a pro model oh no no this is just gemini like their text model for okay yeah the lom correct me if i'm wrong but the official name for nano banana is also like gemini something right it's like gemini image three point something Yes, yes.

13:20Yeah, no, you said you saw Sundar, the CEO, as well as Demis Hassabis, speak on stage. What would they say? What happened?

13:28Joey:Oh, yeah. I mean, so, yeah, they had the keynote. They did that. And then there was like a kind of fireside chat side things that you could go watch. And so, yeah, I know for Demis, it was AGI coming in. Oh, wow. Three years. They pushed that timeline way back. yeah that was his time yeah well okay yeah and there was definitely a bigger push you know his was interesting of like it was a you know debate of like the ai do me doomers versus like optimism realistically optimistic or something um but basically it kind of came down to like how you talk and frame uh when around ai and the risks it brings, but also the benefits that it could bring, which is also why they kind of did do a big push on like AI for science.

14:20Joey:And I think it's a whole separate, um, division now or focus on Google. Yeah. I'm glad they're doing that. Cause I don't think, um, open AI and some of the other competitors are probably not investing as heavily in science. No, I'm not talking about it enough. And that was like his original, uh, you know, a lot of the original team might stuff with alpha fold, you know, mapping proteins and stuff. So like, you know, that's something it's like if you can get rid of diseases or cure diseases, like no one. Exactly. Yeah. And if you can cure cancer, then you can go make a little bit more slop and it kind of just evens out.

14:56So they understand it better than we do. And yeah, I think they're like initially, if you remember a couple of years back when like the first chat GPTs were coming out, they're like, yeah, at this point, we're going to be able to like figure out nuclear fusion and cancer research and all the stuff that needs heavy, heavy supercomputers. They're like, yeah, we can do it now. And then it kind of just went away. And then we had a bunch of mean generation capability and, you know, a lot of job displacement, unfortunately, like a lot of companies bet on AI and had to fund the capital expenditure. So they It was just all been negative since then.

15:39So I'm hoping like with Demis' sort of push back into goodness for humanity, maybe, maybe, just maybe we can solve something huge here that will really benefit mankind. Yeah.

15:50Joey:I mean, yeah, I think there's like a big just, what is the benefit to me thing with AI and all this stuff? And it's like kind of been hard to see, especially with like a lot of displacement, job loss, all of that stuff. And yeah, I think if you could just be like, oh, well, you can improve human health and well-being and cure terrible diseases. That is a clear benefit to society that is hard to dispute. So about 12 months ago, when we started the podcast, they were saying that, hey, AGI is going to be here very soon, but ASI will take a while. So general intelligence is what I think is defined by an AI system that is as good as one human being that can learn as well as we can, new skill sets and so on.

16:32ASI, artificial superintelligence, it's like a bunch of human intelligence put together into one system. So it is super intelligent than us. And if they're saying AGI is now three years out, I'm guessing it's more like five to six years out. And ASI is probably decades away. It's probably a good thing. It's probably a good thing. We don't need this stuff right away. There is enough disruption as it is. as you saw just a couple of days ago meta had i think the biggest layoff yet right 10 of the workforce something like that and they were also like everyone that was staying yeah like there's leaked audio from the zuck that yeah we're training on everybody because you're you're you guys are really smart yet we don't need you here i haven't dug into the full post but also the ceo of click up they did a layoff and i the i didn't read the whole thing yet but like the gist was

17:27Joey:like you know people that stay you know ai should make them 10 times 100 times more effective but like i'm also looking at compensation tiers that are like a million dollars you know if you are the person that like leverages ai and then could do like 10x 100x more and then get compensated appropriately for that. Yeah, there's a really interesting divide in the talent pool, like in the job market now, and I'm kind of starting to notice it a little bit more, and it's becoming more and more pronounced, is that there is an oversupply of talent on the, like there's a hard fence in the industry on AI jobs and non-AI jobs.

18:09The side of non-AI jobs, there's an oversupply, And so you're fighting and wading through like thousands of applicants and so on. On the AI side, there's not enough people. Like they literally can't find people. So they have to overpay and get those million dollar contracts out to hire whoever.

18:25Joey:And the other about AGI, I think he was asked, and I'm going to obviously paraphrase and try to remember the best I can, but he was asked, like, how would you even know, like what, you know, if it does achieve that? And Demis' one of his tests was like, if you had the AI model and you just gave it world knowledge up to like early 1900s, like 1910 or something, would it be able to figure out, you know, right? Right, exactly. Like things that Einstein discovered and other things that humans discovered in that time for, you know, after that time period. That was like one of his benchmarks for like.

19:03Yeah, it all goes back to what Yanlakun is doing now, right? Which is that when we train a model, during the training, it absorbs and learns everything. But then after training, the clay has hardened, right? So you can't retrain it unless you train a new ChatGPT version or so on. Like it has to be an entirely different model. So how do you build in mechanisms for a model to keep learning obsessively over and over? The same way you and I, like, hey, do you remember a time, Joey, before cameras and before editing and before color science? Like, you learned all this stuff, man. So, like, how do you make an AI do that?

19:42And that's what Jan Lakuna is working on is he's abstracting away the notion of tokens, essentially. So instead of like data captioning and images and videos being the tokens, how do you have a more abstract form of neural network that just relies on vectors? And like there is I'm forgetting here, but essentially he is making a more abstract version of a neural network where it's multimodal by default. And it's much more than that. it is also able to absorb inputs much more easily into the network. Yeah, what I could do is I'll do a little bit more research and then you can ask me about it on one of the episodes we shoot later.

20:28Yeah, I don't use notebook LM. It just hasn't served me high utility yet. You will when you're researching this. Just watch Jan LeCun's. Okay, the best thing to do is to watch some of his YouTube stuff and go to the Gemini summary.

20:43Joey:Yeah, that's also helpful. Yeah. Last one about the dumbest thing that I thought was interesting and that also has turned out to be really ironic in the last 24 hours is he talked about how they were using Genie, the generative world model that we talked about where you can spin up a world and navigate through it. But they're using that to create worlds to test the Waymos in with super fringe case studies to see how they would react. One example was if the Waymo's driving in a forest and a forest fire breaks out and it's surrounded by flames, what would it do? How would it behave? And he was talking about these one in a billion fringe scenarios that's not going to usually happen on regular training data in the real world.

21:38Joey:I was like, okay, interesting. Fast forward to yesterday. Have you seen this? The Waymos in Atlanta have been going full send into flooded streets where they have now stopped Waymos. Okay, for a second I thought that was an AI-generated image. Now, that's terrifying if it's real. No, this is real. The Waymos have been driving into flooded streets in Atlanta. So they need to spin up some more Genie models to test out the fringe cases. Yeah. It happened, and they drove through it. Sometimes I feel bad for the Waymos, but then they're machines. Who cares? I did take my first highway Waymo at the event.

22:22Joey:I texted you about my first Waymo fight. So I was in Santa Monica, Joey's neighborhood. And if you don't know, Santa Monica is like the unofficial capital of Waymo in L.A. And there was a Waymo next to me. And I was like, what happens if I mess with it? And just so you know, this was for educational purposes only. I don't recommend that you actually mess with Waymo's disclaimer. Okay. so i took my car and i just lightly started to go into its lane and at first it kind of just slowed down and then i just like huh it's pretty polite and then it caught up and then this time i was a little bit more aggressive and i was just like like just did a real quick jerky move and honked at me and i was offended i was like hey don't honk at me see the waymo honk but i totally deserved that honk but i did want to check how responsive their uh driving system was and it was quite responsive someone brought this up but i felt the same way where like if i like crossing the street i feel way more safe like or don't think as much if i'm jumping in front of a waymo absolutely the waymo will most likely stop whereas a human driver might be on the phone or whatever yeah yeah distracted yes the waymo's got a bunch of sensors i was at a group yeah i guess i should also probably disclose like the trip was paid for by google uh and it was you know a lot of it was like the builders groups there's also a lot of kind of content creators and people that are messing around with the Google products.

23:50Joey:But I will say they were genuinely interested in getting user feedback and making these products better. And they are aware of what happens online and what people talk about. So they're always looking, even if it's just bug fix or things, but also just what use cases there are and how can the models get better. And also explaining why they throw a bunch of stuff under Google Labs and then kill products sometimes because they're like, yeah, sometimes the products suck or they just don't take off and just want to see what works. which sometimes is good and sometimes makes figuring out which Google product to use for your problem.

24:23Yeah. I mean, end of the day, they have to run a business and the product has to have some type of revenue promise, right? So they can't just keep making R &D things happen all the time.

24:32Joey:Yeah. And then also try to figure out new things and new tools and uses. So yeah, other two quick things I want to touch on that were interesting and relevant to our audience. One is Flow. So Google Flow, which is their web app that you can use for video creation. It's like kind of their narrow version of like free pick for like video creation. A couple updates there. One was obviously the new Omni models built into it. They added a character system where you can create characters and then call them up. I was pegging them with questions of like, is there something in the model? Or when they eventually roll out the API, because there isn't an API yet, will there be something that like treats characters differently as other reference images and basically the answer I got was like no basically what flow is doing is just sort of structuring the data a little bit differently under the hood but there's no special omni model that it's using so it's basically like anyone could kind of they just built a clever interface and have some stuff helping with character consistency under the hood yeah there is no a grand master plan to tie everything underneath with like a master workflow basically there's no special model that flow is using that anyone else wouldn't have access to via the api it's the regular model just with like uh some extra stuff built on top to to direct it the master plan thing though that you mentioned reminded me we never finished our avatar talk the thing with avatar and comparing it to Sora and Sora characters, you can only make one avatar of yourself.

26:11Joey:You can't make avatars of other people. I asked them, well, is there a plan where you could, if you give me permission with Addy, can I pull up Addy's avatar if he gives me permission and make videos like we did with Sora where it was a permission process, and if you're okay, that people could use your avatar in their own creations. and they said, you know, maybe, but it wasn't on the roadmap right now. So I don't really know what, I'm like, if that wasn't on the roadmap, then what are you going to do with avatars if it's just yourself and you can't, I can't share it with anyone. I don't know, I just think it has huge YouTube implications.

26:47Like a lot of the faceless YouTube channels that make a ton of money, if they added a synthetic face, they would make more money. Stuff like that, you know?

Read the full transcript

26:59Joey:yeah you know actually around here that's that's actually not bad that's probably and and with the quality level where it's at like i don't think they really care that it's not cinematic and it's not meant for our world i think if it's good enough for like a talking at youtube video it's more than good enough for them yeah i mean i'm sure they'll prove the quality on it um but But yeah, I mean, the use cases of us of like, We need aging, de-aging, we need costume, hair, makeup. Our needs are up here, man. Nobody's doing that anytime soon. Okay, so back to Flow. Cool thing they added, which obviously seems to be a trend with a lot of these tools, is an agentic workflow.

27:44Joey:So there's sort of like a Gemini sidebar. And so you can just chat with it and be like, hey, generate 10 shots of a man walking in the forest. or if you have a bunch of shots, the example they gave here of like a daytime scene and you could say, okay, change all these shots to nighttime and it'll just batch process and generate these new shots just through this basic chat interaction. I love stuff like that. Gosh, I'd hate to call it like agentic. It's probably not, but it's more like assisting you in the creativity process. It's agentic in the sense that like it is trying to figure out, like if I just ask something vague, Like, I want a, I think I tested it.

28:26Joey:I said I want like 10 different shots of the same person walking in the same field. And I just gave it the text. It spun up and generated an image of the guy, an image of the field, and then started making the shots using those images as inputs for consistency. And I just told it the one thing. So it had that agentic understanding of like, okay, to make something consistent for videos, I need to have consistent inputs. I got to make the inputs first because they don't exist. Also, like a real creative would never just go make that storyboard, right? All those shots and then just go right into production.

29:02Like that's their iterative cycle, right? They'll start there. They'll modify shot three, modify shot seven, keep going. And I think that's still so much faster and gives them so much more options than like every little storyboard from scratch.

29:15Joey:Oh, yeah, exactly. Yeah, if you're just like, oh, we copy and paste this prompt and redo it. Copy and paste this prompt and redo it. It's like, oh, no, if you just want to batch change something in a shot you already have, just tell it and then have it do it. Obviously, it's still charging you credits for every single thing you do. So you will eat through your credits faster. But, you know, I think this sort of agentic workflow where you're talking to stuff and it's doing and setting up things for you is where a lot of these tools keep going. And then the other thing in flow, and I'm curious to see how people use this is, and this was sort of a trend with Google I.O.

29:47Joey:in general, was building more vibe coding and tool building capabilities into more tools. So Flow has this creative tool builder inside Flow that is basically like a very white version of a vibe coding tool set where you just kind of describe what tool you want. And it sort of spits up the applets. Yeah, that's powerful, man. like a motion tracking one. And can you download other people's tools that have been previously created? You can remix. Yeah. There's already like a directory of like publicly created ones that you can copy. This was my idea for iOS. Like Apple should have rolled out something like this where it's like vibe coding for idiots in iOS.

30:35So you can build little apps that only exist on your phone.

30:38Joey:Yeah. Something. Yeah. It spins up a widget really fast. Yeah. Yeah. This is like that idea. but it's like lives on your computer, in flow. I'm just curious because it's like, you know, it's a, would this turn into the gateway drug of like you mess with something here and then you start going into full-on cloud coding. You're spending$3 ,000 on cloud coding. The next day, yes. Yeah, anti-gravity, Google's version of cloud code. Are creatives thinking that way where it's like, I want a tool that does something that solves a specific problem. It's an obvious idea now that we say it, but for them to actually go through it, but I'm sure there's a ton of engineering under the hood that's happening for user-generated app application on the fly.

31:19So that's a first, especially on the creative side of things. Yeah.

31:25Joey:Okay. And then this will be my last one. But this one, I think, was the most impressive or coolest surprise thing that I saw. And it's called Google Pics, which is a confusing name because there is also Google Photos. But Google Pics, it's basically, it's turned Nano Banana into like a Canva-like editor. So you can generate an image, you can generate a flyer, you know, with Nano Banana that has like text and images and design elements. And normally if you did that and then you wanted to change something, you'd have to like redo the whole thing. And then it sort of does this like, again, this is one of those things where she's like, they're doing something under the hood.

32:07Joey:It's not a special model, but I'm not quite sure how they figured out how to in-paint in Nano Banana and blend the edges. So it turned something flat into many, many layers, and then you can individually edit. Yeah, this was not the best recorded demo, but at the booth, basically from scratch, it generated this rooted future flyer thing. And then you could see the mouse hovering over every element, and it turns every element into something I can click. and then I can reprompt to say what I want to change about it. And then it'll just edit that one section and not touch anything else. But the outputs are like, feel cohesive.

32:47Joey:It was kind of wild. This was another test one I did where I had already prompted it. I think I said, I clicked on the text and I said, make the text fancier. And I clicked on the person on the left and I said, put them in like a spaceship in an astronaut outfit and the person on the right in a, um, I don't even know, safari outfit. And then here's the output. And so it changed. The text is the same. It changed the text. It kept the person's likeness, but it changed their outfit. Didn't change anything about the layout. Yeah, it's still there. Even that tape thing on the top, the tape's still there.

33:23Joey:It's still blended in, but you know, it changed the, uh, underlay of the wording underneath it. Oh, dude, graphics design is getting so easy. I mean, like if you compare or if you complement this with GPT-2, which is really good at graphic design. So you take, you know, your first pass at GPT-2, you generate the thing that you want, and you're like, now I want to kind of change it. You bring it into Google Picks, and you're doing layer by layer, element by element adjustment without. I don't know. I'm curious if they'll let you. I hope so. I guess they would. they're like also the complete opposite of what it's doing has nothing to do with real photos per se yeah yeah look man they're good they're yeah they're good at making stuff sometimes they're not the best in name and stuff like yeah no that's that's amazing look we've covered io in the past but this feels like a really eventful one and a slightly less evil one if we're pivoting into Yeah, I also didn't touch, I know people are upset because they also announced that they're basically shifting main Google search to like AI centric search first.

34:35Separate topic, but I mean, how are they going to make their money? All that comes from AdSense, right? They probably already thought about all this. Moving on.

34:43Joey:Runway Aleph 2. Runway Aleph was the one way as model. That was like one of the original edit video, video to video, give it a video, just tell it what you want to change. It was okay. How was your experience with that? It was okay. I thought the quality lacked, but the promise of Aleph was one of the first, if not the first video-to-video model, predated Kling and some of the newer ones. And I thought, yeah, this is absolutely the direction we should go. We should not have to generate everything from scratch. We could just generate something that's roughly there and then iterate and iterate and iterate on it.

35:20some of the generations that I ran had like high frequency noise and like blobs and things like that which I thought hey it's just the first generation it's going to get better but the promise and the vision for runway what I thought was really really strong.

35:36Joey:Yeah similar boat I've like I've always had a hard time getting good outputs out of runway it's just never really like vibed for me. All of everything I tested it would just either change too much about the video warp too much stuff too much would be soft or fuzzy it was just never really usable. Aleph 2 definitely feels like a huge bump in quality I'm not sure what the resolution is but just even from like lighting composition you know that the tone map looks a lot of that is subdued and just looks more and more natural. One of the improvements is you can get more specific about what you want to change so it has a slider where before you kind of had to describe your change or you can kind of change the first frame, but this slider lets you pick any frame in the video and then modify that frame.

36:24Joey:So you can give it a better starting point and guidance of like what you want to change. And then you can upload images or you could also change this frame and describe it with what you want to change. So it's like an image reference insertion on the frame that you'd want. Yeah, and obviously you could just download this frame and then like yeah but that's like round tripping whatever you don't want to be doing that really dial in whatever you want to change yeah but yeah at least they haven't built in here you can modify it here and then change it so that's the new workflow uh i did the same test where i gave it that the video of the um metropolis sculpture first image generated from the first frame kept too much yeah yeah yeah lacma railing and balcony so i didn't use this frame I ended up just getting one where the camera tilted down and used this bottom frame as the guidance.

37:20Joey:And this was the output. Yeah, it's so much softer and lacking so much detail than what we saw with Omni. And then this was the same. Let me change the shot to a dog or change the dog to a robot. This was the robot that it came up with. And I was like, sure, cool, works. No. and then there's the output. Oh, the leg disappeared for a frame of time. The legs warping. It's still moving like your dog, but not a robot dog. Whereas the Omni model was really the mechanics and the rigging was like a robot dog. Yeah, this is like they're feeding it through open pose and just extracting the anchor points and then just attaching it to the new dog.

38:08Whereas the Omni model was really doing something different. It fundamentally translated that animation into a completely robotic animation.

38:17Joey:Yeah. So you see here, like those are backward bending legs in the front. Like that's hard to do anyway. And yeah, like moving like that, that's how a Boston Dynamic robot spins like on all fours, but not a real dog probably won't do that. So from an animation quality standpoint, Omni nailed it. Aleph, not so much here. Having said that, it is still really successful in painting because you can see her hand on the leash and just overall crowd work and all that. None of that changes. No, I mean, at least it's not messy with anything else. Yeah, the shadow swims a little bit. The left paw, the right paw just kind of disappears.

38:56It's tough, man, because Runway's probably been working on this for, I don't know, six months to a year. They release at the same time as Omni. Obviously, Google is way more funded and has a bigger team, and they're going to release a more superior product. Unfortunately, the thing is they are coming out at the same time, and we're going to compare the two against each other.

39:15Joey:I mean, I felt like this is probably a thing where they, you know, it's like, when do you release it? And it's like, okay, well, like Omni is coming out and like touting doing the same thing. So like, we better drop this update. You know, if you have success with Olive, let me know. I'm always curious because I know some people like it and have used it, and I've just never really. I'll just say one shout out to Runway that the fact that being like a small company, they're not a Google or an open AI or, you know, they're still hanging with the big boys. Right. So Cristobal is doing something right.

39:43Joey:All right. Thanks for everything we talked about at denoizepodcast.com. If you'd like to meet us in person next week, we'll be at AI in a lot. Joey's going to be a little bit busier than I am. So I'll be walking the show floor. Come say hi if you see me. Yeah, I'll be buried in the back room. But let me don't forget to be there. I'll be around at the after party stuff. Thanks, everyone. We'll catch you in the next episode.

From the publisher

Joey returns from Google I/O with hands-on tests of Omni, Google's new video world model, comparing it head-to-head with Runway Aleph 2 on the same shots. Plus: Demis Hassabis puts AGI three years out, Google Flow gets agentic workflows, and Google Pic...

More from Denoised

All 101 episodes
“Nano Banana for Video”: The Simplest Way to Understand Gemini OmniDenoised · 40 min
Listen in VO