In short
Podcast Summary: Denoised - Episode: Why Veo 3.1's New Insert Feature Changes Everything
Overview In this episode of Denoised, hosts Addy Ghani and Joey Daoud delve into significant updates to Google's Veo 3.1. They explore various features like video ingredients, video markup tools, and improved frame extension. Additionally, they discuss other industry advancements, including Runway's new apps for visual effects (VFX), Netflix's Eyeline research on video reasoning, and NVIDIA's DGX Spark.
---
Key Topics Discussed
- Veo 3.1 Updates
- Ingredients to Video
- Allows users to input images, locations, and objects directly into the video, enhancing character and location consistency.
- Similar to previous versions, but with improved control over video outputs by tweaking ingredient ratios.
- Video Markup Tools
- Features an annotation marker, enabling users to draw on videos and specify where elements should be inserted, enhancing user guidance.
- Frame Extension
- Introduces an "extend" feature that utilizes the last second of video for context, improving temporal consistency and reducing abrupt changes in video generation.
- Cost-Effectiveness
- Using Veo 3.1 via Google's Flow platform is more cost-effective compared to accessing it through the API, offering better credit-to-cost rates.
- Runway's New Applications
- Runway Apps
- Custom-built single-purpose workflows designed to simplify user interactions by minimizing the need for complex prompting.
- Examples include apps for changing weather, backgrounds, and lighting in a scene, highlighting user-friendly AI integration.
- Advancements in Video Reasoning
- Netflix Eyeline's VChain Research
- Introduces a method for reasoning in video generation, allowing the model to evaluate and adjust its outputs during the generation process.
- Aims to enhance the realism of generated videos by incorporating temporal awareness and contextual adjustments.
- NVIDIA DGX Spark
- Introduction and Purpose
- A new powerful box aimed at handling large workloads related to AI and video generation.
- Priced around $4,000, it offers a compact solution for high-performance computing needs in creative workflows.
---
Key Takeaways
- The incremental updates in Veo 3.1 demonstrate an ongoing focus on improving user control and video production quality, suggesting a proactive approach from Google in response to industry needs.
- Runway's new apps signify a shift towards user-centric tools in filmmaking, emphasizing ease of use and accessibility for non-experts.
- Netflix Eyeline’s research represents a significant leap towards making AI-driven video generation more intelligent and context-aware, paving the way for more sophisticated filmmaking techniques.
- The DGX Spark serves as a game-changing tool for filmmakers and content creators who need high processing power in a mobile format.
---
Conclusion The Denoised podcast episode provides valuable insights into the latest developments in AI and creative technology within the film industry. By dissecting features of Veo 3.1, Runway’s innovative applications, and the advancements in AI reasoning through Netflix's research, Addy and Joey make a compelling case for the transformative potential of these tools in shaping the future of filmmaking.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Time for the AI roundup. Let's go. all right welcome to the other studio addy yeah hey man i'm on the west side it's nice to be out here a little bit cooler nicer for la tech week welcome yeah thank you we've got an ai event we're going to tonight it is a zoo all right big story this week moderate update to vo vo 3.1 yeah so not a full no it's not vo 4 but a lot of quality improvements and just sort of like product updates incremental updates yeah all right so the big updates in this one uh there's now ingredients to video so this was something that's already in vo2 but basically you can similar to like cdream or nato banana you can give it a couple of images locations people objects and it'll use them in the video so that unlocks a lot more consistent characters or consistent locations consistent things so it just gives you a lot more control yeah over the output as you know with baking it's not the ingredients it's the ratio of ingredients so how much control do we have over that well there was another demo that was interesting from google uh where basically it seems like there's also and this is all in flow specifically the google web app that is sort of the portal to use vo and that probably has the best options for vo outside of you could still do this in the api other tools but they had a it seems like it's sort of like a built-in annotation marker so in the video in vo there's a demo from google's uh twitter page that um they have an existing video and then they just drew a square on the video itself and said add a person here nice and then just out of the person on the video yeah so that's sort of like the best kind of guidance you could give it yeah kind of like the thing we saw with the first image hack but now it's sort of built directly into the product and you could do it in a video not like an image to video it's like you're modifying the video itself like we've seen with runway olive runway olive yeah one does one wouldn't have the i'm thinking of the big muppets um oh sort of workflow but that's different run away off okay yeah that's cool no that's very that's very good uh the other update is like there are there's you know quality of life now there's start frame and end frame they used to just be start frame only so just you know more improvements to that and then the other interesting one is there's now an extend feature which is not new-ish because like other apps have extend features but they sort of would just take the last frame for a while yeah they would take take the last frame and use that as the first frame for the next one but this post says that video 3.1 uses the last second as the driving force oh so it goes further back in time yeah which is cool because you always get that weird issue where you get that stutter working completely like the video could completely just completely change like that's not what i had zero context about what happened before and just be completely random so the fact that takes a full second yeah of the video as context is pretty cool temporal consistency exactly really good good job google you guys know what you're doing i haven't taken this for a spin yet i am really curious i wonder if it's available on free pick i think some stuff i'm curious that's the issue you know with like what's available directly on flow and what's available on free pick but free pick would be in the same bucket as every other company where it's using it through the api i think and looking at the google's blog post the insert elements which i just talked about where you could drag a box of the video no ingredients is your input so ingredients is like here's an image of a person here's an image of something it's like the basic components yeah that's like instead of a starting frame it's like a couple of images that's the ingredients i think you can use that on the api the insert is i have an existing video i'm marking up my video and i want something to change oh like i want to insert something or i want to remove something yeah i think you can only do that on flow it's like video in painting exactly yeah yes so i think that's just flow i don't think that's an api thing that's amazing yeah so another reason also i mean flow is still the the the google ai subscription is still like the credit for credit cheapest way to use vo3 like the credit equivalent of the money you're paying there is cheaper than if you were to buy through the api really yeah if you got a subscription to google ai studio the math if you're making enough stuff the math comes out where at least it was i don't know if it changed with the api rate dropping it was cheaper it was like half the price to just do it in google's platform than to uh pay pay for the api remind me again is uh sora 2 like severely undercutting google or are they at the same sort of level i think they were a bit less but i did the math because uh there's a sora 2 api in comfy ui and so that's my easiest way to just be like what's this cost because they'll just tell you the price right there and so i cranked all of these settings up to the max so i did like hd the max duration was like 12 seconds and then with audio and that was six dollars for 12 seconds that's like a coffee at starbucks yeah yeah wow better not mess it up here we go i think if you were to do 12 seconds for vo you'd probably be a little bit more if you're to max out all the settings like so vo would be over six dollars if you maxed everything out do you think for that duration i mean vo i don't know if 3.1 increased the max duration i think it's still similar max duration maybe five so it has the extend feature but i think what it can generate is still maybe up to eight seconds i'd have to confirm that i think some sort of stuff was undercutting but then when you looked at the actual specifics i think it's roughly the same not as drastically undercut as you would think oh that was also using sora 2 pro which is the most expensive version of which was the version if you got a subscription to open ai pro or chat gpt pro or whatever or access it yeah or we're accessing it through the api okay soar to basic is what you get if you use the soar app which is not lower fidelity a little bit off topic but i don't know if you were the one that sent it to me uh there was uh like a youtube video of a guy who was talking about oh google yeah they have vo4 five six seven and eight and nine it's just a matter of when they're gonna release it and i i just i just had to laugh i mean it sounded so silly everyone can everyone can count numbers yes we can i'm sure there will be a sora 20 at some point sure no but no his claim was like google already has it it's like different versions of there was a weird thing with vo 3.1 where it it was like the weirdest companies like the most obscure ai companies were trying to like be the first to announce vo 3.1 and it was like where is this information coming from and are you like leaking stuff that you shouldn't have been leaking right by being an api partner uh it's like leaving the iphone at the bar kind of thing yeah it's like because you just want to be first to announce something about this but i'm like i don't i'm not going to trust you if you're the only one saying oh hey we got the inside scoop on vio 3.1 yeah and like until it comes from google like i'm not going to believe what you say do you think a lot of the view improvements are being driven from like the aronofsky feature or the michael keaton feature like are real filmmakers having an impact on ai development i would imagine they're taking that into account i don't really know how much that weighs on the roadmap or or these improvements because also a lot of these things were things that existed in vo2 they hadn't loaded them into vo3 yet yeah it seems like the arms race is not well it's still quality right like as we see with sora 2 like the quality bar jumps and then the next one will probably even further so it still feels like in a different bucket for me like yeah vo3 is in its own class vo3 definitely feels more cinematic filmmaker focused where sora feels realistic quality Sora is more like anything yeah make anything is like memes yeah so I I guess the question is like is the arms race really in the quality side or it sounds like there it's a lot more on the workflow and I think it's the workflow and the control yeah because I mean I look I've used flow you know a bit but it's like still I like like workflow wise I like something like comfy where it's just it could kind of a bit quicker and easier if I just need to like spit out a bunch of outputs or I just need to like plug a few things in.
7:51It's, I still find it's the fastest way than a web interface where I'm like typing a thing and hitting generate and then waiting and then scrolling back and trying to find the thing I made from a past version. And the, the organization tools with all of these are still kind of lacking. Okay. So as a professional, you're still leaning towards comfy with like a VO3 API note. Yeah. I still just like, I keep going back to comfy and API. I mean, the problem is I'm paying for the apis and so like i'd rather use the subscription that i already have the credits on so that's the that's the that's my personal hiccup sure like i have the free pick that i keep going back to free pick because i have the unlimited image generation but free pick's definitely the best a lot out of the interfaces but still you know not as good as as running stuff yeah i mean uh yeah like we covered on the show like joaquin doesn't pay us to do say this but it really that that being like such a seamless aggregator of things i yeah i mean i started to pay for it recently after hearing you talk about it.
8:47Yeah, it was a good idea. I mean, it was, yeah, like it was a smart idea to make it unlimited. You don't have to worry about the credit stuff. Right. Yeah. And I mean, I keep hearing FreePick come up more and more as like a central aggregator platform for a lot of people using AI. Yeah. Very cool. All right. Other updates from Runway. They have a new thing called Runway Apps. Nice. So it basically seems like it's more of a kind of like the ChatGPT apps or something where it's like a custom built workflow that just does one thing. But it's sort of a way where you don't have to worry about the prompting as much.
9:18It's like someone, I think I don't, it's not quite clear if they're making them all or if it's an open platform that anyone can make the apps. But like, for example, they have one that's like change the weather app. So you just give it your input image and it just knows like, okay, you tell what weather you want to change it to, but you don't have to like get creative with the prompting. You just like change it to snow and like behind the scenes, it has the prompting structure to give you a good output. Yeah, maybe what's, what it's doing under the hood is it's turning your prompt into uh like it's noting the seed and all of the different values of the generation and then it's creating like a custom preset and storing that custom preset i mean it could even be like a low-rank adaptation of the runway model that it's storing every time uh it's it's pretty brilliant like in a course of like a movie right you're gonna probably have 15 20 different shots that need the weather change to the weather that's in the movie so instead of trying to repeat that for all the 15 shots and get 15 different versions of it you rather just use one preset and get it every time yeah and like and get the output you're looking for and not have to be like let me reinvent how to prompt again to get the specific change that i'm looking for yeah i would totally use this for a lot of the color grading stuff i mean not to sort of eliminate color grading altogether but at least get it in the ballpark of what I was thinking, what the color palette should be.
10:38Yeah, some of the other ones they have, change background, change time of day, relight scene. Yeah, I mean, I think it's a clever idea. I'd be curious, I'm trying to see the demos. I'd be curious to see, you know, because it's like, if you were to try to do, if you're like, I need to change this to rain and also like change it to daytime and you keep reprocessing a scene, I feel like it would just keep warping. Yeah, exactly. It'll degrade over time. Yeah, because the issues I've had with runway olive is like it depending on the type of shot and what you give it it sometimes changes too many details well maybe crystal ball you can come on the podcast and tell us how it actually works but i think this is a clever idea where it's like okay i just need to do one specific thing and someone's figured out a workflow that does that one thing really well i don't know to like try to refigure out how i gotta say like strategically runway seems to have like a such a unique path in the way it's approaching the tool sets.
11:33Google is approaching it in a way that I think a normal filmmaker that's, you know, that that's used to a pen and a marquee tool would use it. But Runway is approaching it in a way that a non filmmaker would approach it and sort of break down what are the biggest challenges with filmmaking and how do I resolve it with presets or tools or workflows and they all sort of have their own dna like if you look at you know luma right ray 3 yeah like they're solving a certain set of problems in a different way uh and then runway aleph and now this you know they're solving it another way nobody's wrong here they're all sort of chipping away at the bigger problem is like what is filmmaking of the future look like yeah which is fascinating that we get to live through this era and actually see all of the stuff get built yeah i mean maybe you know this is just sort of an app thing you run now but maybe it's a node in the future and then you got your clip and then to what i was saying before and when it gets good enough where it doesn't warp the video you can have a stack where it's like oh run these clips through the change weather node to the relight node to the you know uh change lighting to daytime node yeah also you maybe you can mask out like unaltered frame here like don't touch the frame in this region and just do it this region and so on yeah also just imagine the point when this does become kind of real time and you're just like you just have your video playback you're like uh more weather and it changes there like a slider and you're like no less rain right turns the rain off and then it'll be in da vinci resolve like everybody will have it for 300 it's crazy yeah as you're running it on your dgx spark which uh we'll talk about in a second yeah so i you know i follow eyeline closely eyeline is um the vfx side of netflix as a business and look recently they've completely hit like a big reboot button where they uh previously they were more of more less a vfx vendor they're used to do shots and things like that and now they're like a proper research arm yeah as well as r &d yeah like their r &d team is really strong and a lot of them are here in la i think they're spinning up a new team in korea as well as vancouver uh paul de bevick who's at eyeline i think he's the chief research officer there yeah i think that title sounds who was on your panel.
13:50So he's one of the names on this paper. It's called VeChain. The title is Chain of Visual Thought for Reasoning in Video Generation. Okay. So I always get to do this research paper. I'm the least qualified person to do research papers. I don't have a PhD. I barely have a bachelor's degree, but let's do it. So when a video generation occurs, you give it the text and then that goes into a clip encoder the clip encoder is aware of what elements are needed in the video generation and then it also populates the temporal latent space so you have the elements as well as a timeline of when those elements are appearing now for better or worse that's a random draw right like the example they're using here is a rock and a feather get dropped and then the feather slowly kind of does the forest gump feather thing and kind of lands gently and the rock just goes that yeah so if you give this to if you give this prompt to the video generation model it'll give you a completely different outcome than you like the feather could just fall straight down like it would have no notion of a feather falling elegantly So VeChain introduces reasoning and a level of IQ to the video generation.
15:09So it actually has a inference time rendition of what is happening and being able to correct it as the frames are generating. So if you tell it, the feather needs to have like a sway to it as it drops. So after, let's say it generates 10 frames, it'll go back and say, actually, the feather needs to slow down here for a minute and then go to the left. and then it'll keep going. So it's like looking backwards to check its work as it's generating? Yeah, and it's introducing a level of almost brain power to the video generation in real time as it's inferring. It's like a reasoning video model, sort of.
15:42Like with the reasoning models, they're like generating stuff and then they're kind of going back and checking and like thinking like, oh, is this what I should have been doing or looking at for video instead of an LLL? Yeah, it's much more sophisticated than just like going from, you know, noise to diffusion to, you know, latent space than decode is certainly doing that but then if you picture like a like a element of brain power that's like supervising the whole thing as it's happening it's still in the research phase now i'm sure they have a version of it kind of up and running head eyeline that they're playing with because this stuff was generated and shown in this research paper here but it's really interesting work and this this follows up on the uh go with the flow research paper that we covered on the podcast a few months ago that's right okay i was like i remember something sounds familiar where it would check itself again so that was using something called warp noise where you know it's not just denoising noise like the name of our show it's not just about denoising noise but it's about warping the noise into shapes that then the denoiser can pick up on okay right yeah so that will help you really dial in motion correctly so vchain also is all about control and controlling motion and temporal domain stuff um so all of the research i think is pointing to uh like an area where i i don't think like somebody like google or sora would have figured out because eyeline is directly connected to filmmakers like more so than google or sora so they would have the need for the highest level of control yeah and this this looks like a step in that right in that direction exactly yeah i'd be curious what because i mean it's not a it's a way for the models to work so i'd be curious like you you would attach this to a model or just hope one of the models incorporates this into the way that it generates that's a great question yeah it's not i mean i haven't read the entire paper but i feel like this is some this is perhaps an llm engine that ties into video generation models yeah so theoretically if this v chain node was available and comfy maybe you can use it with any video model yeah basically i'm trying to figure out how do we use this now because this sounds pretty cool yeah don't we use this to improve uh video generation yeah yeah like could you attach this to like a wand 2.2 like workflow kind of thing so it would it would tie directly into the k sampler where you know frame by frame the denoising happens and then let's say after 10 steps the v chain does something to correct it and then after 10 more steps it'll do something else and so on yeah so we'll see that'd be cool all right and then uh last one that crossed my radar was uh the dgx park which we have the little gold box a little box which we've talked about quite a bit yeah and was first teased at ces uh it is now shipping nice finally yes how much is it a good question i don't know what the pricing actually i've heard five grand but i don't for starting let me see if there's an actual price posted on it but the thing that it popped on my radar because comfy ui oh i can't imagine how nice hosted that they're like you know they're running it native support four grand starting at four grand for a four terabyte i mean that's a bargain because you're looking at the high-end gpus and video makes at four plus grand price range right yeah see i come here has a blog post saying that um they're supported on dgx bark and let's see performance wise it doesn't beat a full desktop with a 5090 but it can run models and workloads that are too large for even high-end desktop systems uh yeah because it's on the blackwell architecture which is their next-gen architecture it probably has a higher memory than rtx lines yeah but the clock speed on rtx is probably faster so although you can you the problem with most comfy instances on most people's desktop is that their gpu can't store the entire model on its own and then when you have to then split that across two different clock rates or cycles then now you're slowing it down by half yeah so you want to be able to load that like yeah one one 2.2 i think is like 20 plus gigs yeah at least right so most gpus memories is not that high yeah but blackwell architecture i think is okay but not so it could do larger workloads but not as fast i yeah i mean i'm clearly guessing here depending on what the GPUs are driven by two things um clock rate and memory so uh what you know clock rate those are when you know when gamers overclock their GPUs that's what they're doing is they're just making the GPU just work faster even though it's not finishing like a cycle of computation it'll just keep going to the next one and then that's what overclocking is so that clock is directly correlated to how much power it has access to how much cooling it can do because you can melt that thing easily right so a little box like this is probably not going to have a high clock rate because it's a low power device it's probably not going to be cooled as well as like a desktop gpu with a giant fan on it or liquid cool so that's my guess um but having said that it's still going to be way better than any little box that you buy this is the best the best little box yes exactly all right cool yeah i'm curious to see the comfy blog said that they'll have benchmarks in a future blog post so i guess after they test it but yeah i'm curious i mean that seems like the most obvious for the stuff that we're involved in what else would you could you use this yeah i just want to give a quick shout to our friends at nvidia if you want to send us a dgs box we'll definitely take a first spin we'll mess around with it yeah but yeah what else i mean what else would i use dgx for in the filmmaking world what else could you could this excel anything i would imagine in the traditional computer graphics world would benefit from this So, you know, if you're an artist that's using Unreal Engine, real-time renderers, you have a giant desktop that's freaking loud, right?
21:39Why not run it on this thing and you could be on site at the shoot, you know? So like virtual production, real-time, little shoulder rig can have way better visualization on it. Yeah. You know, anywhere where you need localized GPU power. Yeah. I would imagine this thing would fit the bill. Lightcraft is a good example. Maybe they can run higher fidelity stuff on there. On that or process the stuff afterwards faster locally. Yeah. And then even stuff like codecs, like compressing and decompressing videos in real time. A lot of time that happens on site and that happens on giant workstations where now you can probably do a lot of this stuff fast.
22:18Again, I'm just speculating here, but when you have a fast GP on a tiny little box, there's a lot you can do with it. All right. That'd be cool to check it out. Yeah. All right. We're good. Yeah, we're good. Okay. links for anything we talked about at denoisepodcast.com and thanks for listening on spotify we had another five star review bam bam here we go if you're watching on youtube give us a comment give us a like and hit the notification bell wow i never thought i'd be saying that but i'm saying that now you should hit the notification bell all right thanks for watching everyone we'll catch you in the next episode
From the publisher
Addy and Joey dissect major updates to Google's Veo 3.1, including ingredients to video, video markup tools, and improved frame extension. They also examine Runway's new specialized apps for VFX, Netflix Eyeline's research on video reasoning, and NVIDIA's new DGX Spark.
--
The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.




