In short
Podcast Notes: Google AI - Release Notes
Episode Title
How a Moonshot Led to Google DeepMind's Veo 3
Hosts & Guests
- Host: Logan Kilpatrick
- Guest: Dumi Erhan, co-lead of the Veo project at Google DeepMind
Overview In this episode, Dumi Erhan discusses the evolution of generative video models, particularly focusing on the Veo 3 model, which features native audio generation. They also delve into the project's history, challenges faced, and user feedback shaping the future of AI-powered video creation.
Episode Breakdown
- 0:00 - Intro
- 0:47 - Veo Project's Beginnings
- Started in 2018 under Google Brain as part of a Moonshot program.
- Initial focus on generative video models; not seen as transformative at the time.
- 3:02 - Origins in Google Brain
- Emphasis on exploratory research; building on foundational AI breakthroughs.
- Focused on video generation as a challenging research problem.
- 5:07 - Video Prediction and Robotics Applications
- Early ideas centered on video prediction based on contextual data.
- Interest in applications for robotics, predicting future states based on actions.
- 7:45 - Early Progress and Evaluation Challenges
- Progress made, yet challenges in evaluating and scaling video models remained.
- 10:30 - Physics-Based Evaluations and Limitations
- Discussed limitations of physics-based evaluations for generative models.
- 12:18 - Launch of the Original Veo Model
- Original Veo model was launched; set the stage for future developments.
- 14:06 - Scaling Challenges for Video Models
- Challenges around training and scaling video models effectively.
- 16:02 - Leap from Veo1 to Veo2
- Transition to Veo2 focused on quality and immediate user availability.
- 19:40 - Veo 3’s Viral Audio Moment
- Introduction of interleaved native audio, creating a buzz among users.
- 21:17 - User Trends Shaping Veo's Roadmap
- User feedback highlighted the need for audio capabilities and ease of use.
- 23:49 - Image-to-Video vs. Text-to-Video Complexity
- Discussed the differing complexities of generating videos from images versus text.
- 26:00 - New Prompting Methods and User Control
- Introduced innovative prompting methods, allowing for greater user control.
- 27:55 - Coherence in Long Video Generation
- Addressed coherence challenges in generating longer-duration videos.
- 31:03 - Genie 3 and World Models
- Explored connections between Veo and Genie 3, particularly in world modeling.
- 35:54 - The Steerability Challenge
- Discussed the complexities of user steerability and control in video generation.
- 41:59 - Capability Transfer and Image Data's Role
- Analyzed how advancements in image generation influence video models.
- 47:25 - Closing
- Encouraged user feedback and insights for ongoing improvements.
Key Takeaways
- Historical Context: The Veo project began in 2018 and has evolved significantly, driven by a desire to push the boundaries of video generation technology.
- Generative Models: Generative video models present unique challenges, especially in evaluating quality and coherence over longer durations.
- User Feedback: User preferences have actively shaped the product roadmap, particularly highlighting the importance of audio in video generation.
- Research Challenges: Ongoing challenges in video generation include evaluating coherence and the complexity of transitioning from image-based to text-based generation.
- Future Directions: The conversation hints at future innovations in user control and capabilities in video generation, focusing on enhancing user experience and satisfaction.
Conclusion This episode provides valuable insights into the journey of the Veo project and generative video models at Google DeepMind. It emphasizes the importance of user feedback, collaborative research, and the challenges that lie ahead in making video generation more intuitive and impactful.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today we're talking with Dhumi Erhan who is the co-lead of the Vio project. We've sort of been working on this since 2018 or so. Like in the Google Brain team, we formed what was then called a Moonshot program. We were part of the Moonshot program called Brain Video. The mission was to just do video generation. I don't think anyone in 2018 was thinking about generative models of videos as something that would be somehow transformative. Everyone is super jazzed about video right now. And the thing that I think has blown people away is this interleaved native audio capability. It's cool. it's entertaining.
0:35Put it in the hands of people. So that they can use it themselves. And then once we ship them, they'll be like, how did we live without this until now? Hey everyone, welcome back to Release Notes. My name is Logan Kilpatrick. I'm on the Google DeepMind team. Today we're talking with Dumi Erhan, who is the co-lead of the VO project. Dumi, I'm excited for this conversation. Everyone is super jazzed about VO right now. So this is going to be great to hear more about it from you. Sounds good. excited as well. Yeah, obviously we launched VO3 at Google I.O. this year. And the thing that I think has blown people away is this interleaved native audio capability as the video is being generated in addition to sort of the just like good video quality.
1:19We've obviously been pushing like VO2 was super popular. We've been pushing on this and like your team has been pushing on this formerly in Google Brain and now as part of DeepMind for like a super long time. So I'm actually curious like the historical perspective to like get us to this place that we're sitting with like state-of-the-art video models that have audio capabilities and all this stuff. Yeah I mean definitely we've sort of been working on this since 2018 or so. Sort of yeah as you said like in the Google Brain team this was sort of like a Google Brain was a kind of a fun exploratory research group.
1:54There's a lot of stuff that came out of it. I think most famously the attention is all you need paper that's transformers work um and you know it was there was just a lot of like desire to push the batteries so of what we can do with machine learning right like so there was just a lot of desire to just do stuff you know like you know do challenging things and see where that lands right so i don't think anyone in 2018 in the ai world was thinking about generative models of videos as something that would be somehow transformative right like i think it was it was just more like okay you guys can do it i think from the external perspective um but we were sort of lucky that like jeff dean and others they were sort of like yeah this sounds cool uh we're we have a lot of data we have a lot of compute you guys should try um so we we formed what is was then called a moonshot pro moonshot moonshot program we were part of the moonshot program called brain video and it was about essentially the mission was to just do do video generation see what see how far we can oh yeah how far we can push the boundaries of that that work um and and was it similar to like from a from a mental model perspective like the form factor or like the the like outcome of generating the video would have been the same it's like you make an ex a second video or something like that or were you like trying to generate like a different format a video or a different like outcome for a video from that perspective?
3:19So we actually did a quite a bit of like you know meandering in like you know how exactly we are going to try to do this. I think our initial sort of hunch was that we should be thinking about this as more of a video prediction problem. So in the sense of like oh we have some context you know of like what is happening in the world and we try to understand what happens next. And this was sort of like are rooted in a bunch of different reasons. One is I think in principle, that is an easier problem than just straight up generating from text. It's a more well-posed problem, right? Like you have a lot more context and therefore sort of just predicting what happens next, predicting the next frame, if you have a bunch of different phrase before that, is relatively easy, right?
4:01Like predicting eight seconds into the future from that can be harder, but like it's generally easier than just being like, okay, I've like generated a cap, you know? Like that's a pretty challenging sort of, you know, without context. Like what kind of a cat? Where? What? What is the cat doing? Et cetera. So I think for a while we're just excited about like also trying to use these models in robotics applications. And there it's kind of, you know, often useful to think about like predicting the future. Like, oh, you have a robot. It's trying to do something in a situation. It's observing the world.
4:40And that is trying to understand what will happen next, especially conditioned on the robot doing something to the environment. Right. So then predicting the state of the environment after that action is important. So it was both a kind of a research problem about like, how do we generate pixels, but also rooted in kind of the desire to perhaps enable interesting robotics applications. so yeah that's and any of the like early work that has like uh as you look back like translated well to what we're now doing at the end i think about the my analogous example is i i was watching the thinking game a couple of weeks ago and it was interesting to see like how many and maybe this is like a slightly outsider perspective not sort of grounded in the actual research but it was interesting as they were describing some of the problems of like solving alpha fold and how like this like really limited data set and like how you scale rl and a really limited data set like that how much that tracks to like today's uh domains as well um so any anything that has like sort of carried over or like surprisingly from that perspective i mean i think in some ways what has carried over is that like we have made uh an immense amount of progress like there's no like doubt that if you look at the videos that we were generating 2018 and compare them with what we have now you you'd be like okay this is this is not it's not even in the same galaxy sort of in terms of quality and sort of uh the capabilities and all that stuff but the fundamental problems are actually quite there it's still surprisingly there like we still don't really know how to evaluate these models well um there's a lot of like kind of just research decisions i was actually having i was on a flight yesterday and i was i was bored and i was just like pinging a bunch of my team people about like hey have we ever questioned like why are we doing this there's a sort of sort interesting sort of you know model decisions that we have settled into not just the team but the field as a whole and mostly out of convenience like we we have just decided that this is what works now but some of these things that we were doing in 2018 we they could if we scale them up to like the number of tpus that we have today uh possibly could work better we just never bothered to.
6:52So we like, I think unlike LLM's like language models, we like, while I think on its surface, I think video generation, like you could think of it as been solved today. I don't think it's by any stretch of imagination true. And we have also not found like a good way to like the, the sort of transformer architecture or the inductive bias that is almost certainly correct. Like we have a lot of good ideas and the field is sort of is kind of bursting with like, you know, interesting hypotheses and whatnot. But even just like how do we measure progress is still sort of pretty challenging. And that was like big challenge 2018, big challenge now.
7:42yeah how are you thinking about trying to solve because i we were talking off camera about this like way in which we eval video models and we actually had a conversation with um some of the like native image generation folks and like actually a similar problem which is like how do you eval something that is like somewhat subjective in some context um yeah i mean it's tricky like there's all these like automated automated metrics that we've uh developed even as a group back in 2018. We still use them even now, actually. They're kind of good to see if your model is completely failing. So you'll be like, oh, okay, well, obviously this ablation is bad.
8:22I don't even need to sample from this model. Like this metric is... So that's kind of eliminating garbage models, right? Okay, fine. Actually, hill climbing kind of metrics that are useful are not that many. we can sort of you know like and and maybe that's not something we should talk on camera i guess but um but like in terms of like human evaluation i think we do a lot of that i think i think we we're sort of you know we we we try to put our models you know in kind of arenas that exist in the world like have people just rate their preferences um it's it's it's hard i think people have a lot of subjective biases that that kind of that just they just exist right like you can you can just make something prettier like we were talking about crank up the contrast and you can potentially get better scores um on some evaluation with humans is that does that lead to a better model probably not yeah um so similar as i think as you pointed to just like making the text longer uh when you when you when you output a response from an llm so um similar similar issues i think yeah human preference uh evals are tough uh tough if you don't also have a bunch of like truly grounded sources of truth and i think it's what has made some of the like rl hill climbing on text model stuff so interesting is like you do have good like ground truth sources and like it's nice because you can then actually go and help climb on yeah i'm jealous of llms in general because like the world there is tokenized like in some ways like there the way in video generation and video understanding you you basically have to sort of like understand you know that the tokens are not given to you right like you have to basically like figure out what is important in the video and and sort of extract that information and effectively basically and then kind of you know figure out how to predict um and that's just not given to us in text is that yeah uh we're talking about the one of the ways of um like potentially using like physics as a as a proxy for this and you had an interesting perspective on this about like why that actually intuitively doesn't actually make sense in practice for for video models yeah i think i think you know there's There's a lot of work, actually, there's a lot of work in GDM as well about like, oh, let's use kind of physics benchmarks to measure how good video models are.
10:54And I think this is kind of an interesting and exciting line of work. It's sort of tricky to do this in a kind of correct way, I think, in a principle, kind of in a rigorous scientific way. But, you know, people are doing their best. I think my point of view is that like I think a lot of people what they're doing using these models for VO3 specifically is to generate things that are not actually physically plausible right so then like like you know like they're not trying to generate reality tv some people are and maybe we can like measure how physically plausible that is but generally it's all about like ASMR videos of someone cutting a gold block or you know some sort of you know fantastical planet that doesn't exist any or or you know some zombie invasion or alien interviews on the street or things like that the yetis the yetis are yetis yeah yeah yeah all that all that stuff right um it's cool uh it's it's entertaining um it's definitely unrealistic physics yeah um doesn't mean it's bad so so i think it sort of gets to the crux of the problem is that like you you solve some problem of like satisfying the user but perhaps not in the way that like any particular eval would have shown it to you if you were focusing on that yeah I'm curious about this story of how you know 2018 the project starts in Google Brain doing a bunch of this stuff when when does VO sort of enter the enter the stage I actually forgot when the original VO model was launched.
12:34Yeah, not too long ago, actually. Yeah. So VIO was a project launched at IEO 2024. Oh, wow. We did a bit of meandering. I think generally, like, perhaps we started looking at visionation a bit too early in the sense that, like, the amount of compute needed for this was to get to a good model. It took, you know, took a while to catch up with our ambitions. um so i think that's you know like it maybe 2018 was a good place to start we kind of laid laid a lot of the groundwork um but we were like we just i think realistically looking back we would not have gotten obviously like vo3 quality uh like in 2020 um it just wasn't possible but I think also it was good to start early because we tried a bunch of ideas I think that was also one of the benefits of kind of you know working in the brain team and the GDM after that is that you have a bunch of people trying a bunch of different approaches and it's not clear immediately before trying them which would which of these approaches would benefit from scale I think that's the part that's interesting.
13:47It took us a while to get to a regime where, okay, if we do this and we just crank up the number of TPUs in a principled way, we will get significantly better results. So this is how VO came to be eventually. We're like, okay, now we know what to do. Crank it up. Is that a distinct property of video models relative to like the the tech or like other model regimes where like it's it's harder in video to predict like what is the thing that yeah what you should do in order to like actually keep making the model quality scale up i think yeah it's basically like you know in machine learning we have this like this notion of an inductive bias you know you have to make some choice you have to sort of think of it like of like what how do you what kind of assumptions do you make about about the world, about the data, that the model basically would fit that data, right?
14:42So, and I think for text specifically, kind of the transformer architecture, even from 2018, is basically almost the ideal inductive bias. So people have settled that and then they have, you know, that paper has shown pretty conclusively that you can just like scale up, you know, that was the beauty of that paper. And then, you know, again, like for people have figured this out for images later on. There was an image transformer paper, et cetera, et cetera. And then for videos, I was just looking at some document, which was pretty amusing, document from 2018, about 20 % projects for the team. And the first one was, oh, just, you know, create a video transformer.
15:25And it's like, target three to six months. It only took us like, you know, about six years to get there. So I think we were sort of, it wasn't, I think we were like deluding ourselves how easy that task would be. Like, well, text was easy. Image was very easy after. Video should be obvious. But I think it was just a lack of both compute and the right kind of inductive bias, principled way of coming up with a modeling technique, which is really kind of, you know, kind of pliable or whatever, amenable to scaling, I guess. Yeah. So we launched VO, original VO, IO 2024. I think VO2 was December of 2024.
16:10That's right. What was the jump between VO, VO1 and VO2? The jump was, I think, VO1 was a high quality model, but it took us a while to get into the hands of users. We're like, OK, this is great. It's a flagship, state of the art model. but it just took a bunch of time to make it scalable enough compute intensive inference time so for VO2 we really tried to nail this super high quality model but also we can put into the hands of users immediately so that was the thing that we really targeted so like when in december when we shipped vo2 it was like okay it's here and it's for everyone um so i think that was that was sort of a kind of a big shift also in not just in in the vo team but in the kind of gdm as a whole yeah that we sort of try not to what people say uh derisively here is like we try not to ship blog posts yeah um like just you know some sometimes it's okay but like generally like well if we're bragging about a big big cool model put it in the hands of the the hands of people so that they can use it themselves.
17:29Yeah. And I think that was the strategy there. It was like, okay, we're just going to run with it, like make it significantly better. And it was like state of the art for like, whatever, six months after that, because I think we really tried hard and then also have it immediately available in the hands of users. And I think that model also started to have, I don't want to say the emergent property, but like the early seeds of this audio story, which I think is like a huge part of the VO3 story right yeah so we kind of initially planned on we we were we were doing sort of audio or thinking about doing audio for quite a while um uh even sort of there was this like um paper from some colleagues of mine called video poet i think in 2023 24 i forget that had like a model transformer sort of model for video audio video to audio and audio to video amongst other things they were not doing sort of joint audio video generation but they were clearly thinking about it right um in this context it was a good paper i think got like a best paper award at one of these con machine learning conferences so it's very cool um and in vo we were also like okay well audio clearly like something that we want to do and so for vo2 we really really wanted to have it but it just didn't hit the quality bar so like uh at at launch time we're like okay december we want to ship this and you know it was a kind of a late call they were like okay no we really want this to be a good user experience rather than just being first to market yeah so that was a bit of a risky move i'll be like okay we could have been first but with a subpar audio thing uh but what were the what was like the failure it was just like the quality wasn't good or just wasn't matching up with like synced up well i think the sync the sort of the speech uh lip syncing stuff i think that you know that ended up being a pretty differentiating big differentiating feature right like it's one thing to look at like asmr videos of whatever tomatoes being cut but the yeti blog and all that stuff wouldn't have happened without joint kind of av lip sync stuff right that that was the that was the thing that sort of created that viral moment yeah um because we're sort of we like seeing that kind of you know humans even if they faked talk so yeah yeah and so vo3 uh may at io this year from the from the team perspective like was it super obvious that this was like, I think externally, obviously everyone blown away is credible, best model ever.
19:54Was that super clear as you were like training this model and like doing the testing and like bringing it to people? Like was everyone as excited as the external? Sometimes there's this like big mismatch on either way, either we're overly excited or we're not as excited and the external reaction ends up being so different. I mean, I think we were excited, but it's so difficult to predict this. Like sort of, I think in general, it tells you that it's difficult to manufacture a viral moment like as much as we would want this like like it's not i don't want to sort of you know state that like yes we were all in it's going to be you know huge success obviously we we believe in in the project and everything like that but it's so difficult to know like to predict uh what what would like we were we were testing in in like in the dog fooding and the bigger thing before IO.
20:42And all of us, we thought rap videos would be the viral hit. These were the most popular things that we were generating in the tea. Those rap battles between Einstein and Newton or whatever. They're fun. I don't think any of them ever got popular. We could bring that trend to life. Yeah, I know. I'm just saying it was just like, okay, what will be popular? And then it was Yeti stuff. and things like that. So completely sort of, you know, different from what the expectation was. Yeah. How has that changed your perspective? And again, back to this like image generation, native image generation story, actually really interesting where like the use cases that in like failure modes that people were sending us for the initial version ended up dramatically influencing like what the 2.5 version of that story ended up being.
21:37So as you've seen these weird trends happen, has that influenced from your perspective what the North Star looks like or the things you want to eval or test or anything like that? I mean, I think what we've learned is that generally people like the outputs and often many of them want to be in control or want to sort of do some iteration over the outputs, things like that. I think, you know, definitely it's, like, sort of not, like, a crazy idea that we want to be able to give people more control. So something that, like, you were like, okay, how do we enable that? Like, is this sort of a model decision?
22:20Is it a product decision? Yeah. Like, where does that, like, where does that line get drawn, right? Like, you know, you could animate yourself, but then it would have a different voice for you. Right? Like, you could do this now. Yeah. That's awkward. Yeah. Things like that. Like, how do you solve that problem? Things, you know, like, how do you enable people to have a lot of control, but also not, like, have just, like, one zillion kind of, you know, drop-down menus everywhere? where like the whole beauty of the text of video and audio at the same time is that you could just like type inside and get like this is why you know people could generate video and audio after before that with like competitor tools it's hard but it was more like okay get this video and then like upload it somewhere else i tried to do this a bunch of times and then like put the text and then it's it's kind of wrong lip sync and all that stuff so you could theoretically do all of these things before, but it sucked.
23:26Now it's much simpler. So we want to enable people to do more. And there's a lot of feedback about like, I want all of these things. But it's a sort of a kind of a classical sort of product, product kind of roadmap decision now, like what, how do we sort of prioritize and make, make good product decisions and also make the next kind of maintain the same quality and all that stuff. How, how you, you talked before about like the challenge of doing sort of like extending from a frame versus like starting from a text prompt. Is that something that's still like from a like model quality standpoint or like, I don't know, from like a compute standpoint, like is one of these things like significantly more difficult for us to do even with VO3 as an example, if I do like image to video versus text to video?
24:13I think from a compute perspective they're kind of they're kind of negligible in terms of like um uh differences um i think from um image to video is an interesting example because i think people in principle like the idea it turns out it's kind of complicated from a learning perspective usually people want take wanted to they're like okay i want this image and then i i want you to like animated in this certain way. But it turns out that, you know, the model would need to kind of contort almost the image to your text in a lot of ways. Like usually you don't come up with a prompt that makes sense in that image.
24:55You obviously try to change that image. So like, you're like, okay, here's this image and animate me like swimming with a mermaid, right? Okay, but maybe you're not like, you're like just sitting here on the Google campus. Yeah, yeah. Like there's a gap between where you are, where that start is, and where you need to be in order for that to make sense. So sometimes actually that poses a challenge, like from a kind of a modeling and learning perspective, like from what the user theoretically wants, which could be image to video, or it could be what we call reference to video, where it's like really what they want is some notion of themselves in a different scene, not literally a video that starts with this frame.
25:43So again, I think there's a bit of a mismatch between what sometimes academic benchmarks do or even arena do for image to video and what users actually want. How has your worldview changed about... with some of these like prompting paradigms i don't know if you've seen this about like people are like talking online about like json prompting and then the like drawing on the little drawing on the pictures and then being like move here do like move this object or like zoom in or whatever has that is that something that like um is sort of this like now i don't even know where these trends have came from maybe we actually put some of this stuff out i assume this is just people figuring out what works and then like sort of showing us actually some of these use cases.
26:33I don't know about the JSON part. I think I'm confused. Yeah. It's not clear to me. We've done some tests and I think like... So the model is not trained for JSON prompting. It's not meant to be. Interesting. Yeah. Like if it works, I think better for some classes of prompts with JSON prompting. I was pinging people to be like, we should provide some guidance to try to tell people this story but yeah yeah but generally it's if if it does work like that better it's accidentally i'd say interesting um the drawing part is actually i did i do think that was discovered by somebody somebody inside gdm um and don't know i think i'm not supposed to say how that works uh but it's cool it is cool it is cool and it's such a it's such a uh it feels especially like with video and image it feels like such a natural prompting technique you know yeah like i feel like i i just like i i get it like it's like this extra level of control that actually like is so unique to the to like the form factor of image that it feels cool that's right that's right and i think it sort of surprisingly solves a bunch of these challenges we were just talking about like two minutes ago like it is is like oh like how do i make like like how do i avoid having a bazillion in drop-down boxes or whatever, like, you know, like, oh, just like, you know, just image prompt with the stuff.
27:55I'm like, oh yeah, that makes sense. It also helps with, at least I'm lazy. And like, I don't like prompting, being like very precise about what you want in prompting is like this meta problem of using AI products where like, it just requires a lot and like actually just drawing it and being like, no, put that there and move this the air is like very simple to do visually one of the interesting like threads is just around like the length of generation and i'm curious to double click on this and like try to wrap my head around obviously on one hand there's like you know could the model just zero share a single shot like an entire movie or something like that obviously today is doing eight seconds um i've seen a bunch of comments or talked to a bunch of people and like it would be cool to do like 20 seconds or something but from like that from the model perspective like what's this uh the line of the trade-off between like the length of generation relative to control relative to compute relative to like all the other like is there like some quality trade-off on all that stuff um yeah i mean i think there's some sort of compute kind of you know uh constraints like generally like you you you know the longer it is for you to sort of kind of just generate in a single shot kind of like a particular video, the more expensive it is.
29:16And the more expensive it is to like train such models. So it's a trade-off, right? Like, you know, you have to have good reasons for why that you need that. And I think sort of eight seconds, I'm not saying it's like enough for everybody or whatever. Like, you know, people obviously want more. there's also a notion of like how to enable more but without it's again it's a kind of a bit of a user and sort of product problem like many of the prompts that we get are like you know super short five words kind of you know things and it's sometimes actually difficult to just fill in eight seconds of video from that right like you know imagine if there's like dialogue and stuff like that.
30:03Like you just have to come up, like the model has to come up with like, like fabricate, like lots of dialogue, unnecessary, right? Just filler basically, kind of conceptually, if you think about it. So how do you, how do you enable this without like just, just making something that then the users be like, well, that's not quite a way to have in mind. So you're wasting a lot of like, like TPU cycles, predicting something, making something that the user is gonna throw it in into the garbage right again you can make now now you can just sort of either stitch together or internally you could like auto aggressively to just extend kind of these videos uh to whatever an hour um i was joking earlier that i would probably have to pay you to watch just like you know an hour video like you know just from like one a single prompt uh like that uh they're not probably not going to be as interesting uh as as you may think uh they will just kind of converge to some some sort of weird slop um just uh and it's just not going to be that fun um is there something about this and maybe i so you can also correct if my understanding is wrong but like there's something about like long context as a as a proxy for this in text models where it's like the the more context you have like it's it's harder and also like there's multiple things that you're trying to attend to in context like it's actually like a really different like a difficult like coherence problem like do you sort of lose like the consistency of the like the model is if you were to generate an hour-long video it's actually not like thinking about the like in a normal movie there's like a plot and there's consistency and it's the same characters and you actually like lose some of that because the model can't see far enough back into like the context so there's definitely some of that right so so like one of the examples that uh corey one of our leads talks about um it brings up is is um you know generating a video of going on a mountain biking trip right like you're you're you're you're starting on a trail you go and it's a loop right so you go like and you come back right obviously like the last frame should be the same as the first one right um that's sort of the otherwise like it's not a loop so that's kind of almost an extreme example but it's a good one I feel like you can eval that that's a good example in some ways so you can hill climb on that no pun intended I guess you can do that and you see some glimmers of that in the Genie 3 releases they optimize for slightly different things than the VO folks they showcase a lot of that where the virtual camera turns away and then moves back.
32:46It's the same scene. So I think it's definitely keeping that context is important. How to do this in a kind of efficient and scalable way for like an hour? Very interesting research question, I think. Keep you busy on the research side. You mentioned Genie 3. I'm curious, we were talking off camera about this like pixel versus concepts question, which I think has a lot to do actually with like video generation and like understanding the world. And I'm curious to have you sort of double click on that on that problem again and explain it. Sure. So GD3 is, you know, I think it's advertised as kind of a kind of a world model, you know, basically.
33:31Right. So it's a real time. It generates the pixels of that world. and that's very cool and there's sort of this idea and we can do it for like you know minutes long I think and this is there's this sort of idea that you can in some in some maybe not too distant future you can have virtual agents basically kind of you know maybe virtual robots kind of explore generate virtual worlds for themselves and learn how to do certain tasks in those worlds and and then kind of learn policies, RL policies, whatever they are, opening doors or whatever, folding bedsheets, whatever you want your robot to do, and kind of transfer those.
34:16Especially if you have generated a photorealistic environment that looks exactly like this one, then you can do this. And there, actually, going back to our discussion before, there physics matters. You don't want unrealistic physics if you're going to be training kind of virtual robots in there. You do actually want to generate good physics there. So this notion of like, you know, oh, we just want to imbue kind of agents with this imagination, world models, if you want. Another word for world models is imagination. It's been existing for a long time. Like Jürgen Schmidhuber, I think, has invented this idea at least 30 years ago.
34:57how to do this? Like, is it, should we really like generate literal pixels of the world or should we sort of try to understand like, you know, all these robots kind of, should they be generating the idea of like merging into a different lane on the freeway or, or the actual effect of merging into a different lane on the freeway? that's this kind of unresolved research question. Like what do we want to predict? The actual future and the consequences of an agent's action or some representation of that future? And unfortunately, generally we only observe the future or the consequences of the actions, not the representations of those.
35:47There's no ground truth for representations. They're just something that we can make up, but they don't exist. Yeah, you mentioned to expand on this thread of unresolved research questions, but also to go back to some of the things that people are really enjoying and giving us feedback on. I feel like we've already got a great model that people love that does audio. I'm like, what more could we actually do? Is it just like continuing to like scale quality and like find losses of things that like end up being super weird? or like i'm curious like from uh um and i don't know if there's like a capability story as well where it's like certain like we've you know we've landed image to video as this example and it's like you know what you want is native video to video or something like that i don't know how much that makes sense but right i mean um i think we try not to be super kind of short-term kind of local optimum kind of optimist optimizers in the team where like you know i think kind of a contrived example would be audio if we were just took vo2 as is and just optimized for what people told us that they wanted audio would have never made it yeah because basically almost no one actually asked us for it uh but and it's kind of weird now now that we have shipped it like there was this funny twitter kind of a comment where like i think from is it justine moore or somebody like that one of those people were like oh vo3 has ruined everything because now everyone complains about other models where they don't have audio.
37:20And it makes sense, right? Like now, like, I don't know about you, but like if I see it like AI generated video on the internet and it has no sound, I'd be like, well, what's the point? Like I want, like to me, it's a defect. Yeah. Right. So we have a bunch of ideas and we're working pretty hard on that, but it would be something of that type where of course we're going to like improve quality and we want to be like unambiguously state of the art. um in this in the kind of joint av stuff um and we're gonna work with like you know the the variety of products where this has shipped inside google and uh and even outside whatever through cloud like to be like hey like what is it that that doesn't work for you like is it like is it length is it steerability something else but also we want to kind of you know delight our users with things that they don't know yet that they want uh and then once we ship them they'll be like how did we live without this until now our long videos that's yeah the long i'm joking this steerability story i'm actually curious any anything to double click on like um how you're thinking about this from a model perspective like what like i don't know if it's like the the researcher the trade-off in like um like what you provide in the at the prompt level versus like some like like parameters that and i'm not i don't think we provide a ton of parameters right now other than like the formats and the sizes and stuff like that so i i think for vo2 we actually did have more like there were some more of those or like some i think imagine has some of these on the on the image side so is that something that you think about like what do you provide in that versus prompts yeah so we it's it's again it's it's it's you know we we do provide certain things but then suddenly you know somebody discovers that like you could just like draw on an image and then you know somehow that works magically yeah um so that's cool steerability i think you know i would love to have a good natural way of kind of iterating on a on a you know having the user kind of you know kind of interact with the model but i think that that would be sort of you know my way of sort of not solving steerability but be like oh hey well i created this and like well i just want to modify something yeah like how do i do that like i think i think that would be an interesting challenge um and there's sort of interesting sort of research challenges there too like how how to do this in a way that is like doesn't involve like throwing massive amounts of compute right like we need to do this in a kind of economical way um like so there's there's interesting challenges of like you know like like can we provide you with a quick rough draft um i I think this is a consistent feedback from customers.
40:02It's like you have to like to you basically have to pay for the whole thing before you know if it's actually what you're interested in. Can you can you like partially solve this by just like decreasing the or does it not like just make one second of video instead of eight seconds? Like is that is that like a hack to this or is that is that not actually a viable solution? I mean, I think you can. I think it's a question of like, yeah, we could. I think it's tough because, you know, maybe the whole concept that you were trying to do is more of an eight-second concept, right? Got it. Yeah, that's fair.
40:40Like a car chase, whatever, falling into a submarine. You know, if there's like a sequence of events that you're trying to capture, the one second is not going to give you that. Yeah. So it's a challenging problem. Like, how do you sort of get into the user's head and extract what they want and then loop back in a way that makes sense from all sorts of perspectives. There is a sort of virtuous kind of loop that we have right now where we use Gemini's video understanding capabilities to generate, to annotate a lot of the data that we use for training VO3. That makes sense. It's not like rocket science.
41:20Rocket science. Having a good autorator is great. Yeah, yeah, yeah. It's basically like, you know, Gemini can be very good at describing what's happening in the video. Yeah. And this is what we can, this is how we can collect data, basically, to train these models. But there's no, like, to answer a more directly question, at this point, there's no, like, direct kind of, you know, loop of, like, it would be cool. And that's just because Gemini is, like, architecturally so separate from what VO is doing? I mean, I think it would be cool. you can think about it like you know you generate something with video and be like have gemini look at the output and sort of direct you know that would be another way of being like oh that's not quite what the user wanted yeah yeah you know like you know go back this ends up getting talked about anytime we talk about models uh on this on the show i guess um is this like capability transfer and like how like for example like in in video generation world like how us hill climbing on video understanding as an example transfers to video generation we released we announced back in i think it was like three months ago or something uh 2.5 having state-of-the-art video understanding i'm curious how that like impacts on the generation side yeah i mean so the impact is pretty direct for us is that we use gemini and its video understanding capabilities to to annotate our a lot of our data like with with basically with captions of what's happening in those videos.
42:43And, you know, we effectively try to learn the inverse mapping. I mean, that's basically what text to video is, right? Like it's trying to figure out, okay, how can we generate the pixels based on the caption that we have for them? So we use Gemini very heavily for that. It's interesting, most telling me in this example about how there's, and I forgot the of vernacular for this specifically, but there's all this like undescribed, like in the world of texts, there's all this like undescribed detail because it's just like not what's important about a specific scene as an example. So like, you know, there being stairs or a carpet or something, it's like normal human stuff, so we ignore it.
43:27And then if you were describing the world in text as a human, you would sort of ignore that piece, but then actually like the video understanding is helpful if you want to sort of try to capture like, oh, what is this, what's actually happening in some video? Is that, yeah. Yeah. I mean, I think, I think it's, it's true. And it's, it's, you know, one of the advantages of using something like Gemini to, to do kind of captioning of the videos is because you can get significantly more verbose things than if you just ask the human to do it. Like you can ask the human, you can pay them, you know, a lot of money and they will describe to you and, you know, kind of precise detail, every single kind of, you know, whatever object and their relationship and everything like that.
Read the full transcript
44:02But that'll be kind of, you know, relatively expensive. adventure um uh but yeah gemini can sort of give you that and actually people do want that like you know like speaking of control and durability you know there's there's a class of our users like many of our users are more like you know here's five words generate me a scene fine uh we delight them but there's also a significant percentage of our user base who are like write these like in a multi-paragraph kind of you know like i want this particular chair in this location and all that stuff and you know exact scenes right they want the they they have a particular thing at their head and they kind of you know like they describe it to you and sort of you know almost 3d kind of your details of like here's here's where's everything and like in order to solve that problem for them like we need that kind of data right like we we need kind of that precise annotations um and temporal sequences of of things that are happening in the in the in a in video.
45:01And this is where like something like Gemini video understanding is really useful and shines. Yeah. Same thread actually on this capability transfer around native image generation and also understanding like we have a state-of-the-art model that does image generation. How much is that? Like I, you know, video is just a bunch of frames of images being generated. Like is it, or is there, is there like some transfer of capability where they make that better? And that's like great for, for your team by default. Again, it's not quite, quite that direct. I think, I think there is an interesting sort of research thread on, as you were calling, it's like, oh, videos is just a stack of image frames.
45:36That can be quite inefficient. There's a lot of redundancies in frames, adjacent frames in a video. This is why video compression exists as a field, because you can just severely compress a particular video because of redundancies in the video. No, a lot of the times, actually, we're... And we've published a couple of papers on that before, a couple of years ago for kind of related projects on video generation. But like we use a lot of the image data in conjunction with video data to train better video models. It turns out that just adding image data helps significantly. And a caricature of this is that is a kind of caricature explanation of that would be that images, they just happen to have a lot more concepts than your random video.
46:27Like we just can mine or can just acquire a lot more images that are diverse compared to like the kind of videos that we may have in the training set. And it really helps you sort of, you know, just learn very specific concepts. You know, whatever shoe brand or car or whatever it is, you will probably have an image of that. whereas you may not actually have a video patch. It kind of makes sense. So it really, really helps in the interest of diversity and sort of having more concepts in your training set. And then videos, of course, are useful to learn what a video is. It should have used rocket science, like why we use videos to train video models.
47:17But it's not obvious to everyone why images help. So that's why we use it. Awesome. Well, Dumi, this was a fun conversation. People who are using VO3 today, feedback, send it to us on X and Twitter, I guess. Yeah, you can do that. There's like, you know, if you're using the Gemini app, there's thumbs up, thumbs down. We go through them quite religiously. You know, we're in tune with user feedback. So we want to make our users copy. Awesome. I love it. Congrats to your team and all the success of VO3. I'm excited to continue to see all this stuff happening. Thanks everyone for watching this episode of Release Notes and we'll see you in the next episode.
From the publisher
Dumi Erhan, co-lead of the Veo project at Google DeepMind, joins host Logan Kilpatrick for a deep dive into the evolution of generative video models. They discuss the journey from early research in 2018 to the launch of state-of-the-art Veo 3 model with native audio generation. Learn about the technical hurdles in evaluating and scaling video models, the challenges of long-duration video coherence and how user feedback is shaping the future of AI-powered video creation.
Chapter:
0:00 - Intro
0:47 - Veo project's beginnings
3:02 - Veo's origins in Google Brain
5:07 - Video prediction and robotics applications
7:45 - Early progress and evaluation challenges
10:30 - Physics-based evaluations and their limitations
12:18 - The launch of the original Veo model
14:06 - Scaling challenges for video models
16:02 - The leap from Veo1 to Veo2
19:40 - Veo 3’s viral audio moment
21:17 - User trends shaping Veo's roadmap
23:49 - Image-to-video vs. text-to-video complexity
26:00 - New prompting methods and user control
27:55 - Coherence in long video generation
31:03 - Genie 3 and world models
35:54 - The steerability challenge
41:59 - Capability transfer and image data's role
47:25 - Closing

