In short
Eye On A.I. - Episode #231: Paras Jain: The Future of AI Video Generation with Genmo
Episode Overview In this episode, Craig S. Smith interviews Paras Jain, CEO and Co-founder of Genmo, a pioneering company in AI-driven video generation. The discussion delves into the innovative video generation model, Mochi One, and explores themes such as the open-source revolution, the future of creative storytelling, and the potential of AI in transforming video production.
---
Key Themes and Concepts
- Introduction to Genmo and Mochi One
- Founder Background:
- Paras Jain, co-founder alongside his brother Ajay Jain, both with roots in AI research at UC Berkeley.
- Ajay's contributions to diffusion models and video generation.
- Product Overview:
- Mochi One: A state-of-the-art video generation model focused on motion quality and adherence to prompts.
- Offers unmatched precision in generating AI videos, significantly improving motion realism compared to competitors.
- Open-Source Approach
- Genmo's commitment to open-source facilitates rapid advancements and customization within the AI community.
- Allows developers to fine-tune models, enhancing character consistency and enabling unique applications.
- Technical Innovations
- Model Architecture:
- Utilizes a unique architecture termed ASYMDIT (asymmetric architecture) for greater efficiency in training and operational costs.
- Incorporates synthetic data pipelines for improved learning and quality in video generation.
- Challenges in Video Generation:
- Addressing issues of latency and cost, particularly for longer video outputs.
- Achieving facial consistency and character stability across generated videos.
- Future Directions in Video Generation
- Discussed the potential for interactive environments and real-time experiences that mimic video game-like scenarios.
- Emphasis on the scalability of video generation technologies and their implications for immersive media.
- Commercial Aspects
- Paras elaborated on the monetization strategies for Genmo, highlighting the balance between open-source access and the need for production environments for commercial applications.
- Addressed the competitive landscape, emphasizing the diverse use cases for video generation, suggesting that multiple specialized models will coexist.
---
Key Takeaways
- Mochi One's Capabilities:
- Delivers high-quality video generation with an emphasis on realistic motion and responsive prompt adherence.
- Currently supports video lengths up to 5.4 seconds, with plans for longer outputs and greater interactivity.
- Community Engagement:
- Over 2 million registered users testing and creating with Mochi One, showcasing the rapid appetite for AI video generation tools.
- Vision for AI Video:
- A belief that video generation represents the next frontier in AI, moving beyond traditional storytelling methods to create immersive and interactive content.
- Collaboration and Innovation:
- Ongoing partnerships with universities and the open-source community to push the boundaries of AI capabilities in video generation.
---
Episode Structure
- Introduction (00:00)
- Introduction of Paras Jain and Genmo.
- Video Generation with Mochi One (01:45)
- Discussion on the capabilities and features of Mochi One.
- Open-Source AI in Video Generation (04:41)
- Exploration of the impact of open-source on AI development.
- Building Mochi One (06:08)
- Insights into the model architecture and development timeline.
- Technical Challenges and Solutions (Multiple segments)
- Topics include reducing latency, character consistency, and synthetic data pipelines.
- Future of Generative AI (48:39)
- Speculation on the transformative potential of AI in various creative fields and storytelling.
---
Conclusion Paras Jain's insights into Genmo and Mochi One illuminate the rapidly advancing field of AI video generation. The commitment to open-source approaches and innovative technologies positions Genmo as a key player in shaping the future of video storytelling, making it an exciting space for both creators and developers.
---
Stay Updated
- Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
---
This episode provides a comprehensive look at the intersection of AI technology and creative storytelling, offering listeners an understanding of both the current capabilities and future possibilities of AI in video generation.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We were really focused on driving two things forward, being the best of these in the market. Number one is motion, having the best motion, because we feel many of the products in the market today, they don't feel like real video. They feel more like live photos, almost off your phone or GIFs. And our model, Mochi One, has the state of the art motion in the industry. And number two is prompted humans. It does what you ask, and that's really important. We have the fastest time to press pixel in the market for our model. We will start to see videos rendering off of the GPU within 10 seconds. So Paris, lovely to have you here.
0:33Why don't you introduce yourself to listeners? Tell us a little bit about the company, and then I'll start asking questions after that. It's really great to be here, to connect with the wonderful audience here. I'm Horace. I'm one of the co-founders and the CEO of Genmo. We are a company that is working to advance the state of video generation market. We are building some of the frontier models for this. We recently released Mochi One, which is one of the state-of-the-art models for video generation of the market. It provides some of the best motion we've seen out of technology today across the field.
1:15My background is I started my journey doing a PhD at UC Berkeley, working on large language models. It's insanely exciting technology. and I started the company with my co-founder and very remarkably my brother Ajay who was one of the pioneers working on diffusion specifically he is one of the co-inventors of the one of the most important architectures in image generation and so it was kind of a confluence of a perfect set of factors for us to start a company to work on video because we feel video is represents the ultimate medium for communication. Two questions. Runway and Pika have kind of dominated the public imagination in this space.
1:56How did you approach starting a company? When did you start it? And in the face of that kind of competition, and how do you differentiate yourselves? So when we started the company, video generation just was not possible. We started in December 2022. to. Ajit and I both were doing our PhDs at Berkeley. And, you know, he had worked on image generation. He had done the foundations of 3D generation. He had done this project called Dream Fusion while he was at Google Brain. It was the first model that could create 3D objects from text. It was remarkable. And it got so much excitement that like, I remember it was number one on Hacker News for like two days or something like it just massive traction.
2:39I think we hosted the website on our own academic servers ourselves and we racked up like tens like i think a 10 or 20 000 cloud bill in just two days just for the the landing page it was insane because there was so much excitement this was the moment when we were like video or just like creativity was really exciting but first time people were willing to you know um pay for it ultimately something that we just thought was like not possible because the ai technology was was growing so fast um and so So we've shipped three generations of this technology. We actually shipped the first image to video model on the market in January 2023.
3:15And yeah, I mean, it is a competitive field, but the way we see it, we are less than 1 % of the way there to the generative video future. The models today are rough. They make short videos and they're not perfect by no means. And they're hard to control, right? It's hard for artists to actually get good results out of them. There's a lot of lottery pulling is what we hear from our users about just the sea of technology and we were really focused on driving two things forward being the best at these in the market number one is motion having the best motion because we feel many of the products in the market today they they don't feel like real video they feel more like live live photos almost off your phone you know gifs and our model mochi one has the state of the art motion in the industry better than both the companies you mentioned right and number two is prompt adherence it does what you ask and really important because it's very hard to use these models in production or actual use cases if you can't control them right and emoji one is one of the first models that really sets the bar for prompt adherence uh prompt adherence and then the other big issue with these models is consistency um is that how how do you guys do on that has that been solved because i've been seeing some generated videos where the facial consistency is much better than in the earlier models.
4:40It's getting a lot better. We made the decision to open source this model. That was a huge step. I mean, there has been no good open model available on the market. And honestly, frankly, coming from academic roots in UC Berkeley, if there's no open model, it's just a game with a big incumbents. There's no potential for kind of startups, let alone academics to succeed. And coming back to your question about consistency, one of the things that open source enables is people to fine tune the model. They can take the video model and they can say, produce a Craig or a Parse model. Like if I wanted to make a Parse avatar or digital twin, I could do that.
5:19And so people are beginning to take this technology, fine tune the model, and then it gains character consistency. It gains the stability. like it will produce stable and consistent identities ultimately across many clips and that's something that's possible for the first time because of open source that hasn't really been possible in the prior era where you only could describe things with prompts right it's very difficult to get consistency there yeah uh well i have a lot of questions but let's start with uh building the model i mean that's if you're building it from scratch you're not building on top of an open source model.
5:59Where did you get the capital to train a model? How much capital did you need? And how long did it take to train? Yeah, so we spent almost close to a year working on our core motion engine in our model. It's a step breakthrough for the field. It's actually a fundamentally different architecture from how these models have been trained many times in the past in order to unlock motion. We talk about this in our open source release, but one of the key innovations was a new architecture. We call this like this somewhat technical term, ASYMDIT, or it's essentially an asymmetric architecture for video generation.
6:36It actually makes it more cost efficient to train and run. It enables us to actually reach that better quality motion. You know, I did my PhD in large language model scaling, but a lot of my PhD was actually focused on efficiency of training these expensive models. Because they are really expensive. That is abundantly clear. What we feel is it's really important to find steeper scaling laws is what we call this. Essentially what we say here is, you know, for the same amount of compute, just get more amount of oomph and scaling and parameters effectively within that same hardware footprint. You know, in terms of capital, we recently announced our$30 million Series A.
7:13And so that was, you know, that is a that is a large round. But I think compared to what many other companies have done in VideoGen, it's actually fairly conservative. And we are currently at the close to the top of the leaderboard. We're the number two model in the entire market on the artificial analysis leaderboard, which is, you know, third party evaluation by hundreds of thousands of independent raters, effectively, the general public. And we did that with way less capital than many of our peers. Well, can you talk about the initial model? you said you spent a year getting motion down, but when was the official model launched or made open source?
8:01How large is it and how did you train it? Yeah. So we just launched this October 22nd. And we released a 10 billion parameter model. So That's a relatively small model, but for video, that is one of the largest models on the market available today, especially in the movement. But it remains comfortable enough that I think what's really remarkable is we shipped it initially to run on a full DGX server. That's like$250 ,000 of hardware. And within four days, people compressed the model by a factor of 25x. So now it actually runs on a consumer graphics card, something you could go and buy$1 ,500.
8:43dollars. Yeah. And so it's really remarkable. Like everything here goes back to open source, but that's the power of open is that the community is taking this and they're finding ways to make it smaller, make it more customizable, make it faster, just improve its quality. You know, in terms of time, like, yes, it took about a year to develop this, but, you know, ultimately we think of our innovations in three key pillars. Number one is data. You know, ultimately these models are garbage in, garbage out. And one of the things we've been really remarkably proud of to work on is synthetic data pipelines for models.
9:16It's something new. This is really going to be very important for video generation in the future. But we actually find that it's often a faster, more efficient way to learn the rules of physics and nature, ultimately, and how the world operates through just simulation rather than just doing this directly through real data. It's really early. We're still innovating on this, but I think this is going to be hugely important for this field. Something that was never possible for language models, but now is something we can do with video. What do you use for the synthetic data generation? We've been exploring ways to kind of interplay between 3D generation and video generation, specifically because my co-founder has deep roots in the 3D generation field.
10:01And so, I think there's some things that are hard for video models to learn today. One of those is 3D consistency. You know, a test case here is if you walked around in a circle, you end up at the same place, right? And you'll see with video models today, there's actually a gap there. And so it's something that's a specific test case we can monitor. And because it's so specific, we actually find ways to augment it with synthetic data. Just learning also just like 3D structure for objects is also really interesting. To be clear, this is really early. It had a relatively small part to play in mochi one but it's something that's becoming more and more important for us as we go on yeah i mean the reason i ask about the synthetic data is a colleague of mine from the new york times named dale kim built a company was acquired by meta and i i'm not going to be able to remember it now but it's had the word dream in it um and then of course there is unreal engine.
11:02I mean, do you use third-party systems like that to generate the data to train Mochi1? We're testing this internally. It's really early. I think, you know, before I did my PhD, I was in self-driving. And I think what's remarkable there is the state of the art in self-driving is simulation. People train these things literally in 3D engines. Like I remember at our startup, like setting up a massive server cluster and running a bunch of instances of GTA, literally and having our self-driving agent drive around GTA for hundreds of thousands of hours. And remarkably, that actually helped to transfer to the real world.
11:41We're searching for what the right analog here is. I mean, in some ways, our models actually go beyond 3D rendering of the traditional 3D graphics pipeline. For example, one of those areas is like, I remember this was in June or May this year. I remember in an earlier version of the model, we had made a video of a dog swimming underwater. And so when we looked at this video, you could see the sun refracting through the surface of the water as it kind of reflected underneath and simulating how light refracts through like liquids underwater like that. That's actually really complex. Like 3D engines struggle to do this accurately in a physically consistent way.
12:20And it was something our model could do in less than a minute of GPU time. Like this would take thousands of hours of rendering to produce. So I think it's really interesting to see where video generation models have actually deeper physics understanding than 3D engines versus where 3D engines still have additional capabilities over video generation. Yeah. So you built the model from scratch. You trained it. I don't know if you don't want to answer, but I don't yet understand how much it costs to train. and then you open source it and people start iterating. Where are you now in the process? Yeah, so we are continuing to grow the open source community, but we're also continuing to develop our next model.
13:09I mean, in terms of resources to train this, to train Mochi One, it took something roughly on the order of a thousand GPUs, right? So it's a fairly substantial footprint of resources. But, you know, I've heard estimates that, you know, there was a blog I'd seen that estimated Sora took about 10 ,000 equivalent GPUs to train. And so if you think about that, we are order of magnitude more efficient for something beginning approach, you know, that level of Sora quality, I think. You know, we're not quite there, but, you know, I think as I've shared on the leaderboard, we're at the top of the market, right?
13:43And so I think efficiency is going to be a huge, huge thing here, because if you think about these video engines, They're just incredibly resource intensive to train. Right. I mean, what's also interesting, though, is serving these models, which is a huge challenge. One of the things we spent a ton of effort on was reducing the latency to generate these videos. I think it's a really common experience using VideoGen as people queue up a video and then it takes 10, 20, 30 minutes to sometimes finish. Right. It's not an easy experience. And one of the things we did is that we have the fastest time to first pixel in the market for a model.
14:17you will start to see videos rendering off of the GPU within 10 seconds. And so I think not only is quality important, but also just sheer speed. It's very hard to use creative tools if there's high latency. And if we can reduce that iteration speed, that makes the creative process so much faster. Yeah. And the model is available to people through the cloud or are people downloading the weights and building their own instances? How does that work? People are using it in three key ways. We have a hosted product which we make available. This is freemium. Anybody can use it. But open source is really interesting.
14:59With the open model, people are hosting APIs for people to build access to this technology. We work closely with our inference partners to provide high-quality APIs for people to run it. But what was never possible with any video generation company was for you to be able to take the model and run it locally and really like touch and feel it, like understand how it works, pull it apart and look at the internal mechanisms. What open source means is people are downloading this thing. They're finding ways to shove it into the smaller, smaller hardware. I saw somebody working on a port of the models were running on a MacBook, literally, right?
15:33It's just incredible to see the progress there. And alongside that, people are finding ways to kind of hack on it and find new capabilities that you just could never do with an API or with a product. One example here is, you know, I had seen a community member launch video to video editing on our open source model, something that we never we never trained it to do this. But they had found by taking the open model and hacking on it, they could take a video and they could take a video of a person and say, add a hat and it would actually edit in the full hat on their head. And it looks realistic.
16:04It looks like they literally had that in their life. And it was remarkable to see that happen in the open source community with just no additional fine tuning or anything needed. This was just something that somebody could do with one person and, you know, small set of computational resources. Right. But it is available through API or not? Yes, yes. We have third party API partners who are working closely with are offering reference implementation models. So you can call in and build products. And we've seen some really interesting early product use cases built around the API as well. Yeah. Well, you mentioned third-party or inference partners.
16:46What do you mean inference partners? So one of the providers on Thal.ai, RunPod, and Modo Labs are three providers who have been hosting OSI One for their customers. And so people can go there and they provide the GPUs and the computational resources to integrate with and they serve the model. Yeah. In this whole video generation space, so you've got a handful of leaders. Do you think there will be a consolidation around one model or do you think that there are a diverse enough range of use cases that the models will become specialized in different areas. I'm just curious how you see that moving forward.
17:41I think there will be millions of models for video. I think video is just too diverse for there to be one model. You know, with language models, there's only a few reference models here at the leading edge. And I think what's interesting is,
17:58you know, actually, backing up, The way I see this is foundation models are raw materials. In their raw form, it's like crude oil. Like if you or I went and we started drilling and we found crude oil, like you could shower in it, you could brush your teeth. Like, I don't know, like it's not inherently useful. You have to refine it and customize it and distill it to make rocket fuel. And our model is very much the same thing. The raw model itself is powerful. It is valuable. But the best use case are going to be people who take it and find some interesting use case. Like it might be animation or like anime, or it could be like marketing e-commerce.
18:35It could be like, you know, some kind of new storytelling application, right? It might be some style or aesthetic and they fine tune it and they produce the best model in the world for that specific use case. And I think that's how this technology comes to market. Video is just too diverse. Like creativity is unlimited, I think. And in this way, like you can't just have one model ultimately that just is the best of the world in every single day, right? Yeah. and i were as i recall looking at your site you guys are getting close to six seconds uh for a video right yeah so we support 5.4 seconds today that's right and six sex and six seconds is kind of a a benchmark in filmmaking i mean a lot of scenes are six seconds long and you You just, you know, and then you have a longer unbroken shot.
19:33But a lot of it's six seconds or that area. How are people using it today? And what has to happen for the generation to get longer? I mean, you know, up to a minute or two minutes or something. Yeah. Yeah. You know, what's interesting about this is actually the average shot length in film has been decreasing over time. It's this really interesting property that's been happening. And I think that doesn't mean that models should be restricted to shorter and shorter windows. Like at the end of the day, today, we do single shot. Moji One is a single shot model. I think there is a moment where we go across and I think that's really the next big evolution.
20:18And so there you ask the model to produce, let's say, a 20 second end produced, let's say, commercial or something. And there the shot transitions are just decided by the AI based on where he thinks the best place to put them is going to be. The biggest challenge with length, though, is cost. Like the cost of these longer video generations actually goes up what we call super linearly. So doubling it costs more than two times the cost running it. And so, you know, I think I see this technology very much as like crawl, walk, run. Today, this is the stage where, you know, we actually chose, it's funny enough, five points to four seconds because it was just over the minimum length of video you could upload to TikTok.
21:03So for many of our peers, like you couldn't even upload their videos to TikTok because it was just below that threshold. And so we but we wanted to make it big enough that it was useful, but not too big that it wouldn't run on real consumer and hardware. If we made it 10 seconds or 20 seconds or 30 seconds, you would need these big distributed GPU clusters to run it. But 5.4 seconds is the sweet spot that it fits, again, on a graphics card that a consumer can buy. Yeah. And what about consistency between shots? I mean, I understand that you've solved sort of facial consistency as someone's moving or turning or something.
21:45But if you have the model, and maybe you can explain the workflow. If I wanted to generate a model, I mean a shot of me laughing, and then I wanted to have a second shot of me frowning,
22:07what's the process first of all to get the model to generate an avatar that looks like me and then to ensure that the second shot is matching the first shot yeah I mean since it's open source I think the step here is you would collect you know some reference images and video teach the model what is the concept you're trying to teach like what what does what is parse you know who's correct You would represent that and you would fine tune the model. It would become really adept at creating that. And at that point, you would be able to consistently produce your identity. I think this character consistency is really important for creative workflows.
22:48Ultimately, what is a story without actors? What is a movie without people? We were really focused on the first step again for motion was actually about producing realistic looking humans that don't look like claymation dolls. I think this was the first big challenge. A lot of AI video, people kind of just look like puppets almost. They don't feel real. So that was the first threshold. We crossed that with Mochi 1. I think today already it's possible to do this fine tuning in order to produce consistent characters, but it takes resources costly. You need some technical know-how to fine tune the model today.
23:23But it wasn't even possible to do this with closed source models and platforms before. You know, I think in the future, like my hope is that you can get to like few shot prompting. So you can say just like an LLM, you can give it an example and say, here's this like text. I want you to write in the style. You know, it should be possible. I think eventually for these models to say, here's here's this person and this is the character, right? Like this, this, this is the character I want you to create. And here's their backstory. And then it just can consistently produce shots of that person. We're not there yet.
23:54the technology needs this fine-tuning step still to produce consistency. But I think this is where the technology has to head to make it really usable and provide good iteration for creatives who aren't really going to be able to fine-tune this technology easily. Yeah. What's the difference between what you guys are doing and companies like
24:22Synthasia, is that it? Synthasia. Yeah. Where you can provide a lot of reference video of yourself and they can create an avatar and then you can write scripts and the avatar will speak the scripts, which sounds like it's the same technology in a way. But what's going on there? What's the difference? So these are actually two fundamentally different approaches to video generation. I would say there is a category of avatar generation, like you're describing, where they generally take an image or they'll take like a short video. And they'll manipulate like the lips and the eyes and your facial features to make it actually say something and match the audio.
25:09And I think in limited scenarios, this can feel realistic. But if you think about it, that is as far as the technology will go. it will. And now if you want to say, have the person pick up an apple and take a bite out of it, right? Like technology is not going to be capable of doing that fundamentally. It's much more cost effective and you get much more control. But I think with an architecture like ours, what we've chosen to do is we've chosen to build what we call almost like a world simulator. It's a physics engine that simulates, you know, inertia, mass, you know, it simulates optics, viscosity, it simulates optics to viscous liquids, right?
25:46Like it's everything, the cross product of all the physical relations with the world. And so we started with really simple scenarios, but at this point, like it can produce humans. What it won't do is it won't produce dialogue yet. Humans, like they'll kind of look around, maybe they'll speak, but there's no audio, right? And it's not going to be matched to dialogue. But, you know, I believe as these models continue to scale, they'll gain the capability to have better ability to actually, synthesize consistent dialogue. If you want the person to say what it's going to say, it's going to produce it, but it won't just be able to say that.
26:18It'll also be able to do things like pick up an apple, eat it, drive a car, do anything a human would be able to do. And I think this is fundamentally just more powerful and will produce more interesting storytelling in the long term. Yeah. I haven't spoken to Jan LeCun for over a year, but I was following his work pretty closely, and I spoke to him several times when he was working on world models and using video. And his experiments were pretty crude compared to the video generation models today. But how foundational was his work on what you guys do? Or was he going down one alley and you guys went down another?
Read the full transcript
27:16You know, I think it's interesting as language models learn through regurgitation, I almost would say. What do I mean by that? everything a language model learns has to essentially be the output from some other person, right? Like it's somebody talking about a subject, about what they perceive about the world, and that's how these models learn. Now, I think what Yan was trying to push at was ultimately these models learn by kind of operating in the world, right? By actually reasoning. There's no indirection step here, right? Like language models have. And in this way, this is directly learning from the laws of nature, right?
27:55the physics and the nature of reality. And I think that's a really impactful idea. And I think that's ultimately why video generation is a really fundamental capability for intelligence. It represents kind of this. It's not second degree, it's first degree learning from the real nature of what reality is. And I think that's really powerful. So when we think about our models, I think ultimately, like it's left brain, right brain. It's going to be a forever debate, which one's going to be more powerful. but I actually think video generation has the capability to ultimately be one of the most powerful forms of intelligence for this reason.
28:30I think the world model concept is just inherently very powerful, I think, ultimately, because of this first degree versus second degree form of learning. Yeah, but when you're training,
28:47and you know I'm a journalist I'm not an expert but I remember Jan's videos very quickly would blur because the model there were so many
29:06possible futures and it would present all of the futures How do you get a model to follow one thread of possible futures? Because isn't that what you're doing when you're generating video? Absolutely. So we leverage the diffusion model. And so like I shared, my co-founder, Ajay, worked on the DDPM papers, one of the first architectures to kind of realize this diffusion concept of practice. But the idea here is, you know, you're modeling essentially a process to go from noise, you sample some random noise that represents just like some, some concept and you slowly denoise that into a real video ultimately.
29:49And, you know, it's not like a language model where a language model produces one word at a time or one token. Our models actually predict all frames at once. So our models work by concurrently denoising all frames at once. And so what that means is, because you're starting with noise, it kind of removes this problem of this blurring that you described. Like there's multiple possible futures. Like the noise identifies the final endpoint precisely, right, for the model. And so it removes this ambiguity. And it's actually really interesting to bring this up because, you know, when I worked in self-driving, you would have this problem too.
30:23You want to simulate the future. You need the vehicles to understand and predict what, for example, pedestrian will do. And one of the big challenges is people are kind of, you would call them a multimodal in this way. And that what I mean by that is that they do multiple things. They could cross the street one way or they could go another way or they could just go jump or they could go running. Like humans are inherently unpredictable. And so what you would find is it was really important to reason about the probability distribution and then from there you're kind of saying, here's a specific path that they will take.
30:54Right. Like, and I think that it's interesting in that generative models in many ways are reasoning about the probability distribution. I think the power of say a model of diffusion model is that you know it has the ability to produce you know sharp unblurry very realistic samples by kind of following this like you know is a procedure yeah uh the um uh so when you say that you're producing all frames at once as kind of a continuous continuum, right? How does the model break it into frames or does it? So the model follows a setup where it has essentially a compression and decompression step. So it takes video.
31:40Video is huge. It has a huge amount of redundancy. If you look at video, I mean, the frames across multiple sequences basically are very similar. And then on top of that, so that that's temporally and even spatially, like you'll have like just white backgrounds like my wall here right and that there's no information there and so like the first step is actually just to compress that information so we go through a step where we take video um and currently it's about that like you know that you know five and a half second context window and we compress it down 100x both spatially and temporally and it's not just one dimension like we're not only doing spatial compression we do spatial temporal compression that's really important for the model and that's what helps it reason about both space in time.
32:23And once it's in this compressed space, I mean, there's not really any delineation here between different frames, if that makes sense. It's all just one common shared space. And so in that space, we denoise the predictions. We go through this iterative denoise process. So it goes from pure noise and iteratively progresses towards a video that you actually see. And then we decompress that into the final video. And it's with this kind of three-stage architecture that we make it computation efficient to do video generation um otherwise the cost would be incomprehensible right you would have you know more than 100 frames right in that short sequence yeah but do you need to break it into frames in order for it to be perceived by humans i mean Otherwise, you would get kind of, it seems to me, kind of a blur through time.
33:22Yeah, so it ultimately comes back to the frames, the pixel space, we call it, but the frame space, which humans watch it. There are separate dedicated frames, ultimately. And so I think there are two ways these models work. I mean, previously there were pixel space only, like when Ajay started working on this area, you would just do it directly in the pixel space. But this like compression space with this like latent diffusion set up is essentially what this is technically called, has proven to be really important for computational reasons. You know, I think fundamentally, I mean, I think that what's interesting is that there are these two architectures for intelligence.
34:02There's the left brain uses the autoregressive language model, the chat GPT, the code one token at a time. On the other side, the right brain uses the fusion, which is kind of this all at once, you know, you don't decode, you know, temporally one thing at a time, you just denoise it, right? And if you use our product, you'll actually see within 10 seconds, you get like a blurry video, which has noise in it. And so it gets less and less blurry until it is really, really sharp and crisp. And so that's visualizing the step. And these are the two architectures. And it's really interesting that they function completely differently at their core between kind of creative modalities and text and and you you you put them together is that no so we do the diffusion just like all at once okay yeah so it's just interesting that there's this left brain there's this right brain of intelligence and and the architecture that works really well for each one is completely different i think it's actually yeah yeah uh and And then can you adjust the frame rate in your output?
35:04So we do 30 FPS today. You know, I think it's a really interesting question. Some models do 24 FPS because that's what film has done. 30 FPS is very natural. And I think in the future, maybe you'll do even higher FPS. But like we do this fixed frame rate because I think 30 FPS represents kind of like, it's kind of what you get out of your phone, right? It's like the raw asset. If you were to just shoot a video, like that's what it's going to come out as. And our goal was to kind of match that technically, do no less. Yeah. But you could increase the frame rate. I mean, it's an engineering problem.
35:39Yeah, it's really a cost question here. So, like, we could do it. I think, you know, in the setting where people are consuming media, like film, it's almost like interesting in that, like, we've had requests to go less. Do you support 24 exclusively? Because it's more cinematic, right? You know, I think it's interesting. where do video models go like in the world model context like eventually these things end up looking more of a video games right like you can hook a eventually you might be able to hook up a joystick to that video model and actually drive a car in it in that world i think that frame rate's going to be really valuable and really important you're going to care a lot about making that even feel even more real time right um wow yeah yeah that's incredible and then if that were the case you would be generating the environment as you move through time, right?
36:27Yeah, absolutely. I mean, you could imagine exploring new worlds completely that are fully synthetic, right? And it's like this question, like, you know, today our video model can do many things that a 3D engine couldn't do, like complex hair for physics, like I said, like this optics, like some of these things, like a 3D engine could never simulate. And so it's like a whole new space to explore creativity. And I think if you give people the ability to interact in it, I mean, that's going to be really next level. I mean, I think we've talked about a few different directions. Like there's long video generation.
37:00I think that's going to be really important. Controllability, steerability, and consistency is going to be really important. And interactivity, I think, is like this third dimension that I think the industry is going to push on. By open sourcing this, there are teams actively exploring all three of these on top of Mochi 1 as we speak. And I think that's what's really remarkable. And in this way, I think a lot about open innovation versus closed. And I think just the creativity of the open source space is just unmatched. Nothing can do that. Do you think you could share your screen and go to the platform and just take us through creating a five second video?
37:43You offer this playground. Yeah, you can go to genmo.ai. It's a free place for people to start to test the technology. For us, it was really important to get this in the hands of real people. So today we have more than 2 million registered users on this product across 40 plus countries who use this to make video. And so you can write a prompt here.
38:08And I'll show some other examples from the community feed as well. Sure. Yeah, you'll go ahead and cue this.
38:18And so behind the scenes, we're seeing huge load, again, as I shared, people's appetite for video generation is voracious right now.
38:30And so there it's scheduled, and it's going to start rendering. So just like you might render out a 3D graphic or something, our model is doing something similar but it's doing the demoing. And so we'll see this start to stream back just very shortly there. and uh what i i wasn't paying attention what was the prompt so here it's a it's a it's oh i just took this from the community feed i mean i'll show some other examples but this is like an infinite zoom uh on someone's eye right i see okay um and so that's going to take a second to generate here um i'll go ahead and show like what this ended up being i took this from just the community.
39:08But this was the resulting video. Yeah. And it can do abstract styles too, which I think is really interesting. These are just like actual user creations that I'm showing here.
39:24But when we talk about motion, this was really the goal.
39:36can you can you oh yeah that's cool can you click on the room uh that yeah at the uh wow that's amazing so many videos just kind of would show a close-up shot of the mouse and it would just be like not really doing anything so this is what i think is remarkable this model is the motion is fluid and it's a real scene. And that's a one shot off one prompt? Yes, yes, in one prompt. Do you have the prompt for that one? Yeah, it's a video of a rat scurrying through a cluttered painter's room. Wow, that is amazing. This is one aspect too with prompt adherence is like, you'll see each detail is reflected in the final video.
40:18And that's like very important, I think. Let's see if there's another one. I mean, even here, I mean, this is a little bit more abstract prompt, so it's harder to see. But we can see that it's beginning to render out. You can kind of actually see this in this preview here. Like, it's slightly blurry here. And so it's like, this is the process I'm talking about. It's like, the video is like actually just getting denoised until it finalizes into a full resolution video. Yeah. Yeah. If you wanted to give it a reference image, like somebody's face, could it work off that or does it not work that way?
41:10Yeah, we're working on image to video support right now. So it will come soon. So you can provide an image upload and be able to actually convert that into a video. There's some other cool stuff I'd love to share too, actually. Like this is, give me one second. So this is really remarkable. So this came from our community, just completely unprompted, and they developed video editing capabilities. So we do video generation, but Moji One also now supports video editing. And so here, given an input video, which is this top left video, you can provide multiple different edits. In this case, you can ask it to put a hat on the person's head and it happens.
41:48And I think what's remarkable is to make this with a traditional VFX pipeline, this would take a lot of work. You'd have to keep pointing in all the keyframes and everything. It's a lot of work. But it just seamlessly is able to add this hat. I mean, it's completely open source. Wow. And this is what's this Comfy UI is a separate node. it's not something you access on Mochi's page. Not yet. So I think it's hard to keep up with the speed of the open source community. I think they're moving fast. Yeah. So let me ask you, how do you guys make money if this is all open source? You know, I think we're really early in this video generation era.
42:45We're 1 % of the way there. Reality is it's a really high bar. So I think it's really important for people to be able to get started with the technology and open sourcing this model. The first model was really critical to creating the system. Our goal is to support them, though. And I think people will build incredible applications and workflows and engines and things on top of this. And so ultimately, these people will need production environments to run this. They'll need the ability to fine tune and customize, reliably actually build these systems. and we'll be there to support them as they scale.
43:16And I think that creates a really well-structured and incentive-aligned path for this technology. We were facing a world where I think it's unquestionable the big AI giants and incumbents took over the left brain of AI. And I thought it would be tragic for that to happen to the right brain, ultimately. And so open was really critical for it to have a fighting chance. Without that, I think the community would just be completely beholden to your open AIs and other large incumbent AI providers. But I do believe that by doing this, this is going to net grow the creative pie 10 or 100x, right? And create this huge Cambrian explosion of creativity.
44:00Yeah. You guys are open. Is Runway open? Runway is closed source. So it's a closed source product and closed source model. And they have a very good product. I think the interesting thing, though, is I showed just the open community is moving so fast that I think this wins over close. Like if you look at the Internet, right, like where would you be if Microsoft had won the Internet Wars? Right. We would be stuck in some kind of like closed, stagnant, you know, static ecosystem. Right. It'd be like, you know, Internet Explorer, Enterprise Edition everywhere. Like you wouldn't have the beauty of the Internet.
44:40So open is what enabled that. I think that same movement had to happen to video. Yeah, yeah. This is kind of off topic, but the metaverse has not been populated because of cost. And is this relevant at all to that kind of video generation, or is that just computationally much heavier? I think absolutely. I think at some point, like when these video models get strong enough, their 3D reasoning gets stronger. Like you could build out massive universes that you could just join in. I think that would be incredibly engaging. When I talk about like interactivity as one of those long-term directions for this technology, I think Metaverse connects to that.
45:30It's one thing for people to be able to create these artistic and creative assets. I think the next step is let them enter these virtual worlds entirely and just participate in them. The technology is not there yet, but I don't see any reason it's not going to get there. It's just too compelling. Yeah. And you guys are still doing research within the company. First and foremost, we're a research lab. So we're working to develop this technology. Like I shared, I think it's just so early in this technology's development. It's really important to make sure that the right foundation is built and technology is scalable.
46:04Right. I mean, you know, it feels like, I mean, I think some of Moji One's creations feel very realistic. But like, I just have to have that humility. You know, like there is a huge gap to close still. The technology has to advance. I see this as a lot like GPT-1. It's like the first innings. And, you know, I think when people saw that tech early on, like it was just hard to imagine where it could possibly go. Yeah. Yeah. And are you affiliated with any educational institution that has resources or you guys have built all your own, you know, or you're working with your own resources? We have our own research team, but I think what's wonderful about open source is we're collaborating with close to a dozen universities at this point.
46:52And that's growing with faculty and researchers to extend the model. So things like making it cheaper and smaller to run. There's a team that's beginning to investigate how they can get this to even run on mobile devices. I think that's remarkable. I think that's actually something that's really cutting edge. And, you know, that's the power of open. And, you know, we can have these collaborations with so many different academics that just never would be able to participate in this discourse without an open thought. Yeah. Wow. That's really amazing. I'm going to go play with it. Maybe I'll put one of my own generated videos in the final podcast.
47:30But Paris, I'm delighted that we got to talk and I hope we stay in touch. I'd like to see where this goes. Is there anything you want to say that you haven't said? No, I think that's about it. We covered the key points. Yeah, this is a great conversation.
From the publisher
In this episode of the Eye on AI podcast, Paras Jain, CEO and Co-founder of Genmo, joins Craig Smith to explore the cutting-edge world of AI-driven video generation, the open-source revolution, and the future of creative storytelling.
Paras shares the story behind Genmo, a company at the forefront of advancing video generation technologies, and their groundbreaking model, Mochi One. With a focus on motion quality and prompt adherence, Genmo is redefining what's possible in generative AI for video, offering unmatched precision and creative possibilities.
We delve into the innovative approach behind Mochi One, including its state-of-the-art architecture, which enables fast, high-quality video generation. Paras explains how Genmo’s commitment to open source empowers developers and researchers worldwide, fostering rapid advancements, customization, and the creation of new tools, like video-to-video editing.
The conversation touches on key themes such as scalability, synthetic data pipelines, and the transformative potential of AI in creating immersive virtual worlds. Paras also explores how Genmo is bridging the gap between cutting-edge AI and practical applications, from TikTok-ready videos to future possibilities like interactive environments and real-time video game-like experiences.
Discover how Paras and his team are shaping the future of video creation, blending art, science, and open collaboration to push the boundaries of generative AI.
Don’t forget to like, subscribe, and hit the notification bell for more insightful conversations on AI, technology, and innovation!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Introduction to Paras Jain and Genmo
(01:45) Video generation with Mochi One
(04:41) Open-source AI in video generation
(06:08) Building Mochi One
(12:03) Simulating complex physics in video models
(14:35) Reducing latency: Fast video generation at scale
(20:34) Tackling cost challenges for longer videos
(23:14) Character consistency in AI-generated videos
(27:17) Why video models represent intelligence's next frontier
(30:18) How diffusion models create sharp, realistic videos
(34:02) Visualizing the denoising process in video generation
(39:36) Exploring user-generated video creations with Mochi One
(42:05) Monetizing open-source AI for video generation
(45:47) Video generation's potential in the metaverse
(47:19) Collaborating with universities to advance AI
(48:39) The future of generative AI




