In short
Eye On A.I. Podcast Episode #172: Cristóbal Valenzuela - Can AI Revolutionize How We Create Art?
Episode Overview In this episode, host Craig S. Smith interviews Cris Valenzuela, co-founder and CEO of Runway, an applied AI research company that is significantly shaping the future of art, entertainment, and human creativity through the use of AI technologies. The conversation explores the origins of Runway, advances in generative models like stable diffusion, and the impact of AI on creative workflows in various fields, including filmmaking and music production.
Key Discussions Background on Runway and Cris Valenzuela
- Origin of Runway: Founded in 2018, Runway emerged from research at NYU focused on training large models for creative professionals.
- Cris Valenzuela's Journey: Valenzuela is originally from Chile and moved to New York to study at NYU's Tisch School of the Arts. He has a background in art and computer science.
Generative Models and Their Applications
- Diffusion Models:
- Explains the mechanics of diffusion models, which add noise to images and train networks to remove that noise to generate images.
- Emphasizes the recent breakthroughs in image and video generation quality due to advancements in techniques and increased computational power.
- Stable Diffusion:
- Developed through a collaboration with LMU Munich, this model has become one of the most widely used open-source models for image generation.
- Runway has built tools on top of this foundation to empower creatives.
Video Generation Technology
- Challenges: Achieving temporal consistency in video generation remains a significant technical hurdle. Models must maintain object consistency across frames.
- Training Requirements: Video generation requires extensive computational resources, often involving thousands of GPUs.
User Community and Impact
- Runway supports a diverse user base, from Hollywood filmmakers to independent creators, and aims to democratize access to powerful AI tools.
- A notable feature is Runway’s film festival showcasing AI-generated content and the community’s creativity.
Product and Pricing Model
- Runway operates on a subscription model, allowing users access to tools for a fee. Customized model training for specific user needs is also offered.
Implications of AI in Art and Media
- Technological Impact: The rise of AI tools is set to redefine artistic expression, potentially giving birth to new art forms.
- Cultural Evolution: Valenzuela discusses the inevitable transformation of media and the types of expressions that will emerge as technology evolves.
Future Directions
- Continued development of Runway's models and the potential integration of these tools into AR and VR environments.
- The company is poised to grow alongside the rapidly changing landscape of media, with an eye toward enhancing both artistic expression and accessibility for creators.
Conclusion In this enlightening episode, Valenzuela shares his insights on how AI is revolutionizing the creative process, fostering a new era of artistic expression. The core message emphasizes the collaboration between technology and creativity, underscoring that while AI can enhance artistic workflows, it is ultimately the intent and vision of human creators that drive innovation.
Key Quotes
- "Art has always been about technology...these algorithms are another form of technology."
- "I think battling or putting the artists against a technology is not really the case."
Additional Information
- For listeners interested in the intersection of AI and art, the episode provides a thoughtful exploration of creative possibilities and the implications of these advancements on future artistic endeavors.
- The podcast also encourages ratings on Apple Podcast and Spotify to foster community engagement.
---
Stay Updated
- Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
---
This structured summary captures the essence of the podcast episode while highlighting significant insights and discussions about AI's role in transforming creative industries.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Particularly over the last four to three years, we've seen an explosion in quality, consistency and overall resolution of image and video models. And it's been because of a few key advancements in techniques, but also more compute and more data has become accessible and available. So around three years ago, we started working on this idea of creating new kinds of models for image generation. And we do collaborate sometimes with other research institutions. So we did a collaboration with the University of LMU Munich in Germany. And that paper that we published was called Latent Diffusion, that introduced a new technique for generating images on a latent space and conditioned those images to a variety of different inputs.
0:45Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I speak with Chris Valenzuela, co-founder of Runway and one of the creators of Stable Diffusion, the open source text-to-image model. We talked about advances in text-to-video models and what that means for the future. But I had this conversation literally the day before OpenAI introduced Sora, its incredible text-to-video model, so I wasn't able to speak to Chris about that. He did talk about the advancements in techniques, increased compute power, and data accessibility that have made applications like Sora and Runway possible. Join us as we explore the intersection of art, technology, and the future of AI-driven creativity.
1:40I hope you'll find the conversation as fascinating as I did. Can you start by introducing yourself, give your background, particularly your educational background, how you started Runway, how you built Stable Diffusion. Yeah, totally. So I'm Chris Valenzuela. I'm the co-founder and the CEO of Runway. I'm from Chile originally, and I moved to New York seven years ago to study at NYU, where Runway was our research work with my co-founders. We were interested in finding ways of training and creating large models for artists and creatives. And we started the company around 2018, so almost five years ago.
2:27And since then, we've published and made great progress on making sure models are accessible and empowering for creatives around the world. Filmmakers, designers, musicians that use runway to make anything from future films to short films. This is something I've been very passionate and deeply care about for almost now a decade. I've been working on media and research for a long time. So still on the way to go, but excited to see how we're changing and how many people are using Ranoi these days. Did you study under Jan LeCun at NYU? No, I went to art school at NYU, to Tisch. Oh, really? Yeah, I have a son that graduated from Tisch.
3:13Oh, nice. So Stable to Fusion, can you talk about how that came about? There are a few projects that kind of emerged around the same time, and I've never slowed down to figure out how they're related or not related. You know, there's Midjourney and Dali and, you know, various other models. So can you talk about how that idea started and what was happening in the underlying research that allowed all of these to emerge kind of in the same time frame? Yeah, absolutely. So a lot of the research and the things we're seeing these days on AI are not overnight success. These are things that has been on the works for years and sometimes decades.
4:13Everyone is building on the shoulders of giants and building on other work and open source research that has been published for a lot of years now. I think particularly over the last four to three years, we've seen an explosion in quality, consistency and overall resolution of image and video models. And it's been because of a few key advancements in techniques, but also more compute and more data has become accessible and available. So around three years ago, we started working on this idea of creating new kinds of models for image generation. And we do collaborate sometimes with other research institutions.
4:58So we did a collaboration with the University of LMU Munich in Germany. And that paper that we published was called Latent Diffusion that introduced a new technique for generating images on a latent space and conditioned those images to a variety of different inputs. That model and and that set of research eventually became Stable Deficient, which is just a larger version of the original latent diffusion model. Stable Deficient was an open source research model that we've released, and it's been one of the most, I guess, used and widely accessible models these days for image generation. There are other techniques also generating kind of like great images and great videos.
5:40And so we're also working towards making sure that we continue to just push the boundaries of how you can create and the kind of outputs you can make it up these models. Can you talk about diffusion models first of all? I've never actually talked to anybody about what a diffusion model is and how does it relate to transformer models? It's a different research stream although sometimes they overlap on some ideas and a lot of the improvements on video and image quality this day comes from actually transformers-based video generation and image generation models. But diffusion comes originally from a few concepts from like physics and thermodynamics.
6:20And the idea is basically to train a model that's, first of all, adding noise to an image. You have an image, a set of specific like references of pixels. You start adding noise to that, and then you train a network on being able to remove the noise from that image. And you do that millions of times, and the model gets really good at understanding then how to start from random noise, just a bunch of different pixels, and convert an image or create an image or generate an image out of that. It's a process that's been becoming more, I guess, widely used these days because the results have proven to be much better than previous techniques.
7:03Previous techniques were much more like they were using things like guns that were proven to be good. But the thing with guns is they're expensive to train and to use. And they're not as customizable and easy to move around or use as the diffusion models are. But again, I think the diffusion perhaps won't be like the last researchers we use to improve the quality of models. And we already have seen this where transformers have proven to sometimes be better than diffusion for some tasks. Yeah. Is DALI a diffusion model? I think so. DALI is a model that hasn't been released and open source, so it's hard to look into underneath the hood to understand how the model was strained and created.
7:48But you can, I guess, a safe assumption is that there's a diffusion part of generating an image in something like DALI for sure. There might be also a transformer kind of architecture at some point in that model as well. And mid-journey is also a diffusion model? We don't know also either. Everyone has different techniques. Those are different products from different companies and everyone can have different approaches to how to solve it. You can have educated guesses and time to approach if it's a diffusion model, a transference model, but they might be using some technique here and there that might be diffusion inspired or diffusion based.
8:24Are there any other open source models for image generation like stable diffusion? Yeah, there are a few models that the open source community has been putting out. I think there's definitely a lot of interest in trying to improve the quality of open source models. So yeah, there's a few models from different teams and different initiatives that are trying to... Sometimes, you know what happens? There's a lot of not just pre-trained models or baseline models, but also fine-tuned models. You take an existing, for example, stable efficient model and you tweak some parameters, or you fine tune on a particular data and that model gets released.
9:03And so there's actually like marketplaces where you can use some of those pre-trained and fine-tuned models. And that's becoming a widely kind of like a successful effort in open source and community building because it doesn't require the amount of resources and compute and time that creating the first baseline or pre-trained models in the first time, in the first case, like required so it's much easier for someone with like not a lot of resources or maybe just a few gpus to train very uh impactful or great models in terms of quality stable diffusion is the largest open source community in this space is that right so stable diffusion is an open source model um it's it's uh just a research and a paper and sort of weights and code that's been released is people have taken that model and have kind of like tinkered, modified, inspired by it.
9:57No one really has any particular saying on how the community grows. It's more of like the community just evolved on its own and people are building on top of that, for sure. That's how open source works. And that's how open source works for any project these days. And that's kind of the idea. You can build on top of each other and basically find and discover new things. Are there other models out there that are open source that have licenses available? I'm just trying to understand the landscape. There are the proprietary models like DALI and Midjourney, and then there's Stable Diffusion, which is open source, and then these licensees that have fine-tuned models.
10:42Stable Diffusion is an open source project that a few researchers, including Runway, collaborated on. That's an open source model. Runway creates and builds their own models. We don't have a membership. We have nothing to do with any other company. We're just building on top of Runway's like research. And so we have our own preparatory models. We have Gen 1 and Gen 2, which are video models, for example. Those video models are not based on any previous architecture or open source models. They're based on Runway's own research. And those are like set-of-the-art video models, for example, that are able to generate up to like 18 seconds of video.
11:18And those are in the same way as you were describing DALI or Mid Journey, those are models that are running on their own. There's an open source community of people who have taken some of the research we've done, but also some of the research other companies have done and have used that to build other companies or projects. And I think that's kind of the great part of research that you can have both ends, closed source and open source. And to be honest, I don't think there's a there's a peak where you have to be 100 percent of one or the other. we've been in the middle. Sometimes we open-source stuff, sometimes we don't.
11:50But I don't think of it as a binary state where you either open-source everything or you don't open-source anything. I think for us, what has worked really well in terms of research and community building has been open-sourcing some things like we did with some versions of table diffusion and latent diffusion and not open-sourcing others like Gen 1 and Gen 2 that have not been open source. And so it's a different strategy. And I guess we've chosen to be kind of like in the middle. Right. And those, the two versions of Runway that you're talking about, those are built from scratch or Runway's original product before you got into video.
12:30Was that built on top of the open source model? No, we build, yeah, we build all of our products and research in-house. We have our research company. We have a research team that's been working on these problems now for a few years. And so we build on our own kind of like stack. We don't use third party models. We don't use like APIs. We kind of like just work on the core fundamentals of our product and our research internally. And so I think like 85 percent of our team these days is just research engineering, advancing the training of these models and also the user, the use cases and the fine tuning of those as well.
13:06Okay, okay, that's clear. So what I really wanted to talk about is video generation, these models, and you're doing it. There's a company called Pika that's doing it. Are you both using similar architectures? Are they completely different? And then, of course, Google and OpenAI, everybody is coming up with their own version of this. So can you talk about the underlying research for video generation? Is it a completely new stream of research? Is it built on top of the diffusion architecture? It's definitely building on some of the findings that people have found around diffusion models, but also transformers and other techniques for generating images.
14:03Video presents a different problem, which is temporal consistency, how you make sure the models and the videos are consistent across different seconds of video generation. And so making sure that you can close that gap by creating new architectures has been kind of like the core fundamental research challenge these days. You want to make sure these models understand how the world works, the physics, interactions of objects, how characters move in a scene, and that's critical to make good quality videos. Every company might take their different approaches, so it's hard to define or tell that everyone is using one particular thing.
14:40This is more like you know, it's like manufacturing cars. People, you want to build different cars and different cars can have different logistics and factories and the cars that might come out of those factories might also be different based on what you're trying to like create in the first place and so we have a specific manufacturing facility and we make a specific version of our car but other companies might want to try different techniques to create those cars and also output different cars and so really for us it's about improving our own models and kind of like creating the the best car experience that we want for our customers and that kind of like metaphor analogy of cap manufacturing.
15:19But there's no one single way of building a car. It can be different depending on what you're trying to do. Yeah. And right now the videos generated are very short. I've seen longer form of videos, generated videos that are very hallucinatory, you know, dreamlike, make the things sort of morph into other things. Is that what happens if you run the model that you have long enough? Or, I mean, is there, why are you only doing, you know, a few seconds? So you can create up to like 18 seconds of video these days with Runway and probably and very soon longer videos. It all depends really on the artistic like direction they want to take it.
16:09people you might want to create more surreal videos or more uh realistic videos and that's kind of like up to the creator of how you want to take it um but really like it's not limited to time you can always extend that video for longer the way you create those videos though is you can start with chunks of four seconds and then eventually get too longer so you might have seen specific versions of those videos in particular moments in time but yeah you can create way longer videos if you want. What I'm talking about, these videos that people have produced that are very sort of psychedelic because there's no consistency, is that coming from, I mean, does Runway produce those?
16:58What is going on in those videos? Yeah, it's hard to tell without like seeing, I guess, watching the videos that you are are referencing too. There's a lot of different techniques to generating video. Some of them are more surreal, more psychedelic than others. Some people are actually creating just models to create psychedelic videos and that might not be using Runway at all. So there's no one singular way of creating videos in the first place. You can think about it really, a good analogy is to think about it as a camera. You can have different cameras to create different types of videos and scenes and it's up to you how you want to like create those videos using one particular camera and cameras might work differently depending on what you want to make.
17:42But with runway you can make it so it's very flexible so you can make abstract shapes and surreal objects all the way up to like very realistic and almost live action kind of like shots. Yeah and so consistency is the issue right when in as you as you... Temporal consistency, which is, I guess, how do you make sure objects pertain and maintain their consistency across different seconds of videos. And how do you do that? And can you talk about the underlying architecture of your model? You said it's kind of a mix of diffusion and transformers. And just to give you an idea of where I'm going, you know, I talked to this company, Wave AI.
18:28I'm sure you've heard of them. They're actually an autonomous driving system. But they have a world model. And that's why I asked about Jan LeCun because he's, you know, done a lot of research with video and world models. But they have a world model that they train, and then it can plan and generate scenes in a representation space. And then they can decode that into pixels, and it's remarkably consistent. I don't know if you've seen it. But are you working with world models? Are you doing something different? Yeah, we're doing the... So we have a current effort called General World Models that it's trying to solve, I guess, something on a very similar spectrum, which is not applying to self-driving cars.
19:30So the problem with self-driving cars is how do you have systems or cars that can navigate very complex physical situations to the point where they can drive themselves to avoid objects or kind of like situations that are arising on the street. The way you train that and the way you can do that from a world model's perspective is you don't decode the rules of the system. You have a system trying to figure out how the world works and how objects can fact predicting what's going to happen next. That's kind of the key aspect of it. If I throw a ball to the air, I can kind of predict, because perhaps I've seen enough balls going to the air that the ball at some point will start going down, right, because of gravity.
20:16For a model that's really hard to know, like the model hasn't been hard-coded to know how gravity works. But our belief and thesis is if you train models large enough and you give them enough data and with the right kind of parameters, you can actually try, the model will try to learn that. And so world models are models that not only look at particular videos or images or type of data, but they're actually trying to predict what's going to happen next. So transformer models are great for language because they basically do that. They predict the next token in a sequence of words. And by doing that, they are able to generate coherent language.
20:52In world models, you want to do something similar with video frames, for example. You have a video frame, one particular frame, you want to predict the next one, the next one, the next one. And you do that sequentially until the point you can understand what's going on, occlusion, movement, gravity of other objects as well. So our architecture behind the scenes is trying to leverage that idea. A big component of it is just large, very large models in the first part, but also training them in a way where all of these things do make sense for the model itself. Can you talk about the level of compute necessary to train these models and how large they are at this point and how large you want them to be?
21:33what the roadmap is? They're very large and they're very expensive. There's a lot of compute that's needed to train these models. I think one of the biggest expenses in research is just compute because these models are just very large and training them requires vast amounts of compute power. And so we're talking about thousands of GPUs sometimes or even more, tens of thousands, to make sure that the models perform and work at the kind of level that you want them to perform. I think that's only going to be, to be honest, like continue to increase over time. More computer will become more needed.
22:08At the same time though, there's always going to be optimizations and tricks that both chip manufacturers but also software engineers and kind of like the software layer can do to optimize how these models are trained so you can squeeze even more power out of existing GPUs. And I think both efforts are happening at the same time where more GPUs are being kind of like put into the market and at the same time, more optimizations and better techniques for model training are going to be kind of developed and are developed. And so by putting both together, you're going to have larger and bigger models come up.
22:40Hi, good tech solves problems you know about. Great tech solves problems you haven't even thought about. What can the commerce platform trusted by millions of merchants do for you? It's time for Shopify, the commerce platform revolutionizing millions of businesses worldwide. Whether you're a garage entrepreneur or IPO ready, Shopify is the only tool you need to start, run, and grow your business without the struggle. Shopify puts you in control of every sales channel. So whether you're selling satin sheets from Shopify's in-person point of sales system or offering organic olive oil on Shopify's all-in-one e-commerce platform, you're covered.
23:23Shopify powers 10 % of all e-commerce in the United States, and Shopify is truly a global force, powering Allbirds, Rothy's, and Brooklyn, and millions of other entrepreneurs of every size across over 170 countries. Plus, Shopify's award-winning help is there to support your success every step of the way. Sign up for a$1 per month trial period at shopify.com slash ionai. How do you measure these models? Do you measure them by parameters the way you do in LLM? Depends what you're measuring. I think if you're measuring consistency and quality, you can rely on benchmarks and evals to make sure that models perform as you want them to perform.
24:13Although I'm less a fan of benchmarks and evaluations. and much more of a fan of quality with humans. If you want to see how good the model is, it really doesn't matter how many parameters, how many billion set of parameters the model has, or how long it's been trained. If you want to see how good the model is, just show it to a filmmaker and have them use it to see what they can make with it. And that's going to be a much better indicator of if you're making a good job or not, if they like it, at least from a visual standpoint. That's kind of the goal of making these tools for artists in the first place.
24:46I would presume that the larger the model and the more compute applied, the more data and the training, the better the produced video. Exactly. So the larger the model, both on parameter count and the larger model on the data side as well, the better the model will come and the better the model will work. So on parameter count, what is your largest model? I can't really go into too much of the details, but it's on the billions in terms of size of parameters. Yeah.
25:21It's not saying that much, Cristobal, because everything's in the billions these days. But yeah, and in terms of this, you're using GPUs, I would imagine, for the training? Yes, we're using mostly GPUs. These are H100s and A100s for the most part. Yeah. And can you give an idea how many GPUs you have involved in the training? I can go into that level specifics, but a lot. We have a lot of GPUs. Yeah. And so where is this going? You know, you've made a lot of progress. There's some other people in the market that are making progress. Where on the curve are we? I think it's early, you know, I think it's very early.
26:07We're although in the midst of a major change in media, I think we are about to see a new kind of media that we haven't seen before. I don't think we're going to use the same word to describe it. So where we're speaking about right now filmmaking and movie and movie making, we're probably not going to use the same word to describe the type of things that you can make with these models in a couple of years. And that's kind of like a very interesting point because technology will create a new art form and we've seen that in the past. And so perhaps the Tisch schools of the arts in the future might be teaching something which just don't have a name for it yet.
26:43In the same way that perhaps filmmaking for people 150 years ago would have been like unthinkable off. You can imagine that someone will study something like that and now it's like of course an institution and something that's well known. And I think that's the interesting and more compelling part for me that technology is moving so fast and so quickly and in such an affirmative way that it's not going to only change the way we see the world, but it's also going to change the way artists interact and tell stories of the world. And that's going to create on its own a new art form. And so, yeah, it's still the very beginnings of that transformation.
27:21And so right now, most of your users are commercial advertising firms. I mean, who's using Runway for video creation right now? Yeah, we have everything from directors, producers, artists, filmmakers, production teams, people in Hollywood working a lot. we have musicians, we have advertising companies, and also creatives and creators, independent creators, people who are just making videos for fun or for a living. It's really anyone who wants to make and tell stories. That's kind of the goal of building Runway in the first place. And it's a wide spectrum of creatives that spans from professionals to casual creators.
28:06Yeah, we have what we call like a streaming website where you can watch things people have made with Runway and that's open. So we can also follow up if you want to just see anything people have made. And there's also in the application itself when you want to generate, you can also watch and kind of get inspired by what other people are making. So by doing those things combined, you can basically get a lot of inspiration to not only watch, but also create because that's kind of the point. You are now in control and you can not only view a viewer, but also you can be part of the creation process of making video like that in the first place.
Read the full transcript
28:39Yeah. How is it priced? Is it priced? I mean, you're spending a lot of money to build these models, and presumably your backers want to see a return. But on the other hand, if you're targeting the creative market, I mean, certainly Hollywood has a lot of money, but that's not typically where the money is. I mean, we're charging. It's a license system. Very basic. You can basically pay on a monthly, yearly basis and get access to unlimited resources in the tool. And there's also the ability for people to fine tune models. You can come in with particular data sets and we can help you customize the models to work much more specifically to whatever output you want.
29:29And that's something a lot of companies do ask for because they want to have styles or consistency of objects when they're making stuff. And so we offer that as well. Oh, that's interesting. So do they get an instance of the model in the cloud or something that they fine tune? How does that work? Yeah, they work with our team, with our research engineering team. They can provide us with data and we use the data to customize the model that then gets used only by them. And on the issue of consistency, which is the big hurdle, can you talk about technically how you address that? How do you get a model to be consistent over time?
30:15And you said you're up to 18 seconds or so. So is the limit there maintaining consistency or is the limit just the cost of inference or whatever it's called in this case, that it just becomes prohibitive to allow people to make longer videos? No, the cost is, I mean, there's different challenges for sure. One is with cost. Models haven't gone into the space yet of optimizing them to make them faster. faster, we're working on it, and I think other people are as well, and you'll see costs going down for inference and training. So that's only a matter of time, to be honest, that costs will continue to continuously go down.
30:58And I think we've seen already this in like language models where running language models is now kind of like possible even with like your computer and your iPhone or your phone. And so that's becoming like very cheap and very convenient to use. For consistency, though, for creating larger models, just training data matters a lot, but also the techniques and the tools and the systems that you can put in place to make models just have a longer window of context of how you're going to basically understand sequences of video more coherently. And there's different approaches of it. I think people are still researching on the best technique there.
31:35I don't think there's one single answer because it's still very early. Yeah. You know, you started out at Tisch, which is an arts school. How did you get into this deeply technical space? Were you a coder? Did you become a coder? Or are you more on the conceptual end and you're working with people that understand the algorithm? In my previous life, I studied econ. I was working on statistics and econometric analysis. And then I fell in love with programming and software engineering, which led me to falling in love with neural networks in 2015. So really early on, on this new wave of research that we've seen.
32:29And then I've spent two years at NYU, really in art school, but also I took classes at computer science and current and engineering schools. And so I went really deep into trying to merge and combine both computer science, AI research and arts. And I was coming from passions and interests of mine, which is just how to make sure you can build software tools that can be used by artists. that's kind of my the interest that I've always had and so it's more of a at the time I guess it was like unique and strange to be honest to be able to work on art and computer science and research and now with the models and the quality and how many people are using the stuff that we're making it becomes more obvious and I think that's kind of what drives all of what RONON does and how we still grow as a team we are a team of artists and a team of researchers that share and have a common language.
33:24I think that's the core aspect of what we do and how we do it so specially is we have some of the best scientists working in these problems, but they also have sensibility for the arts. And sometimes they're artists in their own right. And at the same time, we have artists who become engineers. And that allows the company to build tools in a much more different way than I think others. This is going to evolve as an art form on the one hand. And then also, is this tech applicable to AR and VR? Or would that require an entirely different approach? So I guess to answer the first question, how it's going to impact art and media.
34:06You know, when photography first came around, actually when daguerreotypes were first invented, people just didn't have the word to describe it. And so they used this idea of a mirror with a memory. Because you've never seen anything like that. You weren't able to see a technology like that. And it's only by being exposed to it and using it that you eventually become comfortable with exploring it creatively. And I think that's pretty close to where we are right now, where people are still in the mirror with a memory phase of AI. And only by having artists and creatives using it, you're going to see how it's going to impact their own craft and their own art practice.
34:40And that's a lot of what we need to do is get this into the hands of more people. Will, is this tech applicable to AR and VR or will it require an entirely different? I think it will, for sure. I mean, if you think about pixel generation in a VR and AR environment, you're also generating pixels. So the way you can generate them can be different. This is a new way of generating pixels if you just think about it in a more fundamental way. hey, you can find ways of applying or this thing that will have been applicable to also generating pixels in a VR environment. I think it's still early. I think I've seen a few experiments here and there, but it definitely will be an area of growth where people are going to start merging and combining more of those two spaces.
35:24For people that are using your models, what is the interface? How do you manipulate what you want? It's an interface that allows you to do things with a combination of different inputs. You can use text, for example. You can use language, another language, to describe something you want to create, and the model would create it for you. You can also use an image as a reference, so use an image and then use perhaps a language description to animate that image. You can also use sometimes tools. We have something called Motion Brush that allows you to move particular objects by brushing over them and then defining movements in particular directions.
36:03and so it's a there's no one singular input or way you can make them it's much more it's super expressive and you can take different routes on how you want to make that that work are you using the tool you you you were in in art school are you an artist yeah of course i mean these are the tools i've always wanted to use so i'm a user of everything we we build and um that's kind of like also common in in our team as well that um a lot of the our team members just use the tools every day because these are the tools they've always wanted to use. Yeah. And your use, I mean, is it just for the creation of art or are you using it in producing, you know, some commercially?
36:46No, it's more of like self-creative expression. I'm not making movies and I'm selling those movies. It's more of being able just to like make something and perhaps take an idea out of your head and put it out. But I like just being able to, you know, that's kind of the goal ultimately of making something like Runway. You want to make sure that if you're thinking about an idea and you want to express that idea, technology should be able to allow you to do that without any kind of like hurdles or processes or time constraints. I think it's a unique opportunity to be able to like think about something and immediately see it like visualize or get it done or be able to like experience it.
37:26And so that's kind of the goal, allowing people to do that way more often and more constantly, if that makes sense. You mentioned Hollywood. I don't know if you mentioned it, but advertising agencies and then creators. Which is the biggest community that you see? Definitely folks in Hollywood and production teams have been like a big source of growth for us. But as you know, musicians these days are using Runway a lot. We have major musicians and artists using Runway to make music videos and also the tours, the visual for their tours as well. I think those are probably going to continue to grow a lot.
38:08But also I'm excited to have this be used by maybe folks that never thought of themselves as artists in the first place, because they were far away from being able to access any of the outputs of these models or the things that you can make with it. And so I think long term, I'm also excited to see the next generation of artists that are going to emerge from it. And I think we'll still have a long way to go to make sure that happens. Yeah. And how big is the community of people using it? Just sort of, is it thousands, tens of thousands, hundreds of thousands? In the millions. In the millions, yeah.
38:43Yes, multiple millions. and and in terms of uh the the uh the generation of the images i was at i didn't i don't know if you went to neurops this year i was there okay they uh university of toronto uh their drama department has an ai lab and they had uh there was a creative day where they had uh you know, exhibitions of different creators doing things with AI. And they had a poet from New York. And as he read his poetry, there was a screen behind him that would morph and blur and images would appear that were related to the words either in a literal sense or in terms of emotion. Could Runway do that?
39:49Can you run live on conversational prompts? You can do live interactions. I think it's, I mean, I don't know what model they were using, but we've seen a lot of interesting experiments using models in a real-time environment for performances, for dance experiences, but also for music. I think it's a whole area of research and work. But yeah, there's definitely people who have done it in a real-time basis. Yeah. What's the most impressive video that you've seen come out of Runway? You know, every two weeks, there's something new that just blows my mind. So that's kind of the idea, really, that something that you thought was going to be just impossible very hard immediately becomes visible.
40:41And then the standards and the way you can project how the technology will move starts to move on its own. So, yeah, every two weeks there's something that just blows my mind or some major artist or creator or filmmaker who's using runway. And I think that's kind of the point. We need to do more of that. Do you have some place where people can upvote so that the most popular videos are easily discoverable? Yes. So we have a few things. First of all, we have a film festival every year in New York and LA that highlights and showcases the best of the best of AI filmmaking or people making stories and videos with AI.
41:20That film festival is coming now in May in LA and New York. and we already have thousands of submissions from all over the world of people making stuff with Runway. The first version of this, we did it a year ago at New York. We had traditional filmmakers come in and speak about how they've been working with Runway and other tools, and we also highlight and show some of the stuff people were making with what we're making. So that on its own, I think it's a great summary, and I can send you more of that. You can also watch the previous winners and to get a sense of where things are and where things are going to go.
41:53and then we have like ongoing competition actually this week called Jane 48 which you get 48 hours to make a film and we also have thousands of submissions for people over the world to showcase what you can do with it and people can vote on the ones they like the most yeah well I'll be sure and go to the New York version of the festival it sounds fascinating so you have of the video generation, you have two versions of the model or two generations, is that what you said? Yes, we have two generations of the model so far. So Gen 1 was the first generation, Gen 2 was the second generation. And then of course, more generations are coming very soon.
42:36And every generation introduces new quality improvements, control, and a few other tricks here and there. Yeah, when can we expect the next generation? I mean, and what is the cadence of putting out new journeys? Is it getting easier? I mean, there's no particular specific cadence that we keep on repeating. I think we just want to make sure the models are safe and reliable and controllable. But the next generation of models will come very, very soon. At the same time, this is impacting artists and creatives all over the world. And so highlighting their work, for example, in things like the AI Film Festival for us, is just fundamental.
43:12all. And so yeah, if you want to go to New York, to the New York one, you should definitely try to try to make it because it's going to serve as a, I would say, historical moment in time to, for us to understand this new type of media that we're like discussing, and a new type of artist that will emerge out of using these technologies. Art has always been about technology. I mean, from the very beginning of like the first ever paintings in caves, people were using some sort technology. Pigments are something to describe the world they're seeing. A paintbrush is a technology. A camera is a technology.
43:46These algorithms are another form of technology. And so I think battling or putting the artists against a technology is not really the case. I think we should be much more thinking about how does technology augment artists in the first place. And technology on its own doesn't make things. It's people using the technology that make things. And so really embracing it and thinking about it as another paintbrush, another creative tool in your stack is perhaps the best way to think about it. Of course, sometimes you get tensions or afraid of trying new things because you've never tried them before, but that's just part of the process.
44:23Once you start trying it and you get used to it, you understand that it's all about your own intentions and goals. And so yeah, I've seen that before people are becoming afraid of it but again people were afraid of cameras people were afraid of a lot of other things that now are commonplace yeah yeah okay uh well chris this is just fascinating really that's it for this episode i want to thank chris for his time if you want to read a transcript of today's conversation you can find one on our website I on AI, that's E-Y-E hyphen O-N dot AI. And in the meantime, remember, the singularity may not be near, but AI is changing our world.
45:10So pay attention.
From the publisher
Join host Craig Smith on episode #172 of Eye on AI as he sits down with Cris Valenzuela, co-founder and CEO of Runway, an applied AI research company shaping the next era of art, entertainment and human creativity
Cris shares insights into the origins of Runway, highlighting how the company supports creators across various domains, from filmmaking to music production, by leveraging the power of AI.
We dive deep into the fascinating world of generative models like stable diffusion, discussing their development, applications, and the future of creative expression through AI.Discover how Runway's AI tools are breaking new ground in image and video generation, transforming artistic workflows, and opening up unprecedented opportunities for creativity.
We'll also explore the broader implications of AI in art and media, discussing the balance between technological innovation and the preservation of artistic integrity.
Tune in and don't forget to rate us on Apple Podcast and Spotify if you enjoy the episode!
This episode is sponsored by Shopify. Shopify is a commerce platform that allows anyone to set up an online store and sell their products. Whether you're selling online, on social media, or in person, Shopify has you covered on every base. With Shopify you can sell physical and digital products. You can sell services, memberships, ticketed events, rentals and even classes and lessons.
Sign up for a $1 per month trial period at http://shopify.com/eyeonai
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(01:56) Cris Valenzuela: Background and Runway's Genesis
(05:48) Understanding Diffusion Models
(08:01) Discussion on Mid Journey and Other Models
(10:19) Open Source vs Proprietary Models
(13:07) Exploring Video Generation Technology
(17:52) Achieving Consistency in Generated Videos
(21:22) Compute Requirements for Training Models
(25:07) Scale and Parameters of Runway's Models
(29:49) Customizing Models for Specific Needs
(33:51) Applicability to AR and VR
(37:32) Runway's User Community and Impact
(40:21) Runway's Best Work




