In short
Podcast Notes: The Rise of Generative Media: fal's Bet on Video, Infrastructure, and Speed
Podcast Overview Title: Training Data Description: Training Data features conversations with leading AI builders and researchers focused on the evolving technologies of AI and their implications for technology, business, and society. Episode: The Rise of Generative Media: fal's Bet on Video, Infrastructure, and Speed Hosts: Sonya Huang, Pat Grady, and Sequoia Capital partners.
Guests
Founders of fal - Gorkem Yurtseven, Burkay Gur, and Batuhan Taskaya.
---
Key Themes and Discussions
Generative Media Landscape
- The generative media sector, particularly video, has been historically overlooked compared to text and image generation.
- The market for generative video models is rapidly growing, with applications in education, advertising, and entertainment.
Importance of Infrastructure
- Fal's Role: fal is developing the infrastructure necessary for the generative media boom, providing access to over 600 generative media models.
- Current infrastructure challenges include:
- Different optimization problems compared to LLMs (Large Language Models).
- Need for compute efficiency due to the vast requirements of video models.
Technical Challenges in Video Generation
- Video models are compute-bound and experience rapid changes in optimization techniques, with noticeable shifts every 30 days.
- Differences between video models and LLMs:
- LLMs are more memory-constrained due to the need to handle extensive token interactions.
- Video models must handle numerous concurrent tokens and computations, making GPU computational efficiency critical.
Insights from the Demand Side
- Growing demand for AI-native studios, personalized education, and programmatic advertising indicates a shift toward generative video.
- Similar to the early days of CGI, initial skepticism is giving way to acceptance as generative media establishes its place.
Model Diversity and Ecosystem
- The open-source ecosystem for video models is thriving, providing more diverse models than text-based models.
- The "half-life" of top models is approximately 30 days, indicating rapid innovation and deployment of new models.
Notable Use Cases
- Examples of companies leveraging fal's infrastructure include:
- Adaptive Security for dynamic training content.
- Faith, an AI-native studio app for biblical storytelling.
- Major brands utilizing generative video for advertising.
Future Implications
- Generative media is anticipated to enhance existing IP (Intellectual Property) and democratize content creation.
- Ongoing collaborations between AI-native studios and traditional studios are expected.
- The potential for entirely AI-generated films is on the horizon, with predictions of short films (under 20 minutes) being feasible within a year.
Challenges and Considerations
- The transition to a world filled with generative media raises concerns about content quality and human creativity.
- The need for careful management to avoid an "infinite slop machine" scenario, where quantity overshadows quality.
---
Key Takeaways
- Fal's Infrastructure: The company is well-positioned to lead in generative media by focusing on the technical aspects of running models efficiently.
- Video versus Text: The generative video market is primed for growth, with unique challenges and opportunities distinct from text-based models.
- Education and Personalization: The education sector stands to benefit significantly from generative media, allowing for personalized and engaging content.
- Rapid Innovation: A fast-paced environment with frequent model updates reflects the vibrant nature of the generative media landscape.
---
Conclusion The conversation covered the technical, market, and societal implications of generative media, underscoring the importance of infrastructure and innovation. As technology continues to evolve, the dialogue between traditional media and emerging AI-native studios will likely shape the future of content creation and consumption.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We recently had our first Generative Media Conference, and Jeffrey Katzenberg, former CEO of DreamWorks was there and he made a comparison. he said this is exactly playing out how animation when it first came out people people revolted against it it was all hand drawn before that and computer graphics it was new and there was a lot of rebellion against computer driven animation and something very similar is is happening with ai right now but there's there's no way of stopping technology it's just going to happen you're either going to be part of it or not
0:53In this episode, we sit down with a team from Fall, the developer platform and infrastructure powering generative video at scale. Fall is a place that developers can go to access more than 600 generative media models simultaneously, from OpenAI, Sora, and Google Veo to open-weight models like Kling. We'll discuss why video models present fundamentally different optimization challenges than LLMs, why the open-source ecosystem for video has a thriving long tail in ways that text models never did, and why the top video models have a half-life of just 30 days. The team also shares insights from the demand side of the video model equation.
1:27We discuss what's happening in the app layer, from AI Native Studios to personalized education, what's happening in Hollywood, and more. Enjoy the show. Borkai, Gorkam, Batuan, thank you so much for joining us today. I want to start with the problem space that you decided to tackle. So FAL is a developer API and platform for generative video and image models. Video is massive, obviously. It is more than 80 % of the internet's bandwidth, and it follows that generative video is going to be similarly massive. But there's not that many companies that are focused on this problem. Why do you think that is?
2:01Yeah, in a way, generative image and then video was an overlooked market in this current phase of AI. In my opinion, for two reasons. Number one, there wasn't a very clear industry use case that people were going after. There wasn't wipe coding that automates software engineering or there wasn't search, which seems LLM market is going after or customer support, anything like that. also number number two the investment on the research side wasn't wasn't as big three years ago and then that ramped up a little bit slower than lms but still considerably since then and now the the models are much more capable much more useful and real industry use cases compared to what it was three years ago it felt like a toy use case this was just going to be for fun on the side and it's going to be a small market at the end.
3:03And now we can see that it's going to be a massive market with very unique use cases and customers compared to the LLM market. Like if you actually go back to like, as we were experiencing it, I think that was an interesting time. We were working on some like Python compute infrastructure and then these models like DALI 2 had just come out. And then soon after that, ChatGPT had come out And then Lama had come out and we were just like, we were initially, we didn't know that, you know, image and video market was going to get that big. We were actually just curious about like running image models much faster.
3:46That was like our initial entry point. And then we saw like the initial growth. We had a few customers and they were growing really fast. We were like, what the heck is going on? And then, you know, a few customers later, we actually thought, hey, we should double down here. And around that time, also, the other thing that was happening was people were over-indexed on language models. This, like, story of AGI was being told. And, you know, that attracted all the dollars. That attracted all the talent. So everyone was just, like, working on that where we thought, like, we had something niche, like, growing fast.
4:22You know, don't tell anyone. And then we just like started focusing on that. And soon after, like as we got more familiar with the models, we thought the... I remember, I think we changed our website copy to say generative media, like generative media platform. And then it was only like two or three months after that, Sora was announced. So we were definitely like ahead, but we really saw the whole like future kind of coming with like better image models, video models, et cetera. So, yeah, we made this early bet. I mean, you guys have a front row seat to the sorts of new experiences people are building.
4:59I think the market's only going to expand from the media market that we know today. Yeah, absolutely. I think like Alkwara Karpati tweet, you know, no good podcast without it. He did one like recently where he was talking about like why he's excited about the, you know, media models. And one of the things he said was that, like he also mentioned that, right? people are visual and we're gonna we have more so much more video than than text like wall of text and he was saying like um he was making a point around like education and a lot of the like content you consume just to like like learn things i think right now the model quality is just like like relatively it's just so much like it's so much worse than what it can be where like you could actually like have you know i do a lot of learning on chat gpt but it's like through text but if it actually rendered a video where it could compress like a concept right instead of you know 10 000 characters if it could do it in in like 15 seconds uh it'd be so much better i think like there's there's sort of like uh the quality bar where if like it's it's gonna go up and and if you have once we have that we're gonna have even more more penetration so it's a it's really a function of the quality right now and we're just like in the very early beginnings totally education market almost untouched right now with video generation and there's so much potential there and it's just waiting the quality the predictability to get there and i think it's gonna have a lot of potential totally i mean you guys sent me that generative video bible app i think it's a much better way to learn some of the lessons from the bible and it's it's you know if you're capturing consumers attention right where they are you know i think it's i agree with you We're just at the beginning.
6:47So Fall is an infrastructure company, and so we're going to structure today's interview. I love infrastructure companies in terms of the technical layer cake. So we're going to start from the core inference engine, compilers, kernels that you've built. We're going to go up to the model layer and then the workflows and then end with some observations on the markets and what people are building. Sound good? Let's do it. Sounds exciting. Okay, let's do it. The inference engine. Batuan, how old are you? 22. You're 22 years old. Okay, say around your background, I think it's super badass and makes complete sense why this company is so hardcore.
7:21I started working on compilers when I was 14. So in a way, I have a lot of experience on that front. You know, it's not just that. But I started working on open source projects. So my first contributions were around tooling around the Python language. And then I started to slowly contribute back to the Python language, core compiler, core parser, and the core interpreter itself. And became one of the core maintainers of it. uh i think at the time i was like the youngest core maintainer of the language and this kind of gave me like a unique appreciation of compilers and how flexible they are so when we first started like working on serving these image models at fall the the main idea was okay there is like these like three different image models three different architectures but this is surely gonna explode there's upscalers there's like you know there's gonna be video models we were predicting that and we didn't want to go optimize a single model, put our eggs into a single basket and then became invalidated when the next model comes.
8:19So we started building this inference engine, which is a tracing compiler that traces the execution and essentially tries to find common patterns that are fitting within the templated kernels that we do. So our bread and butter is like spending, we have a 10 % performance team that's spending all their efforts into writing kernels that are like 95 % there, but like generalized with templates. So we trace the execution of a model and find common patterns that could replace these templated semi-generic kernels to specialized kernels at runtime and optimize the performance of these models. And we found this technique to yield pretty much superior results from anything that's out there in the market.
8:56And this led us to claim number one spot on performance on all the benchmarks. And another big thing about this is we specialize in doing this sort of kernel-level mathematically correct sound abstractions that let us maintain the same quality of these models, which is a very high bar when you're in this media industry and when you really care about the output that you're getting. What's different between optimizing a diffusion model versus an other regressive LLM? In autoregressive LLMs, your bottleneck is how fast you can move all those giant weights from memory to SRAM because you have a$600 billion parameter model and you're trying to predict the next token.
9:34You're doing the attention for all the tokens, like a couple of tokens before that. in diffusion models you're trying to denoise like thousands tens of thousands of tokens for a video at the same time doing attention of it so you're essentially saturating all the compute bandwidth of these GPUs you're not necessarily bound on memory bandwidth but like the computational operations that you do are like fully saturated so you're trying to find better ways to execute around the GPU this could be like writing more efficient kernels or this could be overlapping you know softmax with gems that you do like it's essentially like you're trying to use all of the power of the GPU, leverage it in a way that like, you know, that gets you all the capabilities.
10:12So it's a different binding constraint. It's on the compute versus the memory. And what's the intuition for why LLMs are relatively memory constrained and why video models are by comparison relatively compute constrained, but not as large in terms of just sheer number of parameters? I think it's scaling issue, right? Like in terms of like, if you scale these video models, 600 billion parameters with the same dense architecture, you're going to have to do attention with all those full like 100 like let's say a single video is 100 000 tokens and you do this attention step or like you do this like denoising step 50 times and every 50 times you do like attention over all these like 100 000 tokens it's insanely insanely expensive so i think the constraint there is like just how fast you can do the inference and the same applies with lms at like larger batch sizes but like at like the traffic patterns that people do like you know the batch size are not that much and you're mainly constrained by memory bandwidth so people do optimizations like speculative decoding and other factors to like reduce that overall yeah what exactly goes into being at the top of the leaderboard in terms of performance because i would imagine there's other teams that are also very smart people and you know this is this is my olympics and so what exactly goes into i imagine people have like very similar ideas on the techniques and different optimizations they can do so i don't think anyone cares about it as much as us we are literally obsessed with jarrington media we're literally obsessed with these models we have a team that's like just focusing on this like so far like it seems like and from nvidia to other inference players everyone is like super obsessed with language models everyone's trying to get like one more tokens per second on like deep seek benchmarks whatever and like we're like on a on a different lane we have like competitors but like no one close to us because i think we assembled one of the best teams we found out like the best way to optimize these general models and we just focus on this this is like a purely focused thing right like at the end of the day you're constrained by the hardware.
11:56There's nothing unique about it. But like, we're just like three months ahead, six months ahead. Like when we benchmark Torch, like the latest version of Torch against like, you know, our inference engine from a year ago, we're clearly underperforming because like Torch caught up. The same thing is going to happen with other players. You're going to always like, the lead that you can maintain is three months, six months ahead at most. The thing that matters is just focus. If you focus on it, if you purely put all your energy into it, I think that's like there's, it's very hard to get outcompeted by others.
12:25Because models are slightly changing each month, each release. So it's still the same general architecture, but there are slight differences where we can go in and optimize where that's different. And no one else is paying that much attention to it. Also, hardware is changing as well. So we were able to adapt to B200 earlier than anyone else. And we were able to run video models much faster basically throughout the year because of that obsession with running video models with the latest hardware. Yeah, got it. What are the hardest technical problems that you think you're solving? So one thing people don't appreciate it as much is we are running 600 different models at the same time.
13:06We have to be running them. We have to be so good at running them that we should be running a single one of them better than as if someone else is running a single model. because when a foundational lab is running models, maybe they have a single version of the model, maybe they have a couple other different versions, and that's what all they care about. We have to be better than them at running those models, and we have to be doing all 600 at the same time. So on top of the inference optimizations that happens in the GPU, a lot of optimizations on the infrastructure level needs to happen. We need to manage the GPU cluster in a way that's efficient to load and load these models at the right times.
13:45We need to route traffic to the right GPUs who have the warm cache of these models. We need to be smart about choosing the right kind of machines, which kind of chips are running what kind of models. And the customer traffic is changing all the time and we need to adapt towards that. So on top of the inference engine, the overall infrastructure is also really, really hard beast to manage. And so far, we've done an incredible job at that. Would you add anything to the... I think that's a pretty fair explanation of what we do. I call this distributed supercomputing. I don't know why people don't like that name, but...
14:23I like it. But the idea is we are at 28, 30... This was a month ago. Now we are probably at 35 different data centers. And you have these heterogeneous groups of compute that split across with their own different specs, different networking, whatever. And you're trying to schedule workloads as if it's like a homogeneous cluster that you got from hyperscaler. It doesn't work like that. So we built, like we spent the last three years building that tractions over it from our own orchestrator to building our own CDN. Like we go back to like, you know, fundamentals of web and we built our own CDN service, deploying racks to call us, like, you know, just like routing traffic.
15:01So we built all these technologies to essentially make sure that we can tap into capacity wherever it is and schedule our workloads, which is like very different than like a traditional enterprise LLM usage pattern, you know, like the use case that we have are like so much more spread across so much more you know like uh more consumer facing and like when you consider that like there is so much investment going into making sure we can tap into this like scarce capacity of gpus yeah you mentioned hyperscalers and you know i hear distributed computes and i hear managing giant clusters and i naturally think that's somewhere where hyperscalers should have the incumbent advantage um why do you think that you've been able to so far execute them on the core engine?
15:41There's two things about the core engine, right? There's the inference part where none of the hyperscalers have any expertise. This is a net new field. I think this has been only happening for the past three years, inference optimization. So it's like a brand new lane that we have been outcompeting anyone in our field. I think that's like a pretty much answer of its own. And the second one is infrastructure. I think right now hyperscalers are very busy with their traditional pattern of, oh, we have this data center capacity, we'll just deploy GPUs and we don't care about trust. This has been changing recently.
16:11Even Microsoft is going buying from NeoClouds. There's an interesting pattern happening because the GPUs and the demand and the growth of GPUs doesn't fit to the patterns of these hyperscalers, the growth patterns they expect. So I think at this age, not even hyperscalers have that big of an advantage of scale because they're going buying GPUs from NeoClouds. The tables have turned a bit. Yeah, it almost helps to also be like slightly earlier, like in the company journey, right? Like if you're a public company, you also have to kind of abide by like what the market's expecting of you. So like the other thing is that there's a huge price discrepancy with hyperscalers and neoclouds, right?
16:54So like it's maybe sometimes 2x, 3x more expensive to use things, you know, through hyperscalers. What's driving that? Well, I think one is like market pressure, right? And also like there's added kind of operational expenses that hyperscalers have for like having, you know, better, they just have a better service, right? Better uptime and better SLAs and all of these things add up. And then on top of that, there's kind of an established like cloud margin, right? And, you know, the market expects the cloud margin to be a certain level. Whereas if you have a three-year-old neocloud, you're a private company, maybe you don't have as much pressure.
17:38And there's like, assuming infinite demand and limited capacity, you can actually, hyperscalers can keep their prices high and they will fill out the capacity and also get slightly better economics. Whereas neoclouds compete over the whole infinite demand and that pushes the prices down. Perfect price competition. What does it take to run image versus video models? Well, like you guys started the company around stable diffusion moment. It was the field was mostly image at the time. How does running video models compare to image? Let's actually do text image video. Even let's compare all three of them.
18:18And so for let's say a SOTA LLM, like we don't know, let's say DeepSeq or something like that, where we know the numbers. Running a single prompt, like 200 tokens. let's say it takes 1x of teraflops. I think it's tens of teraflops, but let's call that unit 1. One image is around 100x of that. And if you are doing a 5 second video, 24 FPS, that is around 120 frames. So 100x from one image. So you are already 100x of 100x. So you are at 10 ,000x for a low standard definition video. And if you want to do 4K, that's another 10x. So 10 ,000x compared to a single 200 token LLM input. So it is a lot more compute intensive in terms of amount of flops you are doing.
19:22Yeah. in general like when we started with image the infrastructure was relatively easier to do because it's like it takes three sec or it took 15 seconds back in the days it takes 15 seconds to generate an image you don't need to necessarily shave like that 50ms 100ms you have overall in the system and then when we went to video it's like even like easier because it takes like 20 seconds 30 seconds to generate a video the way that has been happening in the past like couple months is real-time video where you need to stream 24 FPS videos over like a network link from these GPUs. That's where we actually spend like some of our time.
19:57We started this progress with like speech-to-speech models a year ago. We started optimizing them where we were able to like reduce the latency of our system with like globally distributed GPU fleet. When you send a request, we route to the closest GPU, minimize our own overhead and then do stuff like, you know, we pick the best runner, whatever, stuff like that. So we are now applying those same optimizations we did to real-time video. And we actually see like really good, interesting demand there where people want to experience this stuff like as they type, as they prompt. And that's where like some of the infrastructure technical challenges differ from traditionally running image and video models because image and video are similar-ish, you know, like just more compute expensive.
20:35But like you actually need to care about infrastructure stuff when you go from like less than a second generation time for some of these models. Yeah. Another interesting thing is like image models especially, you were able to run them on a single GPU. Like the parameter counts is actually much smaller. So that actually makes it like a little bit easier for us as opposed to LLMs. And then with video, parameter count is going up. Right now, I think we're around like for the open source ones, I don't know, 30 billion parameter. Whereas, you know, we hear rumors about, you know, GPT-4 being like in the trillions, GPT-5, maybe more.
21:15So that's another, that's like, you know, on the flip side, it's a little bit easier, but it doesn't mean video models are not going to grow, right? There's, you know, rumors around numbers for VO, numbers around Sora. So like, there's also a increase in parameter count. So that you're going to have to kind of, you know, use more distributed computing. But if you're just, you know, eight nodes, or one node or eight nodes, you kind of have a slight advantage. Yeah, totally. Okay, let's pop one layer up the stack to the models. Let's do it. So one thing I think people don't fully appreciate at the media space, and you mentioned this, you alluded to this before, is that there's a very, very long tail of models that are actually used in practice.
22:00And so I was hoping you'd give people a sense of on your platform, how many models are people actively using? How is it distributed? And like, why do you think there's such a long tail of models being used? compared to the alarm space? This is actually one of the things I would say three years ago, people got it wrong. I mean, jury's still out, but people, right after the chat GPT, people start talking about omni models. There's going to be these giant models that they're going to be able to generate video, audio, image, and code, text, every type of token. This might still happen, I think, but it's more clear that you are better off if you optimize for a certain type of output.
22:41Even this is true for code generation, definitely true for image or video output. So that's one thing when we were pitching three years ago, everyone, like that's one feedback we got. Oh, there's going to be omni models and there's going to be a single way of running these. It's going to be hard to create an edge on the modality. but turns out it's not true and it actually makes sense to have a technical edge on the modality. And this is one of the reasons why there's also a variety of models because still the best upscaling model is just doing upscaling and the best image editing model, even the best text-to-image model is different from the image editing model.
23:25So all these special tasks require their own model. It might be the similar model family or similar architecture, but at the end of the day, it has its own weights that needs to be deployed independently, and that creates the variety in the ecosystem. I think this also applies to language models, where even in the same modality, there is different families of models with different tastes, different characteristics, different personas. and this happens with like language models still like the code that cloud writes is very different than the code gpt5 does right like and like we see this happening but the good thing about here is like there's like these three four different personas on top of like different like categories upscaling editing video text video whatever stuff like that so it gets you like you know close to 50 models that are active at any point in time and then you have a very long tail of models that people still choose because they might like the person of that better.
24:23Yeah, totally. Speaking of model personalities, what are some of the most popular models on your platform? What do you think are the personalities of them? So one thing that's been true since the beginning, the popular models change all the time. So there's always new releases from different labs that take over the other and it's always a moving target. But that being said, there's two types of models usually preferred by our customers. Usually there's one big expensive model that has the best quality on video generation. This could be Vio, this could be Kling, this could be Sora. And then there's usually a workhorse model, which is cheaper, smaller, but good enough.
25:08And people usually use that at higher volumes. I would say this has been true for the past almost two years is that there's an expensive high quality model that keeps changing. There's a cheaper, good enough model that keeps changing. But overall, this has been constant. And is the workhorse model for prototyping and then you run it through the big expensive model for the final product? Or what do people use the workhorse versus the last? It's for higher volume use cases. And depending on the application you are building, you might encourage different, like lots of variations of the same output maybe, but it's very application specific, I would say.
25:48Yeah. There's also another dimension, I think, that's like kind of happening in real time right now, which is based on like different use case you want to use the model for. So like when OpenAI released GPT image editing, that model had just like superior text editing, text generation and editing capabilities. and for things that require like a lot of text, people started going and choosing that model versus the other models. So it also tends to correlate with like different capabilities models are bringing and also like what they're good at, right? So like Kling, for example, people really like it for visual effects types of workflows because they had that kind of data in their dataset.
26:39as opposed to some other models. For example, C-Dance is very good at detailed textures and artistic diversity, things like that. So it's really a matter of also this sort of use case dimension that models excel at. An interesting metric that we saw on Q2 and Q3 was the half-life of a top five model was 30 days. It's very, very interesting to me. where like, you know, these models are continuously shifting, like the top five of the models are continuously shifting. Tough depreciation schedule for the model providers. Hopefully they are building on top of the work that they already done. So it's, you know, additive to the end, but yeah.
27:26Yeah, I'm teasing. And the model probably is in a more turbulent state right now than what the end state will probably be. What do you guys think is the most underrated model? Like what's your personal favorite? it i i usually like cling models for video um but this kind of has been changing because they don't have sound uh for sound we have vo3 and sora they they are the only ones a lot of people are working on it so would love to have more more variety there as well image models i like revs model and like flux still holds like a very nostalgic even though it's been a year value for me you know like I still go back to Flux.
28:08There's like variations of Flux models now that I like. I'll go with mid-journey, which is not on file. It's not available on API. I just like how they navigated the space, I think is very interesting. Like they kind of brought this like photorealism, which was, you know, that was like a very big deal at the time. You know, no model could do it. And then now they're more like this artsy model, right? Like it's no longer, photorealism is kind of cracked and like no one cares about it. So now they have this like niche, very artistic like visuals, which is very cool. Yeah. I'd love to chat about the marketplace dynamics a little bit.
28:47So I understand your business as a little bit of a marketplace where you aggregate developers on one side of the market, that's the demand side, and you aggregate model vendors on the other side of the market, that's the supply side. And the model vendors are both proprietary APIs, model labs that view you as a distribution partner and then also open models that you host and run yourselves. And so maybe talk a little bit about for the closed model providers, you have partnerships with OpenAI Sora, with DeepMind on Veo. What's in it for them? Why do they choose to partner with you? We were one of the first platforms that accumulated the developer love.
29:29And following from that, these developers work at big companies. So they started working with us and we really built the platform for simplicity and being able to get going really fast. And because the thing Batuan mentioned, the half-life of these models is really short, people usually work with many different models at the same time. So we were able to claim that we have this big developer base that love the platform and not tied into any single model and here for the platform. And model research labs see this and they use the platform as a distribution channel and tap into the developer ecosystem that we built.
Read the full transcript
30:17On the other side, this helps us with the next model provider because they see all the developers, they want to be on the platform as well, which attracts more developers on the platform and creates a very nice positive flywheel for us. Yeah, it very much is a marketplace business. And for developers, it's a single choke point to be able to access multiple model vendors. And to your point on the model space is changing so quickly, I think they really do value that choice. We call it Marketplace++ because we get to provide infrastructure to the research labs as well, also to the developers. So there's additional benefits which ties into the flywheel effect that we are creating.
30:58So it's Marketplace plus other services next to it. How do you position yourselves to get, you know, in some cases, day zero launch access, sometimes exclusive launch access to models like Kling and Minimax? How have you done that? Yeah, throughout the last two years, we were able to build a very robust marketing machine as well. And this is our connection point with the developers who are on the platform. Every time we release something, this creates another opportunity for us to introduce a new capability, introduce a new model. And model developers also see that. And we usually do co-marketing together.
31:36And part of that co-marketing, we get exclusive release access for a certain period of time, sometimes forever. We have a couple of competitors that are on the smaller side. So model developers want to work with the biggest platform out there. And increasingly, that platform is ours. And we get to have these exclusive benefits with the model providers. That's awesome. Why do you think it is that the open source model ecosystem has been so vibrant for video models? It almost feels like the text models are just consistently a generation behind whereas in video you know there's there's so much that's happening in the open source realm video and also image editing image editing as well why do you think that is it started with stability they first open source stable diffusion and got insane adoption and almost the same team then started black forest labs and they knew the power of open source how it helps them create the ecosystem and with image and media models the ecosystem actually matters when developers are training lauras they are building adapters they are building on top of your model it really brings free marketing but also creates stickiness so that the developer there are still people who are using stable diffusion models because they like that ecosystem because it was so open yeah and so the flux team saw this from their experience at stability and they had a very smart strategy of having at least some models that are open source some that are closed source and a lot of video model providers that came after is following the same playbook because you can have a very robust ecosystem it gives you a lot of advantages in terms of marketing in terms of developer love and i think it's going to keep keep going like this yeah i want to add on to that like the domain is also very interesting like i think in the visual domain like ecosystem actually matters more like i think when when like lama2 first came out there was like many fine tunes out there but like if you actually downloaded it and start using one like you can't i mean you can't tell it's a fine tune you can't tell like the difference like you can't really you know if you're using a uh i don't know like a control net like the concept doesn't even exist like it doesn't you know language models are a lot more general like generalized so you can't really understand understand like the difference if you if you were to actually fine tune it right so so it kind of just ends up being very monolithic as opposed to like in the visual realm it's just like um any small adjustment you make to the model it can actually uh you know uh it can actually have huge implications right and so and so it's just it's just uh you know very uh fertile ground for like a lot of a lot of customization yeah i mean speaking of mid-journey david holtz one of his one of his quotes that i like is you know he's curating the aesthetic space uh with mid-journey and i very much think you just have this combinatorial explosion of styles yeah aesthetically and i think that's the reason why some of the i think some of the models on your platform are fine tunes of other models yes and like the the thing is like even if you add a lot of diversity of aesthetics.
34:58Like if you actually train on everything, like if you have trained on too many, you may not be able to like actually get the exact, like there's so many times you want the exact aesthetics and then you may still have to like fine tune the model to get exactly the output you want. Whereas like with LLMs, that's not really like how you operate. You don't exactly want a particular outcome. It's like a different, it's a different problem. So this is a lot more, you know, it's very subjective. So like you kind of have to do these like post-training things on top of the models. Sora is another good example.
35:34Like Sora 2 is very fine-tuned on like social-looking stuff, right? And so, you know, you could probably, you know, you can have tens of different styles and you still want to probably push the model towards that direction with post-training. Yeah, absolutely. It all depends on the use case too. A customer support chatbot does not need personality. You want it to be as vanilla as possible, but we are talking about filmmakers, marketing teams. They all want to add the personality of their style or their brand. So they want to have greater control over the outputs, whereas maybe in LLMs, that's not necessarily true all the time.
36:19If you have an agent, if you are doing cogeneration, there's no equivalent of style and personality. Yeah, okay, that's a good segue for us to go one more layer up the stack. Let's go to workflows. What does the average developer workflow inside Fall look like today? They are using many different models, first of all. So we looked this up recently, our top 100 customers, they are using 14 different models at the same time. These are sometimes chained to each other. So one text-to-image model, one upscaler, one image-to-video model, all part of a same workflow or like a more complicated combination of this part of a same workflow or different models used in different use cases.
37:05I think that's the most interesting part, the variety of the models people use on the platform. We do have a no-code workflow builder as well. We built this in collaboration with Shopify, and this is usually very good for their PMs, their marketing teams, the non-technical members of the team who are playing with these models. It's really good for trying different things, really good for comparing different models, but eventually this makes it into the product as well. You can reach to this workflow through an API. It's been very popular recently and more and more people in a typical software engineering organization is now interested in image and video models.
37:51So the users of this platform has been increasing. Okay, so the average workflow is not just text to prompt. It's not create a five-minute commercial that does it. If I wanted to create a five-minute commercial, what would the workflow be? Yeah, so for this reason, people actually prefer open... Like that's one of the reasons why people prefer open source models, because they get to have more control over the model and they can add things here and there to steer the model towards the outputs they want. And when we go talk to studios or more professional marketing teams, they all love working with the open source models because of the pieces they can replace and control they can add into it.
38:35And then these workflows are usually like the ones, if you've seen any like big conf UI workflows with many different nodes, it resembles those where each different piece of the model can be replaced to create more control for the creator. Got it. Yeah. And I think like what we have, like our workflow tool, it's not the final form of like, there's almost like another layer of abstraction, maybe on top in terms of workflow. And like, as we talk to like these studios, we actually figure out like there's so many ways of just like there's so many ways of using Photoshop. Like there's no single workflow.
39:12In fact, like there's probably like based on your role, right? Like you're a marketing person or you're an animator or whatever, like you have different workflows, right? And so I think that is also emerging. Like as more and more like professionals are actually starting to use these tools, like you see the emergence of like very particular workflows, right? One of our favorite creators is PJ Ace. he actually like shares his workflows uh online and every time like he posts things you know every month he actually has like a different kind of workflow it's it's really driven by like the new models like based on based on new model he may have a completely new new workflow next time i think i think once like we sort of reach some sort of uh i guess like some sort of productivity and, you know, some professionals actually adopting these tools, there will probably be more sort of standardized, like best practices around using these abstractions.
40:15But like, you know, it's not, I don't think anyone knows like the final form yet. And it's like every day we see new things and we try to like update our product to make sure like it caters to those people. Totally. One of the workflows I'm seeing somewhat commonly is, you know, you have an idea for high level what you want and you type that in and then um and the aesthetics that you want and you iterate on the aesthetics from an image model and then use that image model with the aesthetics you want to then generate a series of images which then form the storyboard yes so to speak and then it cascades down exactly from there and then the video models kind of interpolate in between them and it's funny because that's actually how you know that's how you know pixar and all these companies work right in terms of storyboards and so i think it was a cost thing in the beginning.
41:01Yeah. Like, that's why they had to do it like that. But, like, it actually also makes sense, right? Like, it makes sense in so many ways to do it like that. And, yeah, they call that stuff pre-production and then, you know, post or production, right? So pre-production is all the tooling around storyboarding, et cetera. Like, that's what everyone does, like, even today. Even though it was, like, a very cost thing, now it's more of a speed thing. And AI makes the workflow, you know, very interesting where you have everything laid out and let's say a new model new text image model comes out they built it in such a way that okay you can press a button and now all different combinations are going to be generated with this other model and then you can like generate all the videos again we've seen those insane workflows you want to update one thing and the whole thing is going to cost like a thousand dollars to rerun it again but these individuals like they spend a ton of money on creator platforms.
42:01I've seen bills like half a million dollars just spent by a single individual and maybe even more when it's a small production studio, stuff like that. So it's pretty incredible. Totally. Wonderful. Okay. Speaking of studios who are building on your platform, let's go our final layer of the stack. Let's talk about customers and markets and then what the future might hold. Maybe what are the coolest things that people are building on your platform today? And are they what we would think of as traditional media businesses, or are they net new businesses? It's all over the place. What's so exciting about this space is that it just goes across all of the markets you can possibly imagine.
42:47I'll give you some more um i guess long tail stuff first because because it's it's super fun and interesting there's a security company that's building on top of fall and they basically have these like trainings and the trainings are generated on the fly and and the content is all dynamic obviously they have some scripts i'm guessing to to kind of fit like the curriculum but like the the content you get, you know, per person is all dynamic. Is Brian Long's company? Yeah, this is Adaptive Security. Yeah. Yeah, they do some really cool stuff. I think that's one of the, like, most unique use cases.
43:26You can see how that translates into, like, rest of education. I think that market is, like, kind of picking up. Another one, I think, like, you know, this is more common use case, I guess, is, like, like AI native studios. You mentioned like the Bible app. Yeah. That was one of my favorites. It's called Faith. It's one of the like highest ranked apps on the app store. And yeah, they have like stories for each of the stories from the Bible. And they're like really well produced. And, you know, this sort of category of AI native studios, either in the form of, you know, applications or like they're doing like, you know, feature films and, you know, series and things like that.
44:11That's a huge category. So I would call this like maybe new media or like AI native media and entertainment. There is also a lot of like design and productivity, like out of our public customers, like Canva is one of those. Adobe is one of those. So they're integrating kind of like in this, you know, in this older tooling, They're integrating new models. Ads is a big one. So, and ads kind of come in many flavors. Basically, there's like the UGC style ads, like the stuff you see, like there's a person, you know, demoing a product. That's like a very big category. So, AI generated versions of those.
44:55There's also kind of like older styles of ads, right? More professional looking, higher production. production uh maybe you saw the coca-cola ad yeah that came out recently yeah yeah uh so that's that's like a kind of a higher production um you know style of ads but but you know what we're excited about is also like programmatic ads right so where you can do personalized um you know to the degree of like like literally individuals um you know yourself being the ad or in the movies whatever so like that's that's also a big like growing use case yeah i'm most excited for the education use case i think that you know ads is ads is you know the backbone of the of commerce and the internet and so like like super compelling business case but education is a market that's like so important and has never really had that many compelling business cases behind it yes and part of the challenge with education i mean the challenge has been the bottleneck to creating high quality content at scale that's actually ideal for the learner.
45:56And so I'm personally most excited about education. Same. Like, I really love the education use cases. And I actually think that like ChatGPT or like just, you know, LLMs in general, I think they are already solving it in a way, but it's not the right form factor. Like, if you actually want to fully realize like the sort of power that these models are bringing you actually need to go into the visual space because then you know it's so much more compact it's more approachable and and yeah i think once we actually crack like visual learning like through these video models that's when it's gonna you know really just like impact people do you think that the advent of generative media is going to increase the value of existing ip so like mario brothers nintendo disney pikachu all these things or do you think it's gonna to lead to the democratization of the creation of IP?
46:52I love this question because it felt like, I would say six months ago, it felt like this was all happening too fast for Hollywood, the IP holders to adapt and be part of it. And from our viewpoint, we thought, all right, these AI native studios, they're just going to take over and Hollywood is just going to be too slow. And this is going to just go past them and they're going to be left behind. But this summer, something changed and we've been talking to a lot of usual suspects from the Hollywood. We recently had our first generative media conference and Jeffrey Katzenberg, former CEO of DreamWorks was there and he made a comparison.
47:35he said this is exactly playing out how animation when it first came out people people revolted against it it was all hand drawn before that and computer graphics it was new and there was a lot of rebellion against computer driven animation and something very similar is is happening with ai right now but there's there's no way of stopping technology it's just going to happen you're either going to be part of it or not. So we are seeing a lot of existing IP holders are now taking this very seriously. And at least for the medium term, I think they are pretty well positioned because they have the technical people who are actually really interested behind the scenes in this technology.
48:21They also have the IP, but they also have storytelling and filmmaking know-how. You still need quite large budgets. Maybe things are going to get cheaper, but in the medium term, filmmaking is still going to be expensive. Yes, AI is going to make it maybe a little bit cheaper, but we need these deeply technical people who know filmmaking, who has the IP, who know storytelling to actually in the beginning be part of this. And I think they're going to play a big role in the next coming years in the AI ecosystem. Yeah. When there's infinite content generation, it almost puts a value on the things that are finite.
48:58And I think, you know, for those of us who grew up with Power Rangers or Neopets or whatever, there is just this nostalgia element and this finite supply of IP that really resonates with us. The opposite is true too also. There's a lot of new, like we had little toys of these Italian Reynolds characters. These are characters with no IP, no one owns them. They are completely AI generated from the internet community. And once you have cheap generation of content, very different permutations of it things that people like catches on and it becomes part of the zeitgeist so totally the opposite there's signs of opposite being true as well yeah exactly how do you related question how do we prevent like the infinite slop machine state of the world you know there's this you know version where we're just connected to this machine that knows how to personalize stuff for us and we're just you know uh we're just hooked up to the infinite a slot machine and there's a version where there's, you know, human creativity and artistry and things like that involves.
50:02Like, how do you think the world plays out? I think humans eventually, like, converge on the things that are more meaningful in general. Like, I don't know. Like, no matter how much slot we fill the world with, I think, you know, taste prevails and people are drawn to experiences that are personal and human. And I just think that that's going to happen. One interesting example of this was when Meta announced Vibes and then OpenAI announced Sora 2, the reception was very different. And one of the reasons in my mind was Vibes was positioned as this slot machine kind of thing where, you know, they didn't have the product out at the time, but it was just like these AI generated, like, just you have no relation to the characters, etc., right?
51:03Like, it was kind of this, like, detached thing. Whereas, like, Sora really made it about friends, right? Like, cameo, and, you know, they were very vocal. And now you can cameo your pets. There you go. It's huge, right? So yeah, I think this connection to friends and pets and things like that, that actually made, and Sora was also like, they were being very personal about it. They were very adamant about like, hey, we want to make this about friends. We want to make this about these connections as opposed to Influenstrap Machine. So I think that perception was also, I think, a good signal that like there's ways to uh work make this technology work uh you know in a good way absolutely okay i'm gonna get your perspective on timelines and what's feasible today and what's feasible to come i guess do you think that we'll see hollywood grade feature film length films entirely generated by ai and and if so on what time what does entirely generate by ai means is it like no human involvement or no human filming so editing is okay editing yes absolutely human editing but no no human filming i think less than a year we'll have like you know advanced video models with combined with the storyboarding that people have been doing you'll have feature grade short films uh like less than 20 minutes i think that's that's a fair estimation like even today i think you can like do really great films it's just like not enough investment of time is going into these but like with enough investment of time and the model quality i think will be there i think already there okay and you think it's photorealistic you think it's anime you think like what categories do you think are more likely to happen sooner i think photorealistic is like what everyone is like targeting but like anime would be a cool one right like it's like you don't see that many anime specialized models why not i think it's there needs to be a market for that clearly i i think it's going to be animation or or anime or cartoon like one like not photorealistic like as far away from photorealistic as possible maybe even like as fantasy as possible because filming photorealism is cheap and doable already like that's not what costs money when people are making movies it's the non-photorealistic stuff that's actually expensive and um you know even if you look at the animated movies some of the my favorite movies are animated the toy story series how to train your dragon shrek ratatouille and people like these things not because it it reminds them photorealism it's the storytelling that matters and this this created a new medium i think ai is going to be similar to animation and how that brought a whole different angle to filmmaking yeah i think feature films are hard like because yeah with photorealism like you you you typically i mean people usually like the movies that their like favorite actors are in whatever actors actresses and that's you know so so it's like it's like one step removed from that's the thing that costs money to get the actors yeah exactly so so that's the you know we first need to build a connection to this ai you know ai generated character before we can turn turn into a film yeah um but but i think like yeah i think i think it's among like different kinds of content like shorts you know uh i think italian brain rot is an amazing example right it was first like these characters and then it became a roblox game uh and making i don't even know like you know a lot of revenue so so yeah i think i think like ai native stuff is uh and shorter uh form content is is probably gonna be very very big we saw this with vfx where like the vfx effects like one of the most expensive parts of like producing these videos or films is like got like ai fight very very quickly because it's very easy for ai to do like explosions right like or a building collapse it's like almost perfect now and i think it's just going to continue along on that dimension and maybe facial expressions are going to be hard yes and very hard you don't have to do face facial expressions that's going to be okay but now they can do gymnastics yeah gymnastics are important good thing we have a lot of footage of olympics um what about you mentioned roblox at what point do you do you think we'll have interactive video games that are generated in real time yes i think so i i'm very excited about it actually like i think i think like in one world i think the sort of next reasonable step for text to video like if you if you think text to video is the continuation of text to image i would say like a text to game is the continuation of text to video um because you know with it with a game you would you know you would essentially making the video interactive right that's that's kind of what that means and i actually think that there is a world where this like hyper i know hyper casual games exist but this is like another level of hyper casual where it's actually discardable uh i think i think we're not too far away from that i i actually feel like pretty um pretty bullish on having like these you know one-time uh playable games like very short games um i think that's probably going to happen i think that's a good use case for world models other than any other great use cases.
56:38But I think it's going to happen. What about AAA quality games? Will these models at least assist and change the development pipeline of those games? Yeah, I think they're already impacting. At least LLMs are impacting conversations. There's dynamic conversations, things like that. I think pre-production stuff is impacted already. um i think like kind of side quests like ip stuff is impacted right like where you have the assets and you can make a mini game i think people are using it actually not very public but like that is already happening i think like using for a triple a production or like generating that with a model that's like i don't know at least like three four years ahead uh for me and and yeah it's it's i mean that would be insane if we can actually do that um but but you know along the way there just like the just like the uh you know video space i think along the way to the triple a there's like many other things i think those are going to be very big yeah the video model space has just exploded in terms of options quality etc as you look ahead towards what's needed to get us to the promised land for everything that generative media can be.
57:56Do you think that there's, you know, future R &D breakthroughs that are needed on the horizon, like fundamental R &D breakthroughs? Or do you think we're very much in the engineering scale up leg of the race? I think the architecture needs to like slightly change, at least like if you think about like scaling these models by 10x, 100x, I think the architecture is a big bottleneck right now in terms of the inference efficiency, right? Like the more compression of the video space, then that's definitely needed. like we saw this with image in which most used to be like much less compressed and then later like or like you were operating at the pixel space and then we introduced latent space and then like even inside that latent space you took like 64 pixels and made them a single pixel uh and now like with video we are compressing on a time dimension where we are seeing like 4x ratios why not like 24x or whatever like you you need to like increase that like compression and like i think that's gonna be a big driver of improving both inference efficiency as well as training efficiency but like i think like any model like i think at this age that we're operating any model you take on the generating media side we're far from being like scaled up engineering wise like i think there's not enough investment being put into or like it just started happening in the within the past six months like google showed this with like their models and how quickly they were able to catch up they didn't need to innovate that much it's just like they have the resources they can put more more effort into it but at the same time smaller labs are able to demonstrate this because like there's so much like unique and noble stuff that you can do at the data level to train these models so i think that's also like helping contributing and there's the factor of you know like outside like you know mid-tier labs that raise like 100 to a billion dollars that's also trying to come up with models releasing them open source or like contributing to ecosystem yeah that's what's so exciting about this space, there's so much more work to do.
59:43So far, the research community did the simplest thing possible. They captioned images and trained the model on text prompt. And now we are doing video image editing that requires a lot more data engineering to create the data sets. But luckily, seemingly, we have a lot of abundant free video data. We are going to run out of compute before we run out of video data so that means there's a lot more work to do and a lot more room for improvement i mean earlier on like ger cam's math also indicates that like if you want to get to 4k video real time that is like i mean that that means like i don't know 100x maybe more in um in like compute or architecture something has to something has to give to to get us there right and yeah like right now a lot of models are like not that usable like for for uh for professionals especially right or or even for like consumer right like if you're if you're sitting there like for the best models you still have to wait like 40 seconds I don't know sometimes you have to wait two minutes three minutes like that's not really acceptable in a world where like we want everything like on demand so so yeah i think i think something needs to change yeah and probably pace off like hardware getting faster it's not enough i think if that's the case you know it'll take much longer we'll have longer timelines so i think architecture needs to get better awesome thank you guys you made a very high conviction that on generative media as a theme i think way before it was obvious.
1:01:24I think we are just at the start of I think what's going to be an explosion of generative media. And it's been really cool to hear about everything you've built from the kernel optimizations and the compiler all the way up to the workflows and what you're seeing from customers with new and old media alike. And so thank you for joining us on the show today. Thank you. Thank you so much. This was a lot of fun.
1:01:56Thank you.
From the publisher
fal is building the infrastructure layer for the generative media boom. In this episode, founders Gorkem Yurtseven, Burkay Gur, and Head of Engineering Batuhan Taskaya explain why video models present a completely different optimization problem than LLMs, one that is compute-bound, architecturally volatile, and changing every 30 days. They discuss how fal's tracing compiler, custom kernels, and globally distributed GPU fleet enable them to run more than 600 image and video models simultaneously, often faster than the labs that trained them. The team also shares what they’re seeing from the demand side: AI-native studios, personalized education, programmatic advertising, and early engagement from Hollywood. They argue that generative video is following a trajectory similar to early CGI—initial skepticism giving way to a new medium with its own workflows, aesthetics, and economic models.Hosted by Sonya Huang, Sequoia Capital




