In short
Latent Space: The AI Engineer Podcast - Episode Summary
Episode Title
A Technical History of Generative Media
Episode Description In this episode, hosts engage with Gorkem and Batuhan from Fal.ai, a leading generative media inference provider that has recently secured a $125 million Series C funding and surpassed $100 million ARR. The discussion traverses the evolution of Fal.ai from dbt pipelines to diffusion model inference, highlighting significant models that have impacted image generation and speculating on the future of AI in video production.
---
Key Points
Introductions
- Hosts: Podcast hosts welcome Gorkem and Batuhan from Fal.ai.
- Background: Fal.ai transitioned from a Python runtime in the cloud to generative media optimization.
Company Overview
- Scale of Fal.ai:
- 2 million developers on the platform.
- Hosting around 350 models (image, video, audio).
- Recent funding and revenue milestones achieved.
Major AI Models and Their Impact
- Key Models:
- Stable Diffusion 1.5 marked a pivotal change in generative media.
- SDXL led to significant revenue growth.
- Flux models introduced commercial viability.
- Market Trends:
- Rise of generative video models through partnerships with companies like Luma Labs and Google DeepMind.
Technical Discussions
- Inference Engine Architecture:
- Focus on CUDA optimization and kernel reusability to enhance performance.
- Performance and Latency Considerations:
- Importance of optimizing latency for user engagement and generating results quickly.
Model Partnerships
- Closed Source Collaboration:
- Partnerships with closed-source model developers to improve performance.
Infrastructure and Scalability
- GPU Infrastructure:
- Discussion on serverless GPU architecture and optimization techniques.
- Prioritization of H100s and Blackwell optimization.
Architectural Trends
- Shift from Distillation to MMDiTs:
- Exploring new model architectures and their implications on performance.
Generative Video Models
- Market Positioning:
- Fal.ai's growth driven by the demand for generative video capabilities, including advertising applications.
- Revenue Streams:
- Discussion on the monetization strategies for open models versus closed models.
Future Trends in AI and Generative Media
- Potential Developments:
- Future of generative media, including the rise of open-source contributions.
- Data Collection Initiatives:
- Importance of collecting data in the evolving landscape of image and video models.
Hiring and Team Structure
- Hiring Initiatives:
- Fal.ai is expanding its team, seeking talented professionals in various technical roles.
---
Key Takeaways
- Fal.ai's Evolution: Transitioned from a basic cloud service into a sophisticated generative media platform.
- Performance Optimization: Continuous improvement in latency and performance is crucial for user retention and engagement.
- Market Dynamics: The generative media space is rapidly evolving, with partnerships and model releases significantly impacting competition.
- Future Opportunities: Strong potential for upcoming advancements in generative video and the importance of aligning with open-source initiatives.
- Hiring Focus: Emphasis on building a robust team to support ongoing innovations in the generative AI space.
Closing Remarks
- The discussion highlights the importance of adaptability within the generative media landscape and the potential for growth as technology and models continue to evolve.
For further insights and updates, listeners are encouraged to follow Fal.ai's developments in the generative media sector.
Full show notes available at [Latent Space](https://latent.space).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:27Hey everyone, welcome to the Layden Space Podcast. look at my own notes, but you were optimizing run times. Yeah, it was first like we were building a future store and then we took a step back and then we decided to build a Python runtime in the cloud. And that evolved into an inference system that evolved into what Fall is today, which is a generative media platform. So we optimize inference for image and video models and audio models, but we do a lot more. We try to own this whole generated media space for developers, basically. Yeah, amazing. And we can talk about that journey. I wanted to also introduce Batuhan.
1:06We're newer to each other, but you've come to some of my meetups before you're head of engineering. Yeah, I lead engineering here at FAL. You know, glad to be here. And what's your journey? I met Burkha in 2021 when they were just starting a company. And like just before the seed round, you know, Burkha and Joachim, we met online. We're both Turkish. So I think that was the connection. we just met and then they said, oh, why don't you join us? And I was one of the core developers of Python language. So I had a really good experience with developer tools around the Python language. So I started coming here to build the Python cloud, which evolved into this inference engine and the generated media cloud that we are building today.
1:44And now you spend less time with Python and more time with, I don't know, Kuda. Custom kernels. Exactly. Yeah, I remember the dbt file when the modern data stack was out. Can you guys maybe just give a quick sense of the scale of FAL? So you just raised$125 million. Seriously, we can talk about how I passed on one of your early rounds. We can go through that. How many of the developers, how many models do you serve, and maybe any other cool numbers? Yeah, we have around 2 million developers on the platform. And for the longest time, we required GitHub login. It recently changed. So I'm assuming everyone who has a GitHub account as a developer.
2:24and we have around 350 models in the platform. These are mostly image, video and audio models. It used to be only image and then we added audio and the space evolved into video as well. And yeah, that's pretty much the scale. We just announced our C-Round and we've been growing a lot in the past year and it's still continuous. Yeah, you had a very nice C-RC party. Yeah, thank you. And you guys are over 100 million in revenue, right? This is not, you know, just developers kind of kicking the tires. That's correct. Yeah. That's great. When you say 250 models, I think, what percentage of all the models that you could serve is that?
3:06Because, you know, especially in... There is an infinite amount of iTunes post-trained versions of these models. We are trying to serve the models that fix a gap, you know, that fill a gap in the stack. So we don't add a model that's like significantly worse in any aspect compared to other models that we have. We are trying to bring unique models that solve a customer's needs. So that's like, these are 350 models, you know, there's like 20, 30 text image models, but like one of them excels in logo generation, another one excels in human face generation. So like every model has a unique personality, but if a model is like significantly worse in all aspects, we don't add that to the platform.
3:39So there's like infinite amount of models that we can add. And do you rely on your own evals or just like what the community is? We mainly rely on our own evals as well as, you know, like we are in the community. So we also follow the community very well to see what is going to be the thing that's going to be in the next generation of apps. So if we think something, like we have a good intuition, if we think something is going to pop up, we just add it. Yeah. To my knowledge, you haven't published your own evals, right? No, we don't publish. It's internal. And then the community is Reddit, Twitter?
4:05Twitter, Reddit, you know, Hugging Face, seeing how popular the models are in Hugging Face and other demos. Okay. I just want to give people a sense of where to get this info. The best part of the job is the day of a model release, the adrenaline rush that comes with it, the whole team trying to scramble something together and release it. And it happens every week. Every week is exciting. Can we do maybe a brief history of like the models that were like the biggest spikes maybe in usage? You know, you kind of, I think everybody knows stable diffusion, you know, and then you have maybe like the flux models and then you have Black Forest Labs.
4:39You have like all these different milestones. Theory-wise, I think the biggest, like the initial hit was Stable Division 1.5, which is when we actually pivoted into this new paradigm of fall, Generator Media Cloud. We started hosting it. We noticed like we had the serverless runtime and everyone was running the Stable Division 1.5 by themselves. And we noticed it's terrible for utilization and they are not optimizing it. So let's just offer an optimized version of this that's ready for API to be scaled and doesn't require people to deploy Python code because we want product engineers to start using it.
5:06We want mobile engineers to start using it. So we started offering Stable Deficit 1.5. It was very popular. The fine tunes around it was very popular. Stable Deficit 2.1 came. It was a bit of a flop. So it didn't like, you know, got that much attention. And then SDXL came, which was like the first major model that brought like our first million in revenue. If you consider that. And with SDXL, obviously like the small fine tuning ecosystem also like tried exploded. People started fine tuning their faces, their objects, whatever. And generations with this, like LORAS started to become very popular.
5:38And then after Stable Deficient XL, there was like a bit of a quietness around it. You know, ST3, there was like some drama around it. And the team at Stability left to start Black Forest Labs, which released Flux models. And that was the first model to reach the barrier of commercially usable, you know, enterprise-ready great models. Where in the first month of Flux models, we reached from like 2 million to 10 million in revenue. It was like a big jump. Next month, we were at 20. Like, it just started going from there. And then video models started, came around. You know, we partnered with Luma Labs.
6:07We partnered with other video model companies in China. We partnered with Kling, Kuaisho, Minimax. And with these models, it created another market segment. That was a big jump. And the final biggest thing was VO3, where it actually created this usable text-to-video component, where before text-to-video was a very boring, soundless video where you would get enjoyment out. Whereas now it's such a great experience. You can create all these memes that you're seeing online, all these ads. So that was another big jump for us, partnering with Google DeepMind 4.0.3. Yeah, actually, that's a really good history of generative media, that SoundBite.
6:46So I wanted to double-click on that because obviously we can dive. I think everyone's interested in video, but there's a whole history of the image side that I wanted to cover first. Just definitely wanted to start with was just the decision to pivot. I think I just want to double-click on that. You know, it's not a trivial decision, but obviously the right one. At the time, I would say like a lot of people were hosting stable diffusion, right? So it wasn't obvious that you can just build an entire company around effectively just specializing in diffusion inference. What gave you the confidence?
7:14What were the debates back and forth? Yeah, a couple of decisions we had to make there. We could have evolved the company into more towards GPU orchestration. And like essentially we had this Python runtime. We were running it on top of GPUs like that could have been the company. but we saw every single person, every single company who are using what we had like a little SDK to run Python code on GPUs, they were doing the same thing. They were deploying a stable diffusion application, maybe using some LORAs on top of it, different versions of it, in-painting, out-painting, things like that. I mean, it was very wasteful.
7:51We decided, okay, this needs to be an API where you actually optimize the inference process and everyone benefit from it. And like, Like you can run it multi-tenant, you know, the utilization is much higher than. So that was the decision number one. And then obviously after stable diffusion, I think like four or five months later, LAMA2 came out and there was a decision point again. You could do language models. Exactly. And a lot of the inference providers at the time, there were maybe a couple of them and they all went all in on language models. And we decided language models, hosting language models is not a good business.
8:28At the time, we thought, OK, we are going to be competing against OpenAI and Anthropic and all these labs. Turned out that it was even worse because the killer application of language models is search. And you are competing against Google at the end. And Google can basically give this for free if they can because it's so important for them. And, you know, it threatens their business right away. And with Imogen Video Models, it was a net new market. We weren't going against any incumbent. We weren't trying to get market share from someone much bigger than us. And we liked that aspect of it. We thought we could be a leader here.
9:09It was a niche market, but it was very fast growing. So we chose to be a leader or to be a leader in this fast growing niche market rather than trying to go against Google or OpenAI or Anthropic. So that was the decision we made. And it turns out it's a good one because we are able to define the market we are in and educate the people and grow with it. And so far, it's been growing fast enough that we were able to build a whole company around it. Yeah. And I think you noted at AIE that, you know, now there's a generative media track. Generative media specialist investors. Thank you for calling it generative media, by the way.
9:46Yeah, I mean, obviously, it's a thing and people care about it. And I do think it's going to change the economy. And as a creative person, I think I also wonder what's going to do for us. Just so I want to keep it technical and keep thinking about the pivot, because I think it's still one of the most interesting pivots I've seen in the AI era. You were not CUDA kernel specialist at the time, right? I come from a compilers background. So my job was optimizing Python bytecode interpreter to make stuff faster, which is performance engineering. And yes, I don't think at the time there was that many CUDA kernel specialists either.
10:18so it's like we were at like the right time you know it was like uh the actually like the space was actually so so much worse than what we have today where like the running basic like stable diffusion 1.5 was like a unit with convolutions and the convolution performance on a1 was like you're getting like 30 percent of the gpu power if you just use raw touch because no one cared about it so there was like so many low-hanging fruits that we started to pick up and started optimizing and it kind of evolved evolved evolved right now it's like much more competitive space with like nvd has like a 50 person 100 person kernel team that's writing kernels you're competing against that at the time no one really cared about so it was like a good uh new field for us to go thrive and there's no community effort like a vlm not exactly when these models were first released like no one in the world has ran them in production like it just didn't exist it's like a research output exactly yeah it was stability yeah you had your maybe local gpu maybe you had like a single GPU that you rented from the cloud.
11:13And basically this was a research interest rather than a product interest. And no one at Meta, no one at Google had run this in production. So we also thought this is a good time to start that company around this and actually spend time optimizing it as much as we can, because if we can get millions of people to use this, there's a lot of economical value to be created there. Can you talk a bit about how much of a performance boost you get? Because I know when I met you guys, you were about a million of revenue. You were like, well, we're writing all these custom kernels. And maybe part of it is like, okay, how many kernels can you actually write?
11:48Sure. You know, as you support all these different models. Like, what's kind of like the breadth of them? Like, are you writing kernels that you can reuse across models? Like, how much work do you have to do on a per model basis? It's really evolved in the past three years. You know, when we first started, there was a single model, Stable Division 1.5. So all of our kernel efforts were, how do we make Stable Diffusion 1.5 as fast as possible? You know, you go from like 10 seconds with PyTorch. At the time, there was not even like a Torch Compile, Torch Inductor, whatever. So you were going from 10 seconds to maybe like two seconds on the same GPU.
12:19And like we started with that. The next thing, like, you know, with adding more models, you know, like Stable Diffusion XL was a different architecture. Pixar was a different architecture. All these like different architectures started coming around. we said let's build an inference engine which is what we call a collection of kernels parallelization utilities diffusion caching methods quantization all that stuff combined into one package and so we built this inference engine uh the same time pytorch 2.0 was released with torch inductor and torch dynamo to do like torch compile which is like essentially a way to trace the execution of of your neural net and generate writing kernels there are few like that are more efficient and i'm a big sucker for just-in-time compilers i used to work at pipi like a just-in-time compiler for Python.
13:00And we said, this is a great idea. Let's apply this, but a more specialized, more vertical way for diffusion models. At the time, it was units. Now it's diffusion transformers, which are significantly different than your autoregressive transformers in terms of the profiles of how compute-bound it is, what sort of the kernels are taking the majority of the time, if you're doing bidirectional attention or causal attention. So we started doing that. And now what we have today is an inference engine that's applicable. that gets you like 70-80 % of majority of the models on diffusion transformers and we still have like a lot of custom kernels for a lot of models to squeeze out because they're still small.
13:35Every model wants to make an architectural difference. You guys see this on like, you know, even for stuff like, you know, QAn, DeepSeek, whatever. People want like, even if we know an architecture is the best, they want to tweak it a little bit just to make sure, oh, we're releasing something cool. So we saw this and then like for that, we have to write like custom kernels for custom RMS norms that people are doing or whatever, like stuff like that. So we have a decent amount of kernels, like over 100 of custom kernels. This doesn't include the auto-generated ones. You know, we have templates of kernels that generates like, you know, for thousands of different shapes, problems, spaces, whatever.
14:05But like, if you consider those, you know, like we have tens of thousands of kernels, obviously at runtime that we are running and dispatching. But that's pretty much like the depth and breadth of it. And on average, a model on file runs 10x faster than I would self-host it. Like if I just take stable diffusion, right, and I put it. I know that this might be a bigger discussion point. Do we consider speed as a mod? this comes to that the existing open source industry goes so fast where you know if you go to this might have been true three years ago now PyTorch is already very very good for H1Ns right?
14:34What about B200s? When you use PyTorch with B200s black belt chips you're not getting the best performance so our main objective and our main goal is for whatever GPU type you're using these diffusion models we're going to extract the best performance at any point in time it could be 1.5x it could be 3x it could be 5x for certain models it could be 10x it would be a bit of an unfair thing to say, oh, we're going to make everything magically 10x faster. No one in the world can do that. We are lucky that this is a moving target and open source community, everyone catches up. But at the same time, new chips come out, new architectures are released.
15:07So we are always ahead of what's possible, but then they catch up, but we have to stay ahead of it. And that's how we can create differentiation because it's a moving target, because there's so much going on. Whenever something new comes up, we are the first one to optimize it. first one to adapt our inference engine to it so at that time the fastest place to run it that helps with margins things like that but eventually people do catch up i think it's very hard to create this differentiation over long term if there is no new architectures if there's no new chips but luckily there is all the time yeah yeah and i think with image specifically you cannot stream a response so to speak so when you have a language model it's like you're kind of bound by like how quickly you can read so even with like rock it's like it's impressive to show a thousand token a second but it's like i'm not reading that fast right so it can go slower yeah versus with images it's like you just need to see it that's why mid-journey now has like the draft mode for example it just gives you this like very low quality resolution yeah but at least you can see whether or not it's going in the right direction how much of that is actually true for like your customers what do they care about the most like it's latency that like what's the range of latency that matters.
16:20Yeah, latency is really important. One of our customers actually did a very extensive A-B test of like they on purposely slow down latency on file to see how it impacts their metrics. And it had a huge part in it. And it's almost like page load time. When the page load slower, you know, you make less money. I think Amazon famously did a very big test on this it's it's very similar like when when the user asks for an image and you know iterating on it if it's slower to create they are less engaged they create fewer number of images and and things like that yeah it's the same learning that amazon has like every you know 10 improvement in speed yeah exactly it's uh the the elasticity is high the other thing i wanted to also dive into you know like um i putting my a little bit of the investor hat on one of the reasons for file success is kind of with not within your control which is how when and how people release release open uh models for defer diffusion which like at the time it was just stability and like there wasn't there was no chinese uh you know output i mean what we did have other image models but they were not great yeah and like and so you you made a bet on when it was just like wasn't super obvious but then the other thing is what you're touching on the diffusion workload is very different from the language workload and the language workload is being super optimized whereas diffusion is not.
17:40So you just like had kind of no competition for a while, which is fantastic for you. 100%. And like the open source, we benefit a lot from it, obviously. But like in the past six months, a year, we started working with some of the closed source model developers, as well, like behind the scenes, helping them. But they're not sending you their weights. They do. They are. Wow. What do you have to give security, like guarantee them? VR, any cloud provider, like what do they think in like AWS? or Google Cloud or these Neo classes. There's like 50 Neo classes, right? Like we are not that different from any other cloud provider.
18:13And this is why we package the inference engine in a way that, you know, they can self-service and get 80%, 90 % of the performance. So they don't even have to show us like their code. They deploy to our, we have our own cloud platform where our inference engine is available only in that platform. So they can tap into that when they deploy their code and their model weights to us. And we don't really have to look at it. If they want to collaborate with us, which some companies did in the past, where we would just essentially, we have performance engineers acting as forward deployed engineers on their behalf and writing custom kernels for them okay have you disclosed who you're doing this for we we disclosed play hd play ai that was one of those we have like four different companies four major video companies that we are doing this with and one image company that i don't think we disclosed yeah as you can imagine it's a little sensitive for them so yeah i would say like so like replicate started serving VO3 models and we were like, okay, are you just wrapping their APIs or something?
19:07And I think it's not obvious how much integration there is going on and how much it's on your infra or your tech. Just to be honest, some of that is happening too. VO3, I think everyone is... It's just an API wrapper. You have a dedicated pool that you can serve to your customers with different speed SLA guarantees, whatever. That's how it would work for something like VO3. But your objective is to be one-stop shop, but then also you can do inference better than some of these other like phd google is obviously hard but like with other vendors like our goal is like helping them run inference because these are research labs that doesn't necessarily invest heavily on inference optimizations scaling up infrastructure that's like another challenge that we can talk about like at launch days like some of these models like in their website they just like explode and file api is working fine because they deploy them we can scale up to like thousands of gpus instantly so there's that aspect to when we pitch this value prop as well as a distribution that we bring to them, it's a no-brainer for them to just deploy their model to file and use it for both the file marketplace as well as their own distribution channels.
20:07Yeah, a couple of follow-up questions. Just on PlayHT, just because you mentioned it, music, audio, is that a different workload than normal diffusion or is that the same? I can't really comment on their architecture, but some of them are autoregressive models in the open source world. Some of them are autoregressive. Some of them are diffusion-based. So there's notorious ones that's known for diffusion. As you guys can guess, one of the biggest companies So it's similar workloads, but at the end of the day, our inference engine is very versatile, and our performance team is very versatile. With PlayHD, we had very deep collaborations where we had three engineers at some point helping them optimize their inference process as well as infrastructures to get them ADMS, end-to-end, time-to-first audio chunk, which is a very impressive thing for real-time text-to-speech workloads.
20:54And then the other known hard problem is serverless GPUs, which is a thing that everyone has chased a lot and many people have failed. What can you say about what you've done there to make it happen? So for example, Modo has been talking a lot about their GPU snapshotting. But I imagine it's like a stack of technologies in order to achieve the scaling. It is a stack of technologies. The biggest problem with serverless GPUs is are you just repping another, if you have a Kubernetes deployment, are you just repping it and giving people access? or are we actually like multi-cloud? Do you manage your own orchestration chain?
21:25Do you manage your own container runtime? Do you manage like all this like stack? And in our case, like we started with a Kubernetes version when we were just doing it for ourselves. And Kubernetes version at Google Cloud was fine in 2022 when we wanted to get eight A100s. But when we wanted to go to like thousands of A100s, it's not going to work. It's a terrible position to be bound by a single cloud. So right now we work with six cloud providers and we have 24 different data centers in four different countries. and we now do like long-term data center releases as well to manage like the sum of the hardware chain ourselves.
21:55In this world, like we had to build our orchestration layer, we had to build our own distributed file system, we had to build our own container runtimes, all the stack to make sure that the call starts are extremely, extremely fast, which is one of the things when you're scaling up, as well as handle actual scale where, you know, like we are managing over like 10 ,000 plus H100 decolons today. Yeah, and a CDN for caching. CDN, that's like outside of the serverless infrastructure, but like CDN, content moderation systems, like all of these like all consist of the platform. Like there's like so many.
22:23We also do moderation. We also offer content moderation services to the foundation of all companies for them to like moderate their inputs and outputs. Yeah, I see. I see. As a separate product. Yes. From a GPU perspective, do you always need to be on the latest? You keep mentioning H100s. Majority of our workloads are in H100s because price per price, it makes sense. But like Blackwell is obvious. Like we have five people dedicated to writing Blackwell kernels right now to make sure we can like, Like, because theoretically it looks good, right? Like flops dollar wise, it makes sense. But can you reach that actual flops?
22:53No. So we have a dedicated team that's like working with NVIDIA directly to write custom kernels for Blackwell for diffusion transformers to get to the point where it makes per dollar make sense. And then we would start with our own workloads as well as some of our foundational companies. We would ask them, oh, if you want to migrate to Blackwell, here's an inference stack that already works. We are at that point where we should be the ones pushing the boundaries on Blackwells because no one else is doing this work. And maybe it doesn't make sense economically right now, price perf wise, but we know it can.
23:23So we are like working towards maybe like a couple of months away from that point. And then whenever it does, we'll probably switch as many workloads to Blackwells as possible. Just to be super crazy, when does it make sense to just work on an ASIC? I don't think it does. That's like honest opinion. This is like one of the most controversial topics, right? is all these ASICs a great idea? If you're memory bandwidth bound and you can put all of ASICs, is it economically viable at that point? I don't know. But the summits around these cheap designs, you see, okay, what is the overhead of an NVIDIA GAM instruction, right?
24:00It's like 16%. So you're essentially buying a matrix multiplication machine. So it doesn't really make sense to specialize it that much. And some of the B300s are going to have a better softmax instruction that gets like 1.5x, whatever. And like, that might be one way where NVIDIA gets like, you know, better performance out of the majority, like for the majority of workloads, which is like attention heavy stuff. I think it might make sense for like NVIDIA to like add more specialized stuff. But like for us, I don't think it will ever make sense to build ASICs. Just thinking about from first principles that the diffusion workload is very different, but also obviously there's still a lot of changes in the architecture that you need to just do general purpose.
24:40We don't have a single model where we are trying to optimize. We are trying to do it for the newest, the best, like always. The flexibility is therefore really important. I was going to show, I'm going to pull up the Quen MMDIT where there's like this dual streaming thing, which I think SD3 had it. Yes, SD3 Flux. Yeah. Is that the standard model now? MMDIT, so it's also a controversial topic, you know, scaling rectified flow of transformers paper, the SD3 paper came up with this architecture. And then one of our research team, actually, like Simo Ryo, he's like our head of research. he found out that just using MMDITs are inefficient.
25:13You need to mix them. And now there's controversial opinions. Movie Jam paper were saying, oh, MMDIT is complete, unnecessary. You can just use a single stream DIT, whatever. So there's controversial opinions happening in terms of architecture changes, which I understand because everyone wants to do a different architecture. No one wants to do the same architecture because it's late. Otherwise, it's just a matter of compute and data. And these researchers don't feel proud that their model is an output of data and compute. They want to make a novel research change. So I think the architecture is going to keep changing until this paradigm of researchers changing stuff for the sake of changing finishes.
Read the full transcript
25:47I'll talk about a couple other architectural things just to keep it bounded with this topic. The distillation was a thing for a while. SDXL Lightning, you guys did fantastic demos of TL Draw, which we've also had on podcast. Fantastic episode. What happened to those things? How come they're not popular anymore? I think it makes for a good demo. You know, you could build real-time applications. You could build these drawing applications, things like that. But I don't think people could build applications that have user retention long term. People couldn't really build useful things with it. Let me play out what I thought was going to happen.
26:22And then you tell me why it didn't happen. Which is consistency models for drafting. It's like you use your hand to draw things and it creates the draft. Then you upscale with a real model. But that's it. Why can't it be a two-stage process instead of one stage? Yeah. And I think one thing that happened is Flux. That generation of models were not good at image-to-image when it first came out. So you need a good image-to-image model to be able to draw. Maybe it needs to be revisited around this time with some of the editing models, maybe. Like image-to-image and control-less. Flux, like STXL or control-less were very popular, where people were used to do this stuff like sketch-to-image, whatever.
27:03And with Fluxer, I think people cared less about it. One thing that I keep thinking about this is like, is this true for LLMs too? I always default to Cloud 4.1 Opus, right? Even if it's slower than Sonate, it's just like, I know I'm going to get the best quality. Exactly, that's what's happening here. Yeah, it seems like what's happening here as well. Okay, anyway, as a creator, I want fast, quick drafts and then I can refine, right? So I don't know why it didn't happen. More true for video models, right? It used to be like five minutes, four minutes for a single five second generation. Now it's mostly under a minute, but you want 10 second, five second generation.
27:40And then because the workflows of like creatives when they're working with it, they generate a ton of videos and then like pick one and then create a story around it. So when you watch these people actually generate videos, they generate like hundreds at a time and they have to like kind of sit around and wait and then like iterate on it. like the faster speeds mean a lot for for creators yeah it does the other thing i wanted to also briefly touch on before we go back to the main topics is the autoregressive models which you mentioned right like obviously i honestly i still think gemini is underrated because they were first and then but then obviously openly i did the 4.0 image gen and that was a huge thing i actually even wonder if there was a panic for you guys because obviously it's like this is soda image gen and like no one else has it it's not it's not open source you've passed through those eras so many times you know like when stopped worrying about it yeah your camera's like good stories around dolly google yeah i mean i talk about this how like when the lead two first came out as like okay open ai is so far ahead of anyone else like it's impossible for mid-journey and then people caught up within months and then stable diffusion was even maybe better or just as good as the Lee, like a couple of months later, and it was open source.
28:58So like a year later, same thing happened with Sora. Like they put out those videos and that time around, like, I think we were excited because now that people see that it's possible, like this is like actually doable, researchers get motivated and they see the hype. They see that this is possible. So they work on it. And within a couple of months, we had maybe not Sora level, but much better video models. Now we have video models that are much better than Sora. So whenever we see someone actually pushing the frontier, it's a reason for excitement because now that's possible. Other people are just going to do it within a couple of months.
29:36So we don't panic anymore. Is the fact that Entropic doesn't have an image generation model tell you anything about what the larger labs care about? It tells more about Entropic's own like personality than like the in general, like what other labs, because If you look at XAI, if you look at Meta, if you look at OpenAI, if you look at Google, they all have really good image models. Yeah, like Google, in their last announcement, they used the word generative media, by the way, which was a proud moment for us. It's a win. And a lot. And, you know, they focused on generative media as much as their new LLM models.
30:11So some labs definitely care about it, and some labs, it's not a priority for them. Look at XAI. They keep pushing like image. You're like AI slob. Yeah, I know. I know, it's crazy. waifus. In levels of interactivity, you have images, you have video, now you have Genie, this kind of like more world model. You have kind of like gaming applications of that. How far are we from like FAL getting a lot of traffic on those models? Like, is it mostly experimental today in open source? Obviously Genie is impressive, but like, it's a Google model, you know? I have a very optimistic take on this and that may be like a normal outcome.
30:44I think at worst we are going to have very capable video models that come out of world models, right? It's going to be a very controllable video model and the use cases will be similar to what video models are. You're going to create content, but you're able to control the camera angles. You're able to control the video model a lot better than what you can do today. At worst, we are going to get that from world models. And at best, I think it's very hard for anyone to predict what's going to happen. Yeah. Movies and games, like it's going to be something in the middle where you can be part of the whole movie universe that's going to be playable.
31:24So it's boundless possibilities. What's going to happen at best and how affordable is it going to be? Is this ever going to reach mainstream adoption? We'll see all that, but it's definitely technically incredibly exciting and impressive what's coming out of these labs. yeah i need to find the paper again but there was this study on like um video models and like image generation understanding physics yeah and like it could like predict the orbit of a planet but then when i actually had it draw the gravitational forces it was like completely wrong yeah you know and so i think like that's my thing with world models like i understand the creator application which is like you can create consistent world but i don't know if like the other side of people that are like hey these are like the best way to like simulate the world and like get intelligence and things like that.
32:10There's so much optimism around it too, because whenever you talk to someone who's working on robotics, they're bottlenecked by the amount of data they have. And from all these, you know, past three years of AI innovation, we've seen that whenever there's an abundance of data, that type of models actually like improve a lot. And you see, so like robotics, we expect something similar. Whenever they figure out this data problem, those models are going to get better as well. So that's why people are so optimistic. Okay, maybe this solves the robotics data problem. And it's, yeah, boundless opportunities there.
32:47And regarding the example you mentioned about gravitational forces, I think this is still the same problem as, oh, LMS can't do 9.9 plus, like 9.0. Yes, it can. You just need to train it with more data, you need to have a better tokenizer, it is the reason, whatever. It's just a matter of data scale and the underlying fundamental architectures but like i don't think it's going to change that much we just like we're going to put thousand x more data thousand x more compute and we'll get like the best physics simulators and i think this should be possible with the existing signals coming from the data just to double click on video stuff as well yeah uh you had a great slide in at aie where you're like currently 18 of files revenue comes from video models and it might be a that was february so But now it's probably over 80.
33:33No, 50%. 50? Okay. It's like over 50. Yeah, yeah. 100%. Wow. Okay. I guess editing models brought some life into the image as well. So like both of them grew, but yeah, video grew faster. Video is like pretty, pretty significant. And one of the main drivers is open source models where, you know, in February there was Hun Yuan video. I think that was like pretty good. There was Mochi from Genmo, but like the quality still wasn't there. and one from Alibaba was like insanely good model. And they released a newer version of this like a month ago, I think, or like a couple of weeks ago. And now it's getting so, so popular.
34:10And like, we can run this model, like for like 480p, like the draft mode version, we can run it like five seconds under five seconds. So people can have like instant feedback loop. And then when they want to go to 720p, like full resolution is just like 20 seconds. And we were planning to bring it down to 10 seconds. Yeah, that's amazing. And I want to double click on that. For a while, I was kind of bearish on Alibaba because they kept releasing papers with very cherry-picked. And it was like, okay, we're on GitHub. And then you go to GitHub, it's a readme. I mean, you can't see something really change.
34:42They've been releasing new image models, new video models. No, we haven't talked to them. But it seems like, and now there's competing teams inside Alibaba. One is a really good image model, but they released Quent as a competing Quent's image model. So we think one's image model is actually very, very good. If you'll do one with single frame instead of 81 frames or whatever, you get a really good text image model out of it. And this is just because of the pure amount of data that you put from the videos. So now there's Alibaba has two of the really good models from their lab. And then there's smaller labs in China that you might not hear about, but like Stepfun, they released an image editing model.
35:20Hydream, Vivo, there's all these small labs they're releasing. because I don't think training these image or editing models are that expensive. And video models might be like slightly more expensive. My guess is like training these costs like a couple million dollars, which is not that much, especially, you know, like, you know, they're probably backed by some sort of entity, you know, like other than Alibaba, you know, like there's like Stefan, whatever they probably raise a really good amount of money. So training these models will bring you a lot of attention and it's more attention that you would get releasing a subpar LLM because like LLM space has so much more competition.
35:53So just training this for like a million dollars a video model and then releasing it, I think that brings you a lot of attention. It is a hack. When you look at Hugging Face, let's look right now, I'm sure the top models are image models. Like it's, Koen image editing probably is... Probably up there, yeah. Probably up there. Yeah, number one. Number one. Hoonyeon Gamecraft. Is number three? Number four. Then you got Gemma, 270 million. ByteDance had some image stuff. ByteDance has an open source, but they have a really good team. Seed, that's their new lab. they're working on C-Dream C-Dance OmniHuman like stuff like that we have a good partnership going with them to hopefully have their models hosted in the US as well and the idea is I think that the team that they were able to assemble is very good and it's coming from their previous researches whatever ByteDance was doing really good open source stuff on like SDXL Lightning they released SDXL Lightning paper Animative Lightning so I'm pretty hopeful about them yeah first of all hopefully they don't reach out to you when they launch they just drop and then you have to rush.
36:52At this point, people reach out to us because we are the market leader. So they just reach out to us for getting distribution. Yeah, there's a Chinese platform that they always launch on first, which I forget the name of it, but you have to... We also get day zero launch with a majority of these models. So basically, I think the question is always like, you know, you are the ones making money. Stability did not make money from stable diffusion. I think the thing that Black Forest Labs did was very, very interesting in this aspect. They released three different models. Apache 2 licensed extremely distilled model, which is good for...
37:26Dev. Chanel. This is the Chanel version. This is like for four-step generations for like lower quality stuff. They released a dev model with a non-commercial license where their inference partners are, you know, like you're paying a revenue share and this is like a very good way. And then there's a pro version where you can like collaborate for hosting it. And like the revenue share is obviously different for that as well. This is, I think, a very smart choice for labs whose whole premise is releasing models. But if you're a company that is doing a product in the side, you don't necessarily need to make money from the open source models.
37:57You're doing it for getting researchers, hiring people, getting distribution, whatever. So it really depends on the company's goals. For Alibaba's case, they don't care if one model is hosted in their API. It doesn't touch Alibaba's top-line revenue, whatever one makes. So for them, it's a no-brainer to release it and get attention and maybe like get some leads to their Alibaba cloud offerings. But like in general for Black Forest Labs or companies like that, I think it's like a smart move to release like a distilled version as like fully open source and less distilled or like the actual model as non-commercial.
38:29And then partner with inference companies and stuff like that. What's the distribution of usage? So like is 80 % of your revenue like five models? Or are people really using like the long tail of all these open models outside of like the initial launch? I think like there is some power law, but not as much as you would think. And it keeps changing. That's the other part. Like it's not like only a single model that's being used a lot, like month to month. It changes a lot. This summer has been crazy. There have been like just countless amount of new video models, new image editing models. Like the leader kept changing week over week even.
39:07but if you like take a step back and look which models are being used people want to use either the best most expensive video model or they want to use like a cost efficient like good but cheap enough video model so those two models are usually like used a lot and whatever those models are it changes week over week and yeah the one good example is like fox context was released you know like on late may and kwen imuja that was released like two weeks ago and now it's like topping out context dev. It's insane how quickly these stuff transition just because there's a better quality model. And that's the value prop.
39:45You don't have to set up the infrastructure to manage flux context dev. As soon as like co-image is available, you can just switch to that with file. I mean, it seems to me that if some models are open and some models you have to pay a revenue share, you ideally want to move the people off the revenue share models into the open models, right? Like what's that dynamic? It's all price-to-date. I'm also thinking like, okay, we'll do whatever our customers are going to be successful with. We are still like early enough that these small calculations, I don't think it matters. I'd rather people actually go to production and build products with it and be successful rather than like, okay, 20 % here, 10 % there.
40:22I mean, you're doing a hundred million in revenue. Cool. I'll just ask a few more questions we had around just like the, how people really use this. Okay. I'll ask this super obvious question. Yeah. How much is not safe for work? Almost not. Negligible, yeah. You don't moderate everything. Moderation is optional, right? Moderation is optional to a level where illegal content is moderated. And we also track the non-illegal content NSFW moderation. And we haven't seen more than 1%. The models themselves are actually not generating that type of content. Some model providers, especially if you look at Blackforce Labs models, the models are not for...
41:00It's incapable of generating because it's like... Or like it's an yield in a way that is prevented. And the majority of our customer base, if you look revenue-wise, it's like enterprises are more on the higher level of stuff where some of them might be like user space, mobile applications. But for the last six months, nine months, we've been transitioning more and more to enterprise where it's less of a need for them. So what are those enterprises doing? You know, apart from building a general purpose chatbot that can generate images, maybe Canva would be a good use case. But my imagination is a bit limited beyond that.
41:32advertising seems to be absolutely like growing and if you think about it it fits very well and let's talk about video advertising so like i keep repeating this but some companies talk about oh we are gonna change hollywood filmmaking is gonna be revolutionized like i don't think it's that interesting like how many movies do you watch a year like maybe 20 25 movies how many movies in the theater you watch three four at most so if there are like thousands of movies that a year like people won't be able to watch all of these movies like there's just it's still it's a max quality exactly exactly and with advertising it's the exact opposite more content there is the different like you know ways you can create ads there is always economical value attached to it So you can create unlimited number of ads, unlimited different versions of it.
42:24And more personalized it is, more, you know, economical value there is behind it. So with ads, it fits really well to this type of technology because there's no limit what you can create. I'll tell you a side comment about a Silicon Valley trend I'm seeing, which I cannot explain, which is that all these YC startups and all these, they're spending between$10 ,000 to$70 ,000 per launch video. Yeah. In the age of generative video. like they're hiring you know actual creative directors hiring a studio hiring actors uh i was in one of them and like do you need all that when you have generative video like i think clearly roy started talking about generating videos i don't know if you guys know pj ace i think he's like the absolute killer for this stuff he uh launched like a is it a super bowl ad or something he did like a basketball playoff yeah yeah yeah nba finals right yeah he also did our like series b announcement video like we're like pretty close with him and like that it's insane that like what he's able to come up with and how viral it goes or like you know these videos where you spend like hundreds of thousands of dollars right you know like it's just like you just need to create viral content and these general media models are the best way to do it and we're still at the infancy of this right like obviously it might not be professional quality i'd still think like you know human in the loop like mixed like you know content is the way for today but in six months who would know like 12 months i think like 80 of this is going to be generated like we were like we were watching super bowl and we were like saying oh how much of this video is like ai generated it looks like ai generated like it could be right like you can't tell so like i i think at some point you know we're gonna have like 80 90 percent all that's it reminds me of uh i think um who's who's the guy uh fofer from replicate obviously he is the best inspiration for all these workflows he like lay overlaid some kind of nba realistic sort of laura on top of game footage.
44:13So you could play like NBA 2K, but it looks like a real video. I saw that, yeah. Yeah, I was like, what the hell? It's pretty cool. So maybe that's the other part of my question and I wanted to get into Comfy UI, which is how much LoRa serving is going on, right? How much custom? A lot. You know, okay. Is it like majority? That's one of the reasons. Not majority, but like, if it's like 30%, is it like the majority? And everyone trains their own LoRa's or you pick it off of like a LoRa marketplace? That's why open source works very well with image and video models, because you tap into this big LoRa ecosystem, everyone.
44:50I've never seen a closed source model that can create a good LoRa ecosystem. It just basically doesn't exist. Maybe there is mid-journey SREFs, but I don't know if you can consider them LoRa's. SREFs are just seeds, right? Conditioning, let's call it another condition, like a prompt. Yeah, and then only the open source models have these rich LoRa ecosystems, and it's extremely, extremely popular. But even for the oldest models, it brings a new life. When you see these cool LoRa's, we have still a lot of people using STXL with their own LoRa's because they're happy with the quality. It's fast enough, it's cheap enough, it's amazing.
45:27These models are not single-shutable like the language models. Even the editing models, GPT-Image 1 or FluxContext, whatever, QAnImage. if you put your face or like if you put like multiple people whatever you can't get the quality like it's gonna be like 90 there but if you train for like 1000 steps with like six to 20 images you're getting like 99 accuracy with like like the like we worked a lot on like fine tuning the right hyper parameters like writing like distributed trainers distribute optimizer stuff and with those like people can train like their lores under 30 seconds now on the platform run an inference with them in the same job and get like 99 accuracy for the same face character which is one of the biggest challenges that maybe more on the enterprise side they're facing and less on the consumer side.
46:11If you're creating AI slope, you don't really care who it looks like. But if you're actually doing a product ad, you want it to exactly look like the product. Every single pixel on the products, banner, whatever, you want it to look like that. So you train a LaCroix Laura with 20 images and then after that, you have almost a pixel-perfect model. All right, we had to train a Laura for every guest. Then we can make thumbnails. I actually think that's a very good application because it's a nice way to inject brand, but not in a strict style. And we are just entering like post training on video models.
46:45And what's that that's going to mean? Because we didn't have a good base video model that it made sense. But now we have like companies really investing into like post training on 1.2.2 or Hunyuan and like creating lip sync models on top of it, creating like different video effects, camera angles. Seems like there's a lot of possibilities with like creative data sets that people can do. I think in the next six months to a year, we are going to have a lot of like companies that are just built on post-training of open source video models. Wow. Let's talk about pipelines. So we are Confianonymous on the podcast.
47:22ConfiUI is kind of like this community that like, if you're into it, you love it. If you don't know about it, you kind of underestimate in a way, but people create all kind of crazy workflows. One, have you thought about doing pipelines? Obviously you host all the models. There's kind of this question. We do have a pipeline product called File Workflows and you can chain models together, but it's obviously less flexible than Confu where you can only chain different models outputs and not the intermediary stuff. In Confu UI, you can access the latents from one model and then pass it to a latent upscale or whatever.
47:51In our case, it's more limited, but we have a workflows product and we have a serverless Confu product where people can bring their own Confu UI workflow and run it as an API with just like posting the workflow and inputs. And let the models be served by you. Yes. So is that a bullish thing? Is that going to be commoditized by big models? The thing that we saw is as the models get better, like Confu UI was a much bigger thing or like, you know, relatively much bigger thing in two years ago, like a year ago, when the models were like, you know, one of the biggest Confu UI use case was you were like generating an ST1.5 or STXL image and you were fixing six fingers situation, you know, like you were fixing like the resolution or whatever, upscaling.
48:27Now that the models are actually so good, the Confi UI workflows are actually getting simpler for image side. For video, it's still very crazy. Like if you look at some video workflows, it's like there's like 50 nodes, whatever that you're processing. So I think this is still a matter of like how good the models are and how much extra stuff you need to do around it for majority of use cases. For artistic use cases, you still are doing like a lot of stuff. And that's like something we want to support. But like that, we don't see that happening at super scale, you know, like super scale. There's not companies that are spending like$10 million plus on running this as an API.
48:56So that doesn't seem to be happening yet, just because it's like a bit inefficient and the more existing, like it's more reliable to use an existing model than patching together 50 different things because you don't know when it's going to not work. Yeah. But it feels like for things like ads, you want to do, you know, one step, which is like maybe generate the backdrop. One step is like adding copy. Are you saying the models are so good? Chaining of models happening for sure. But I think ConfUI did very well is you can also play with the pieces of the model. Yeah, you're busy saying that's all.
49:25Yeah. So like chaining of the models, like that's what the file workflow product does. It's basically calls many different APIs back to back or in parallel and then creates the result at the end, I think. And it's very popular. We have like enterprise adoption from it, like from very big names. Yeah. Amazing. I was just going to go into the broader topics. The first thing that comes to mind was requests for startups. If you're not working on fail, but you see a lot of things in the ecosystem, right? What's the most obvious thing that people should be working on? more model companies. Go raise more money and train models.
49:56That's obviously good for FAIL. And host them on FAIL. If you're not interested in training models, but if other people are trained, that's amazing. Go raise more money. There's so much money. Or like scale AI for Imogen video models. Data collection, more prepared data sets for video models, effects, different camera angles. Everyone seems to be reinventing the wheel when it comes to collecting that data. I think it's a great opportunity for someone to come in and do this at scale. So it's really interesting because I think this is what Together AI did with Red Pajama, is they actually built a data set for language models to help people create more open language models that they can serve.
50:37So at some point, it actually might make sense for you guys to do that. The image data is a bit more finicky situation in terms of copyright stuff, whatever, but it's an interesting area. Do it in Japan. I think it requires focus it requires like this needs to be like one connected thing to what Gerkem said is image slash video RL that's an unknown unknown for us say more I can't it's like what does it look like can you RL a video model to be a world model you can right like if you consider it like it's essentially the world models are RL video models where you condition it for like you know moving around so what are the use cases for RL image and video models i don't know but that's like if i wasn't working as well that would be something that like that might be fun to explore yeah and is this is specifically for editing because the rl is for the reward is the edit or that's the thing that's that's what you should look for right like what what is the reward function like what is the interesting reward functions that you can apply on top of these base models i see interesting okay got it actually i was really asking about like if i were to build a foul wrapper startup like on top of that because Because you guys are very low level, which is fantastic.
51:47But I also want to give our listeners some ideas, if they're not going to work at that level. I think I'm going to say it again, advertising. There's so much opportunity there. And everyone's still trying to create these horizontal applications where any creative can come in and do something. But a lot more targeted to specific industries, a lot more targeted to different kinds of ad networks. There's a lot of potential there. good and then requests for models you obviously want more open models that's good for you but like any like specialization in the models like i think image image editing was a huge unlock which i didn't foresee until this year yeah where we're like oh obviously we're gonna no one's yeah i didn't like guess it was gonna go this big it's like it's insane like how popular it became and then like everyone started like catching up i was part of a group at neurops that we meet at new reps every year and like they were talking about this at the last new reps and so like it's in the air but you have to be at the researcher level and like everyone moved to video left image behind a little bit so there was like a little vacuum of research on on image but luckily people people saw it that it works very well and then they went back to it it's so much cheaper to train image models like right now like if you if you look if you want to train a sota image model i don't think it's going to cost more than a million dollars it's extremely cheap it's like a matter of data, engineering effort, cleaning.
53:08I think it's a function of data set. Image models are really, really, really affected by the data set that you use. I think one obvious thing that there's a gap in the market is like VO3 is very expensive and the way the reason why people like it is conversation, right? If you can create maybe a smaller, cheaper video model that is less capable but can do conversation and sound very well, I think there's definitely one open source one that we saw was multi-talk it was a post-trained version of one and it's like really really good for conversations but it lost the ability to generalize you know it's only like talking faces at some point versus like vo3 it could generalize and it could do scenes and whatever so i think that there needs to be like some middle ground between these two between talking faces and like you know extremely generalized video models where it's like much cheaper to run but at the same time you know you get like this conversation because it's very memetic you know people there's infinite amount of memes that you can post with this infinite amount of ads that you can do with this but you don't see a world in which you have a video model and then you have a separate maybe audio only model that can generate the audio for this is the worst question right yeah do you stitch together a whole bunch of things or do you better let people did that before v03 but what v03 gets very well is like the timing almost like totally you ask for a joke and you know the delivery and the timing and to laugh and like you know waiting right before like the joke drops like all of that is so perfectly timed i don't think when you do it separately you get it it also matches the human accent sound to the face that is talking right it's like it's an unknown challenge for like other text-to-speech models like you like it feels very natural is vo3 the best text-to-speech model it is also one of the best it is so good like i don't think any model can do what what is the emotional But I would say the counter argument is that we dub movies.
54:57So there's already, you know, obviously you can. It is also the best lip sync model. Like VO3 has the best, most accurate lip sync because it's generating very natively. There's really good lip sync models. I think they're like 95 % there, but VO3 is like 100 % there, like 99%. To me, this is like the single most bearish thing about workflows, right? And ComfyUI and all this stuff, like, because just wait for a bigger model. It's just pure bit of lesson. Yeah, we love ComfyUI. but i mean obviously like when the technology doesn't exist yet you have to stitch together things yeah so a request for engineers yeah i mean what i'm sure you're hiring right you're just 125 million hiring like we just recently crossed 40 people but like for like wow for like three months ago we were like 20 so like for the last three months we have been actually accelerating you know best kernel engineers best infrastructure engineers best product engineers best ML engineers.
55:49If you're the best at what you do, just come join us. I think it doesn't really matter what you do. Just like, we're just hiring the best self right now. Even on the go-to-market side, we are hiring account executives, customer success managers. Because we work with very large enterprises, we got to grow that side of the company as well. Yeah. On the engineering side specifically, how do you think about how many people you need? There's a whole question of like lean AI, it's like, you know, coding agent. Our performance team is like seven people. I think seven people like focusing on performance.
56:17always like some of their, some overlap with our applied ML team, which is taking these models, productionizing them, exposing new capabilities, building fine tuners. So it's like, and then helping customers adapt these models. So that team, I think we can scale to like double, triple the amount because like there's infinite amount of models and like, you know, it's better because like we're going to have more customers with more proprietary models. So just like helping them optimize it. It's just like a really good function that we have. That team scales very well because there's always like independent work that can be done.
56:45Oh, okay. So these three people are working on this new model, trying to optimize that. And it's completely independent from trying to optimize this other model. So we've been hiring a lot for that Applied ML team. Our aim for a team, you know, we're probably going to keep it lean in contrast to the Applied ML team. And the product team may be like, we want to build more higher level components where people can directly integrate to their applications. Because that's even like now. Just SDK or? SDKs, but think about with components. imagine you're an e-commerce website designer and you're not really the best component designer.
57:17So here's a virtual try-on component that you can put to your app. Stuff like that, more higher level components. And this is also coming from the fact that wipe coding has been very, very insane. We see significant... Revenue-wise, it's very small. But we see a significant amount of user adoption just coming from people who are... Just from looking at our support tickets. Maybe they need more support, but there's a lot of people who are coding these applications without that much expertise in the product building. So we want to provide them more guardrail experiences where they can integrate much easier without messing with all the other lower-level components.
57:52That's really nice care of developer experience. So crack low-level engineers. And crack high-level. Well, yeah, crack go-to-market people, crack whatever. I'm always trying to refine the definition of crack. Both of you, you lead the technical side of fail. Like, what's a really hard technical problem that if someone has the solution, they should talk to you immediately? Maybe that's the way to frame it. Write a sparse attention kernel with FB8 on Blackwell and tell, you know, like, if you can do that, come join us. We already have like a good base. Hired on the spot. Hired on the spot, you know, like stuff like that.
58:27I really like picking like all these, like some of these applied ML people, like we just picked them from discords who are working on these sort of jarrington media who are like already interested. We really have a high culture bar too where everyone in the team loves jarrington media. They're obsessed with it. They would have done this if this wasn't their job. We have this great composition. It's not a prerequisite, but it's just naturally happened where we hire these people from Discord, Twitter, Hugging Face, one of our Applied ML engineers had the number one top Hugging Face space with creative workflows, whatever.
58:58So we hired a person who was training LORAs on FAL just because they were training and posting cool LORAs. Just do cool stuff and we'll find you or you can reach out to us. That's the master builder is what I've been calling this person. Why not make it more explicit? So if I go on your career's website, right? It's like apply to Mel Engineer. It just kind of looks like any job description. I feel like there's like this question of like... It does. That's why we have to do a podcast. But I think it's not just about file. I think in general... It is more like if you know, you know, which I know is not the best way.
59:30You know, people know about file already. So it's like we haven't really cared that much, but you're absolutely right. Like we should make it more explicit. If I look at like George Rots, like on TinyGrad, you have these boundaries. It's like, hey, we'll just, if you can solve this, you should probably work here. Like, do you see? I'm adding the bounty today. Right? It's like, this seems like, hey, look, if you can write this kernel, it's like, yeah, you'll just get hired. It is also, but like one thing that we saw, even with like, there's a lot of people who are just like wipe coding stuff and reviewing those.
59:58Like, there's a limited amount of people who can review those, right? Like, so like, how can you tell it's like not a shitty kernel versus like a good kernel? Well, but then you're spending the time interviewing too, right? Yeah, so we have first line of defense with our recruiters, whatever. So there's trade-offs, but I absolutely agree. Maybe we should have a kernel bench version that you can upload your kernel, automatically evaluate the stability, performance, whatever. And then if you do, you get our email unlocked, whatever. Especially email for you. But yeah, great ideas. Come join us. Deal this.
1:00:31Awesome, guys. anything else parting thoughts yeah i love your rant so this was great yeah i'm happy to run but one is a podcast star yeah no congrats on all your success um i should also say it's fun to do karaoke with you guys yes like let's do it again both extremely technical but also like a fun crew that like and i think it's pretty hard to and rare to to see so thank you to see the good guys win awesome guys awesome
1:01:02Thank you.
From the publisher
Today we are joined by Gorkem and Batuhan from Fal.ai, the fastest growing generative media inference provider. They recently raised a $125M Series C and crossed $100M ARR. We covered how they pivoted from dbt pipelines to diffusion models inference, what were the models that really changed the trajectory of image generation, and the future of AI videos. Enjoy!
00:00 - Introductions
04:58 - History of Major AI Models and Their Impact on Fal.ai
07:06 - Pivoting to Generative Media and Strategic Business Decisions
10:46 - Technical discussion on CUDA optimization and kernel development
12:42 - Inference Engine Architecture and Kernel Reusability
14:59 - Performance Gains and Latency Trade-offs
15:50 - Discussion of model latency importance and performance optimization
17:56 - Importance of Latency and User Engagement
18:46 - Impact of Open Source Model Releases and Competitive Advantage
19:00 - Partnerships with closed source model developers
20:06 - Collaborations with Closed-Source Model Providers
21:28 - Serving Audio Models and Infrastructure Scalability
22:29 - Serverless GPU infrastructure and technical stack
23:52 - GPU Prioritization: H100s and Blackwell Optimization
25:00 - Discussion on ASICs vs. General Purpose GPUs
26:10 - Architectural Trends: MMDiTs and Model Innovation
27:35 - Rise and Decline of Distillation and Consistency Models
28:15 - Draft Mode and Streaming in Image Generation Workflows
29:46 - Generative Video Models and the Role of Latency
30:14 - Auto-Regressive Image Models and Industry Reactions
31:35 - Discussion of OpenAI's Sora and competition in video generation
34:44 - World Models and Creative Applications in Games and Movies
35:27 - Video Models’ Revenue Share and Open-Source Contributions
36:40 - Rise of Chinese Labs and Partnerships
38:03 - Top Trending Models on Hugging Face and ByteDance's Role
39:29 - Monetization Strategies for Open Models
40:48 - Usage Distribution and Model Turnover on FAL
42:11 - Revenue Share vs. Open Model Usage Optimization
42:47 - Moderation and NSFW Content on the Platform
44:03 - Advertising as a key use case for generative media
45:37 - Generative Video in Startup Marketing and Virality
46:56 - LoRA Usage and Fine-Tuning Popularity
47:17 - LoRA ecosystem and fine-tuning discussion
49:25 - Post-Training of Video Models and Future of Fine-Tuning
50:21 - ComfyUI Pipelines and Workflow Complexity
52:31 - Requests for startups and future opportunities in the space
53:33 - Data Collection and RedPajama-Style Initiatives for Media Models
53:46 - RL for Image and Video Models: Unknown Potential
55:11 - Requests for Models: Editing and Conversational Video Models
57:12 - VO3 Capabilities: Lip Sync, TTS, and Timing
58:23 - Bitter Lesson and the Future of Model Workflows
58:44 - FAL's hiring approach and team structure
59:29 - Team Structure and Scaling Applied ML and Performance Teams
1:01:41 - Developer Experience Tools and Low-Code/No-Code Integration
1:03:04 - Improving Hiring Process with Public Challenges and Benchmarks
1:04:02 - Closing Remarks and Culture at FAL




