In short
This A16Z episode is about FAL’s H3 Max and why AI video’s “next frontier” is control plus real-time speed. Guests argue that generative media needs “token market fit” (a single person can spend ~10K tokens/month productively) and that post-training + systems co-design now makes video generation dramatically faster without quality loss. They claim H3 Max Turbo can generate ~5 seconds of video in ~1.5 seconds and at ~2x lower cost; they cite ~35x speedup vs original H3. Key examples include Blender workflows using GPT Astra references, and “H3 Max Director” enabling continuous, action-controlled streams with up to ~60 minutes of continuity and ~2 minutes of memory. Notable claims: single-node serving on 8 GPUs; higher hardware utilization (30–40% to 70–80%); controllability via LoRAs (camera, lighting, lip sync, motion).
Guests
Jennifer Lee (A16Z General Partner), Gorka Mirdzevin (FAL co-founder), Batuan Tashkaya (FAL head of engineering).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Potential of H3 Max
1:07 to 3:10
Discover how H3 Max is changing the game in video generation speed and quality.
“and systems optimization made video generation significantly faster, opening up new experiences where video can run continuously, remember previous scenes, and respond to direction as it plays.”
Optimizing Video Generation Systems
3:10 to 5:18
Understand the technical optimizations that enhance video model efficiency.
“But the biggest reason why everything came together for this particular moment was because H3 was the first truly next generation video model that's open source.”
Achieving Higher Video Model Efficiency
5:18 to 7:16
Explore the compounding effects of various optimizations in video models.
“And we have been approaching that roof line more and more, especially like lately, because our entire team has been focusing on how do we get out more video pixels from a single chip as much as possible.”
Next-Generation Video Applications
7:16 to 11:03
Examine the new capabilities and applications enabled by advanced video models.
“For diffusion models, this is just essentially how do you go from running 50 steps to running something like 20 steps, right?”
Surprising Outcomes from Model Launch
11:03 to 14:00
Learn about unexpected uses and creative projects sparked by the H3 Max model.
“Like we run evals, they're like almost the same, right?”
Creative Explosion at the Company
14:00 to 18:03
Learn about the spontaneous creativity and collaboration during a fall project surge.
“Like the whole company gathered around this model and like some front-end engineers started working on like interesting applications.”
Innovative Live Streaming Experiences
18:03 to 20:49
Discover how internal projects led to viral live streaming innovations.
“So going back, we have been like very, very focused towards role models and essentially like action controlled or like, you know, action driven real time continuous streams of video.”
Advancements in Continuous Video Generation
20:49 to 23:18
Understand the breakthroughs in generating continuous action-controlled videos.
“we did this like fall live website to just like demonstrate it because it's like, people need to see how cool this is, right?”
Market Reception of H3 Max Director
23:18 to 26:57
Analyze the early adoption and popularity of the H3 Max Director video model.
“on the platform, on the file platform by like double almost, like a little more than double in terms of like volume.”
Future of Video Models and Hollywood
26:57 to 28:00
Explore the implications of video model technology on the Hollywood landscape.
“So if you were to do this even more efficient, let's call it maybe even cheaper, we would probably run different parts of the pipeline in different types of hardware.”
Show all 15 chapters
Exploring AI in Video Production
28:00 to 29:04
Learn how AI is enhancing video production workflows, particularly in Hollywood.
“So you will have people like Rohan that can stream partially of the experience from this computer, but also having like the director and the control plane more living on the...”
Achieving Controllability in AI Models
29:04 to 30:56
Discover the advancements in AI model controllability for creative professionals.
“is an extremely popular workflow for professional work.”
Innovations in Video Generation Techniques
30:56 to 33:58
Understand the latest innovations in video generation, including camera and lighting controls.
“and then it can synchronize the lips perfectly.”
Hollywood's Evolving Relationship with AI
33:58 to 36:34
Examine how Hollywood is increasingly embracing AI technologies in their workflows.
“Hollywood is our fastest growing segment.”
Future of Generative Media Conference
36:34 to 37:43
Get insights on the Generative Media Conference and its impact on the AI landscape.
“The other half of the problem was legal and data residency, things like that.”
Transcript
Automatic transcript. May contain errors.0:00Jennifer Li:Generative media is along with the coding agent market, what we call is token market fit. Everyone's waiting for a large consumer moment in AI. I believe H3 Max makes it possible.
0:14Gorkem Yurtseven:Were you surprised by the speed up and the gain you could get from post-training this model?
0:20Batuhan Taskaya:We have a version called H3 Max Turbo that's public that can generate like a five second video in like 1.5 seconds. From a cost standpoint, it's also like 2x less.
0:28Jennifer Li:people starting creating these beautiful scenes using an LLM model, GPT Astra, in Blender. And all of a sudden it unlocked the whole new workflow for Hollywood and professional people.
0:42Batuhan Taskaya:We have been very, very focused towards speed, performance, quality, and now we have a really good base model. The next month or two is going to be fully focused on...
0:50Jennifer Li:What happens when AI video becomes fast enough to generate in real time? A16Z general partner Jennifer Lee sits down with FAL co-founder Gorka Mirdzevin and head of engineering Batuan Tashkaya to discuss H3 Max and the rapidly changing generative video stack. They unpack how post-training and systems optimization made video generation significantly faster, opening up new experiences where video can run continuously, remember previous scenes, and respond to direction as it plays. But speed is only part of the story. They also discuss the push toward greater control over camera angles, lighting, characters, and motion.
1:27Jennifer Li:and why those tools could make generative video more useful for professional creative workflows.
1:35Gorkem Yurtseven:Welcome, Gorkum Batuan, to our podcast again. We did the last one last year. This is long overdue, and we have such an exciting model to talk about, which is FALSE H3 Max. The day when it came out, I was calling it, it's really in the league of its own. Like, it's so funny to see the benchmarks where you have the dot of this model on the far left or far right, and then everything else is on the other half.
1:57Jennifer Li:And that graph is actually log scale. It's actually further, but we had to fit it in. We had to do log scale.
2:04Gorkem Yurtseven:That is hilarious.
2:05Jennifer Li:The time portion, the quality is not.
2:07Gorkem Yurtseven:Yeah, for sure. The internet noticed, for sure. There are so many viral tweets about it. Like people really played around with this model. Maybe just give us the backstory of what inspired you to post train this open weight model from Minimax. And how did you get the quality and speed to where it is?
2:25Jennifer Li:First of all, the Minimax H3 model is the first truly open source, very capable, like latest generation video model out there. So even though we work with some of the other model labs to run inference for them, we never had this capability, like had the right to add this capability on top of it. So when Minimax came up with their very capable open source model that is truly last generation, can take references, like very familiar architecture to any other video model, we thought this is a great opportunity to go all in and see what we can do. And again, we did like many different things that we are going to talk about that combined gave the results that you show on the graphs.
3:10Jennifer Li:But the biggest reason why everything came together for this particular moment was because H3 was the first truly next generation video model that's open source.
3:20Gorkem Yurtseven:What is the idea, given like FAO has been known to be like a general media inference serving platform, like what is the idea to get into post-training open web model? You talk quite a bit about it in the blog of combining the system work with the model itself. Maybe talk more about the work behind that.
3:38Jennifer Li:Generative media is, I would say, along with the coding agent market, what we call is token market fit. And the way we define it is as can a single person productively spend a lot of tokens? And the amount is like 10K a month, something like that. So there is incredible amount of demand in the market to generate video, to generate many things at the same time. And a person who is doing this for their daily job, they spend in front of a computer and do this all day long. And they spend thousands of dollars, lots of tokens. And since around April, the whole industry and FAL itself, we've been compute constant.
4:21Jennifer Li:We are growing as much as we are adding compute. Like there are things we do here and there, but the whole industry has been compute constrained. And we've always been looking for efficiencies where we can relieve that a little bit so people can use this more. So that has been the idea behind everything we've been doing since April. And this just came at the right time because this makes everything maybe an order of magnitude more efficient. So it gives more compute for other models or even like more tokens can be generated using HTT Max. I think Batuan would agree on that.
4:59Batuhan Taskaya:Yeah, like just from a system-wide optimizations, which is what we have been doing for the past three, four years, you can maybe make the model 2x, 3x faster while producing the same quality, right? Because it's at the end of the day, same model, same architecture. You have the same constraints. You're just trying to optimize what you can get out of the chip itself. and there is a roof line there. And we have been approaching that roof line more and more, especially like lately, because our entire team has been focusing on how do we get out more video pixels from a single chip as much as possible.
5:31Batuhan Taskaya:And this new set of like post-training related optimizations with system slash model co-design enables us to go beyond that roof line by an order of magnitude. And like we just felt the pressure. We have been working on it on top of open source image models before the video models. We did one version with ideogram. We did one version with flux. So we have been like experimenting with how can we build post-training infrastructure to take an existing model, build kernels and systems design around it to run it very, very fast for a specialized version that can beat anything else that we would get just by running the model itself.
6:09Batuhan Taskaya:And, you know, combination of that plus just like getting a frontier video model on our hands and all this expertise, we were able to go by like an order of magnitude in terms of speed.
6:19Gorkem Yurtseven:Incredible. Let's dig into that. I may get some of the numbers wrong, but...
6:22Jennifer Li:There's like efficiency numbers, there's cost numbers, there's speed up numbers. Right. Not everything means efficiency, but it all adds up to be very efficient.
6:31Gorkem Yurtseven:Yeah, I guess what is stunning to me is there is like a magnitude lower cost and also much faster. I think it was like 35x speed up.
6:40Batuhan Taskaya:Compared to the original Minimax H3 at point.
6:42Gorkem Yurtseven:Well, at the ELO score, it didn't really sacrifice quality. So, yeah, just reveal a little more of the secret sauce behind. Is this more of the type of system work you have done? Did you have to do model architecture change? Is this system work that really brought down the cost and latency? And how about the next generation of chips like GB200 fits into the whole story?
7:05Batuhan Taskaya:It's just a compounding effect of multiple different optimization variables that we have been targeting. The first one is obviously, okay, you go from like the base model to a model that's like post-trained to be like more efficient. For diffusion models, this is just essentially how do you go from running 50 steps to running something like 20 steps, right? Like you're just like trying to optimize that pipeline. But as soon as you go from 50 steps to 20 steps, you lose quality. So you need to target in the optimization scene, okay, I want to improve the quality and then I want to apply the optimization.
7:34Batuhan Taskaya:So we have like checkpoints of this that are significantly higher quality, but obviously slower. so what we initially did was okay let's run our post-training and our L pipelines so that we can improve the model's quality and then apply the optimization stack on top of it so that the end result gets you the same quality or like even like higher quality than the original model but at the same time you're like an order of magnitude faster so most of the gains come from post-training this model to be like compatible that you can run this on like less amount of steps but on top of that you add like all the kernels and systems engineering work that you do that brings your like hardware utilization from like 30-40 percent which is like standard in like many inference workloads to like 70-80 percent and 70-80 percent on theoretical mfu which is like impossible to reach so you're essentially at the roofline of what you can get out and then these models are not just like a single oh you just give a prompt and you get a video back they're actually pipelines underneath you need to take a prompt you need to like run an llm like a very large llm to go expand that prompt to a format that the model was initially trained at generate the video in the latent space and then decode those latency back into pixels and then depending on the workload there might be an upscaling component involved so there's like multiple components and every single component by default is unoptimized there's still like lots to be gained there and like we just looked at it from a perspective of we are going to get the maximum out of every single component this made us go around llms at super high speeds right there is that component but for a different workload.
9:01Batuhan Taskaya:This is not like something like an agent decoding LLM workload where you have very high cache rates, where you have higher sessions. It's a single shot. Give a prompt, you get a prompt back and there's no caching. You're operating at low batch sizes. So there's a completely different set of optimization on the prompt expression side, completely different set of optimizations on the diffusion model, completely different set of optimizations on the VAE that you take from latents to pixels. And you just combine all of these to have an effect that compounds. From a hardware standpoint, going from something like hoppers to black walls, you see something like 2 to 3x improvement by itself.
9:33Batuhan Taskaya:But from a cost standpoint, it's pretty comparable because the cost is also in that league. So I would say it only reduces your wall clock time, but not just the efficiency itself. But it obviously helps if you want to go significantly beyond real time. If you want to generate five seconds of video in less than three, two seconds, then you need some of these latest generation hardware today to unlock that possibility.
9:55Gorkem Yurtseven:Maybe this is a detail of a question. Is the model being served on a single GPU? It is.
10:00Batuhan Taskaya:The majority of the video models today run in a single node configuration, which is eight GPUs, because once you start scaling beyond eight GPUs, the efficiency gets less and less because of the communication overhead. And existing both, like the existing Minimax H3 endpoints, as well as other video models, are probably getting served at single node configuration, same with this. It's running in parallel across eight GPUs.
10:23Gorkem Yurtseven:And do you think there will be more efficiency gains in there That you can either optimize more of the steps in between By sacrificing maybe some of the narrowed down user experiences Let's say the different type of inputs and outputs Or as you're thinking of parallelism Is there more juice to squeeze?
10:43Batuhan Taskaya:We released a turbo version of H3 Max So the initial idea was calling this H3 Turbo And we were like, we don't want to call this Turbo because the quality is like better than the original one, right? Like that's, like this needs to signify how good of an achievement it is. So we released HTT Max, but like a week later we had like, you know, like our team was like, we can run this 2x faster at like 97th percentile of quality. Like we run evals, they're like almost the same, right? Like there's still like, there's like a noticeable, like there's a small noticeable loss in quality, but we have a version called HTT Max Turbo that's public that can generate like a five second video in like 1.5 seconds, which is like insane.
11:19Batuhan Taskaya:and that's also like 2x from a cost standpoint it's also like 2x less so it depends on like how okay you are with like losing quality you know you can go down and like today these models are so cheap and so fast that I don't think people need any faster or like any cheaper like it's already like at a point where from a cost standpoint compared to the frontier itself it's an order of magnitude cheaper compared from like a speed perspective it's more than an order of magnitude faster, and just enable all the experiences. I think we would need to see, I think we would need to see what else levers that people would need.
11:58Batuhan Taskaya:But my bet today is we just need to improve quality more than the speed at these speeds. Let's fix the speed and let's try to push for quality and controllability of these models, which is what we have been pushing in the past two or three weeks.
12:13Jennifer Li:I think controllability is key. When we first did it, we did text to video and then image to video and then references came later, which adds a ton of controllability. And it's basically the default mode, how people use these models these days, references. And then we are now adding different Lora's fine tunes of the base model as well. We are working on like a lip syncing version. We are working on a different camera angle Lora, different style Lora. So again, open source adds a whole ecosystem around the model and it really, really helps.
12:50Gorkem Yurtseven:Were you surprised by the speed up and the gain you could get from post-training this model? Like, I saw it as a little bit of a surprise that like one day, I think it was a Saturday, you'll launch the model and the Sunday, people put it on Twitch and become a real-time model. Like, that's the interesting part of like when you reduce.
13:07Jennifer Li:We did evals, like we spent a ton of money doing evals on our own. I don't know, like tens of thousands of dollars even. And like the results were unbelievable. And then like the plan was to just release the model without doing external evals. And then, okay, we decided let's hold off. Let's not tell people that this is like so much faster and so much better before we have some external validation. So we waited like three, four days to all these other like eval platforms to actually run the evals. So we matched the results that we have externally as well. And that's how we launched it because, as you said, the results were a little too good to be true.
13:52Jennifer Li:And it was.
13:54Gorkem Yurtseven:I guess, well, you're taken by surprise that the real-time use case that came out of it. Or what are some examples that you think this model has, like, unlocked of the experiences the prior models couldn't?
Read the full transcript
14:05Jennifer Li:This happens at Fall once in every couple of months where like the whole company gets hold of something and the creativity just explodes and everyone is just working on a new little app or a different optimization, Laura, whatever it might. Like the whole company gathered around this model and like some front-end engineers started working on like interesting applications. We can talk about our like world model accelerator team, which is brand new. They started working on the live experience. Yeah, the RTC. The web RTC live experience. So like there were five, six different parallel little projects within the company.
14:53Jennifer Li:and like I think we broke a record on Slack that day how many messages were sent in the company because like and like we have a distributed team. We have people all around the world like mostly in San Francisco but it's like incredible when you see like the 24-hour development like when people like work 16, 17 hours and then someone else wakes up and picks up that and that went on for like three, four days and that's when we released all these projects.
15:26Gorkem Yurtseven:Take me into that. It's so interesting because you imagine a model or product launch being planned out, having all these eval vendors being ready, lined up and ship something out and then you let the world or the external users take it and then experiment and build experiences put online. Yeah, it seems like people internally who are very creative just took this job, everything they were doing, like launched experience that got really popular on Twitter. Do you want to tell us about that one?
15:58Jennifer Li:Yeah, of course. One of our engineers, Rehan, just by himself completely started streaming a live stream of continuous generations of H3 Macs from his laptop. Like he was... His computer. His computer, exactly. He was doing some like prompt tricks, trying to keep like a coherent story. And then like he started live streaming that on Twitch. In parallel, Levels.io, a famous Twitter influencer at this point, had a similar idea. And he reached out to us that he has a website ready already. He wants to like host the streaming himself and have a website that does infinite streaming. Internally also, we had another team who was working on a continuous version of HTT Max.
16:47Jennifer Li:So HTT Max is like the rehoused version and levels.io version were independent clips. It's still very fast, but the clip starts, it ends, and then you take the last frame of the clip, try to put it in the next one and try to create a continuous. You need to put some work into like the last frame. There's no memory, like the second clip doesn't really remember anything from the first clip other than the last frame. But internally, the ML team was working on a version where the transition is more seamless. There's like two minutes of memory. So like you're in a scene and when you direct the model or someone else enters the room, it actually like everyone looks at that person entering and the scene is continuous.
17:33Jennifer Li:So internally, we were working on that. And then another team was working on an experience we called File Live for the continuous version. So we had three parallel efforts going on that were all independently going viral on Twitter.
17:48Gorkem Yurtseven:And these were all like spontaneous, like you didn't plan for it at all.
17:52Jennifer Li:You didn't plan for any of them. Yes, exactly.
17:54Gorkem Yurtseven:And they just became products and experiences in the following days.
17:59Jennifer Li:But Tom, let's talk about how we made the model more continuous. That was very surprising to me because I've never seen that actually work on a video model before. Sure.
18:11Batuhan Taskaya:So going back, we have been like very, very focused towards role models and essentially like action controlled or like, you know, action driven real time continuous streams of video. And the problem till like, you know, something like HTT Max was quality was not good enough at all. It was just like, you know, it degraded a lot. It didn't remember the past before, but we built the infrastructure. We built the infrastructure that we can go stream video, have people control it in real time, being able to multiplex it to multiple people, very low latency. And at the same time, our ML team was essentially trying to take every single video model and try to apply this set of optimizations and tricks to, okay, how can we make this generate instead of a five-second video, 15-second video, 30-second video.
18:55Batuhan Taskaya:But you were always below the real-time factor where you were always like, you know, you never could generate like five seconds under five seconds. Once HTT Max unlocked it, the ML team was like, this is insane. Which are like separate teams internally. We have a research team, we have an inference team, we have an ML team. They're like, they saw this and like, this is insane. We can apply all these like set of learnings that we had in previous models where we attempted to do this, where instead of trying to generate a five second chunk, let's try to generate, you know, like a 50, like 10 second video.
19:26Batuhan Taskaya:And then the five seconds from previous one is still attended. we still remember it and like as the video goes up we can like extend that memory up to two minutes and you need to do extremely clever optimizations because attending to a two minute video is just extremely extremely compute intensive and just like it goes up exponentially because uh from like a compute standpoint so like we we did like lots of optimizations there but at the end of the state we were able to okay we can remember back to two minutes which is like generally good enough from a memory perspective and then obviously with like prompt tracks you can still like continuously see, remember more finer grain details above the two minute mark.
20:04Batuhan Taskaya:And you can essentially stream infinitely. We captured it an hour from that perspective. And then that team just like released that model under H3 Max director, which is public for people to use. And I think it's the only model that can generate like, you know, up to 60 minutes, continuous videos that is action control. You can like, you know, start with a prompt, say like there's like an office setting and someone is like, you know, working. and then like 30 seconds later, it just imagines by itself. 30 seconds later, you can like say, a woman walks in through the door. Like it can take the prompt and reflect it immediately, which is the most fun part.
20:36Jennifer Li:And the office is still the same office. The camera can pan back to the original person and the original person is still there in the same state. Yeah.
20:45Batuhan Taskaya:So, you know, we released that and it got like, we did this like fall live website to just like demonstrate it because it's like, people need to see how cool this is, right? This is a new technology. I don't think people are like really aware. And it got also like very viral immediately because we also let people vote on what the next section is. It was like, you know, like a form of - Crowd source. Crowd source. Like the chat was controlling whatever was happening, which is fun. But obviously, you know, we limited on like the options and then they could pick, oh, like a banana enters the office instead of a moon.
21:16Batuhan Taskaya:And it's like more fun. And like people start like, you know, having these. And we start adding more channels and like every channel had a concept. There's like a channel where it's like full chaos. There's a channel where it's like cartoons from like 80s. And like the model is like extremely capable and it just like remembers like so many different concepts and there's like, it has like a big, big memory from like a styles and like, you know, concepts perspective. So it just became like a very fun experience underneath.
21:41Gorkem Yurtseven:Again, like there's so many really incredible experiences coming out of this. Like H3 Max director was just another huge surprise to me. It's like, I found it interesting in the, in the gem media market that you, it's not like, you know, like language model, you have like this linear graph of like just continuously compounding on like, you know, intelligence capability and so on. Like feels like in the field you're operating in, it's always like a few months of like sort of quiet time. But like a lot of things are bubbling. But like in a very short period of time, like everything bursts, like all these things in combination come together of like the base model being good enough.
22:24Gorkem Yurtseven:like you can get the latency down to the point where you can like references yeah yeah um like get the real-time experience but also like apply controllability on top of that real-time experience like this just opens so many you know opportunities of like live experiences where like end user can control what's happening on the screen which is incredible like uh we have imagined a lot of these experiences but never been able to like really play around with it maybe just like tell us more about what you're seeing from the market of like, how are people using like the director capability? Like what are you seeing creators are creating that you haven't seen before?
23:05Gorkem Yurtseven:And what do you think that unlocks as far as, you know, what people can do with this medium?
23:10Jennifer Li:Yeah, it's been like almost three weeks since we released H3 Max and already it is the most popular video model on the platform, on the file platform by like double almost, like a little more than double in terms of like volume. So in a lot of other platforms, it's also becoming the default model that people interact with because it's so fast, so cheap. It just makes sense. If you come to a platform, this is the experience that you want to see. So in terms of like popularity and volume, it's taking over at least from our vantage point. And for Max Director, again, there has been, I don't know, tens of different versions of these live streams.
23:59Jennifer Li:Some of them are still going on and becoming more and more popular. We are trying to work with some AI IP holders, people who have like AI shows on Instagram and TikTok and train a Laura on their style and do a live version of their show. So we have a couple lined up already. So that's going to be very exciting. And like the way people, like if you talk to a creative technologist, prompting with voice has already become something that like they use all the time, like using Whisperflow or the chat GPT voice mode. And now like you can keep talking to the model and it's like almost as if it's a real director in a real movie set directing like the camera, directing people where to go.
24:53Jennifer Li:You can do that. And like our creative engineers started using these models like that. So we'll see like a lot of interesting experiences are built as we speak.
25:03Gorkem Yurtseven:Very interesting. As in like the video is playing.
25:05Jennifer Li:The video is playing, you are like talking to the video and what's being displayed changes accordingly.
25:12Gorkem Yurtseven:That's incredible. And talking about like how the memory piece holds now, like, again, this may be a technical detail, like the capability of remembering what happened in the last scene or in the last couple minutes of scene. Like, are you remembering that through like the frames, the images, or is it like through text?
25:33Batuhan Taskaya:it essentially no it's essentially like it remembers the raw video obviously very very compressed because you can't attend the fall video but it essentially knows like most of the details happened in the past two minutes from its own generations and above the two minute mark it has like think of it as like it has an evolving system prompt on top of the two minute mark from two to 60 minutes where it knows like the overall structure overall detail so it remembers like the last few scenes if you think a scene is like 15-30 seconds then it remembers like the last four to eight scenes. And then on top of that, there's like a continuously evolving, gradually evolving system prompt that like keeps remembering the overall coherence of the world.
26:16Jennifer Li:Everyone's waiting for a large consumer moment in AI. Now it's like good enough and cheap enough that like a truly novel social AI experience can be built on top of it.
26:31Gorkem Yurtseven:Maybe let's talk more about the economic side of this Like what is the I guess one just like talking about serving cost For like same minutes of video With H3 Max And how has it changed your thinking around Like your footprint of like inventory of chips Like how do you want to have like different steps of experiences Serving to the end user
26:59Jennifer Li:But I mentioned this a little bit like everyone talks about how complex the next generation LLMs are, but video models are actually very complex as well because the pipeline has different components and sometimes they require different hardware configuration for efficiency, things like that. So if you were to do this even more efficient, let's call it maybe even cheaper, we would probably run different parts of the pipeline in different types of hardware. another interesting thing would be to run it on consumer hardware for people to run it in their own machines at home like optimizations don't translate 100 % but translate somewhat close to that and then we can do extra work to translate more of it so doing these optimizations in different types of hardware and combining the pipeline in a way that it's even more efficient.
27:59Jennifer Li:I think that's what we are going to do in the next coming weeks.
28:04Gorkem Yurtseven:Amazing. So you will have people like Rohan that can stream partially of the experience from this computer, but also having like the director and the control plane more living on the...
28:16Jennifer Li:Exactly, yeah.
28:17Gorkem Yurtseven:On the cloud. Makes sense. So we talk about all the consumer experiences this model could unlock. And it seems like Botwan is happy with all the efficiency, like, squeeze out of the GPUs. Now we're talking more about how do we, like, improve quality and controllability of these models so that, like, you know, the high end of the market, the Hollywood creators, directors, can take this to the next level. I saw some demos. Coincidentally, like, you know, this model came out the same week prior to Astra. People were combining the Blender experience with H3 Max from FAL. like talk about how it's going to impact the Hollywood world.
29:00Jennifer Li:Using Blender with one of these AI models together is an extremely popular workflow for professional work. Basically, you render a low resolution of your scene, what you want to do using Blender, previous like non-AI technology. And then once you add that video as a reference to an AI model, you basically get close to 100 % controllability. And this is an incredibly popular workflow for VFX artists, people who are doing this professionally, because they want to get exactly what they put into the model. And as you mentioned, a week after we launched H3 Max, people starting creating generating these beautiful scenes using an llm model gpt astra in in blender and all of a sudden it unlocked a whole new pipeline using an llm to create a blender scene and then passing that to the h3 max model or or any video model but it works very well with h3 max because it's extremely fast and you can like try many things all at once in parallel and that unlocked the whole new workflow for Hollywood and professional people and it gets you to like close to 100 % controllability.
30:23Batuhan Taskaya:As I said, we have been very, very focused towards speed, performance, quality and now we have a really good base model. I think the next month or two is going to be fully focused on, okay, how much controllability we can add to these models so that professionals at studios, professionals who want to actually produce content that fits their use cases perfectly can leverage these models. The team has been working on an amazing lip synchronization model where you can just supply the audio, you can supply a video or an image reference, and then it can synchronize the lips perfectly. Same with motion controls.
31:03Batuhan Taskaya:You can just take a motion of someone dancing and apply it to your AI-generated character, and it fits perfectly. And this is like, you can like get these results with like basic prompting and you're going to get like 80%, 90 % reliability. What we are targeting is like 99.9 % reliability in the outputs so that you can actually trust the model did every single aspect of this generation perfectly. And that's like what we've been pushing. One big launch that we had last week was the camera controls, which is essentially you can direct where the camera is going within the video perfectly to the degree.
31:40Gorkem Yurtseven:And this is by like describing in the prompt or like generating the...
31:44Batuhan Taskaya:You just essentially like underneath you give a JSON of like, I want camera at like 0, 0, 0 at T0. I want camera at like 90 degrees angle at T1. Like you essentially supply a structured description of where your camera needs to be at any point in time. And then the model is like perfectly conditioned to regard it as like the only source of truth and it doesn't like hallucinate like where the camera should go. And it's just like, you can essentially reconstruct 3D scenes from a single input, like because the model itself is a very good video model, but at the same time, you know, it's like perfectly adhered to the camera itself.
32:21Gorkem Yurtseven:And this is because the base model itself already has the understanding of the camera angle that you can...
32:27Batuhan Taskaya:It doesn't respect it. It just under, like, you need to tune them all. You need to tune the model to a significant degree. And this is like what enables like at large scale, post-training infrastructure. We now have the infrastructure to take H3 Max, add any capability to it. Same applies for any new model. There's a new video model. We essentially spend most of the time building it as an infrastructure than just one-off training runs so that we can build services around this for not just open-source models, but for frontier close-source models as well. Because we see in the market this is the biggest gap, is just how controllable these models are.
33:03Batuhan Taskaya:first we start with text to video where you put a prompt you get a video back, it was good but you never could describe the perfect character for you, and we had image to video where you used an image editing model and then generated the first scene and then the model was obviously much more fitting, but you still couldn't say, oh I want this new character appear at second tree, you need to put it to the first frame, or you can prompt it, but it was never perfect, and then we added reference to video where you can provide an initial starting frame and you can also provide, I want these characters with these voices.
33:34Batuhan Taskaya:Like, you know, that's also like a big unlock where you can essentially say, this is the voice for this character. And now like, you know, we are adding, oh, within the scene, I want camera to look at this degree at like T0, I want camera to look at this degree at like T3. And then we are adding lighting controls where you essentially say where the light is coming from. These are all compounding on top of each other. Like we just have the unified infrastructure to just apply this to any model at this point.
33:57Jennifer Li:That's incredible. Hollywood is our fastest growing segment. And there's a lot of noise about how AI might disrupt Hollywood, but Hollywood usage was non-existent a year ago. And in the past year, it grew and now it's the fastest growing segment. like Amazon MGM Studios in their conference, they released their NARA tool. It's mostly backed by file infrastructure behind the scenes. And we are seeing incredible, incredible pull coming from Hollywood. And exactly what they need, these like small point solutions rather than generating everything from scratch. They want to be able to extend the video a little bit.
34:40Jennifer Li:They want to be able to change the camera controls. They want to change the lighting and someone has to build these solutions for them. What Hollywood needs and what the creators actually need and what the research labs are working on, there's a little bit of a disconnect there. And we believe we can come in and do these little post-training projects to close that gap. Because we work with all the Hollywood studios and we hear from them what they need. And these are exactly the things they need, these small point solutions that actually make them more efficient, push out more video, and AI can actually close that gap very nicely.
35:22Gorkem Yurtseven:Maybe say in a little bit different way. Like we have been staring at this problem for the last three years as well. Like we see like companies trying to like build a, you know, movie director, like video model, like by either pre-trained or post-trained on the video side. But what I'm hearing is like different people expressing the way they want the output to come out very differently. Consumers talk about it and then like write the prompt and generate the results. very differently from a Hollywood director, which is obvious, right? Like professionals want to talk about like, you know, these camera angles.
35:58Gorkem Yurtseven:They want to talk about the lighting. Like you sort of have built a library or like a collection of post-training, like I would call it data and toolkits that can apply these, any model that you can like grab the weights on so that they are adapted to like a different audience where they can express their creativity in a bit different fashion to control the model when it unlocks a lot of capability underneath.
36:29Jennifer Li:And half the problem was capabilities of these models. We are solving that. The other half of the problem was legal and data residency, things like that. We made a ton of progress there as well. We now have a system of people can apply with their own IP and we unlock their own IP in the models, we are going to grow that and that's going to be a very powerful thing we do with Hollywood Studios. Also, we now have Seadance US hosted as well. We already had previously other Chinese models. Seadance was the missing part. Every Hollywood studio wanted us to have it US hosted. Now that's available. So there are no obstacles in front of these Hollywood studios now, everything is ready.
37:19Jennifer Li:And we believe they are going to 10x, 100x their AI usage in the coming months.
37:24Gorkem Yurtseven:It's such an exciting world for movie lovers, consumers, people who consume a lot of video and creative content.
37:33Jennifer Li:And we have our conference, Generative Media Conference next week. This is our second time we are doing it. Last year, it was mostly consumer AI. There were like maybe a couple Hollywood executives here and there just curious about it. And now it's dominated by studios. New AI studios who are like offshoots of the bigger studios trying to do like only AI shows. But also like the biggest of the Hollywood studios are also there because now they have big plans integrating AI into their workflows, into their existing systems. so you can see the change in the attendance of the conference as well.
38:14Gorkem Yurtseven:That's awesome. Well, for the audience, check out the content coming out of the Gen Media conference. It's going to be very, very exciting. And thank you so much for Kim and Batuan coming on to our show. It's a super exciting time for Gen Media.
38:27Jennifer Li:Thank you. Thanks for listening to this episode of the A16Z podcast. If you liked this episode, be sure to like, comment, subscribe, leave us a rating or review, and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X at A16Z and subscribe to our Substack at a16z.substack.com. Thanks again for listening, and I'll see you in the next episode. As a reminder, the content here is for informational purposes only. It should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund.
39:09Jennifer Li:Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see A16Z.com forward slash disclosures.
From the publisher
a16z General Partner Jennifer Li sits down with fal co-founder Gorkem Yurtseven and Head of Engineering Batuhan Taskaya to discuss what changes when generative video becomes fast enough to run in real time.
They unpack the technical work behind H3 Max, fal’s post-trained version of MiniMax’s open-weight video model, and how combining model post-training with systems and hardware optimization significantly reduced generation time while maintaining quality. That speed has enabled experiments with continuous video, including streams that can remember previous scenes and respond to new directions while they’re running.
They also discuss why the next challenge may be less about speed and more about control, from camera movement and lighting to characters, motion, and lip sync. And they explore what those capabilities could mean for professional creative workflows, where artists and studios need predictable tools rather than simply generating a video from a prompt.
Resources:
Follow Gorkem Yurtseven on X: https://x.com/gorkem
Follow Batuhan Taskaya on X: https://x.com/isidentical
Learn more about fal: https://fal.ai
Follow Jennifer Li on X: https://x.com/JenniferHli
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
