In short
LTX’s open video “world model” LTX 2.5 for total beginners—how to prompt, use multimodal inputs (text, keyframes, audio, video/image references), and integrate into creative workflows via LTX Explorer and API.
Key claims
LTX 2.5 is faster than 2.3 (about 10-second video in under 7 seconds), improves pixel quality via better keyframes and a new decoder, and supports native HDR for VFX pipelines.
Guests
Daniel Berkovitz (LTX Chief Product Officer; 10+ years turning creative tech into usable products) and Alan Yarin (LTX VP of Product; former ML researcher focused on both model mechanics and user workflows).
Notable examples
cinematic rhino charging in NYC; Dalmatian running with a red camper van; audio-to-video using a voice sample; “retake” morph cuts to stitch podcast clips seamlessly; 2D animation in-betweening in Adobe Animate; robotics startup fine-tuning for robot actions (e.g., making a sandwich); avatar/real-time conversational agents; VFX integration without replacing artists (on-prem/IP customization).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to LTX and Model 2.5
1:04 to 2:26
Discussion about LTX and the features of the new LTX 2.5 model.
“Aref's brand radar shows how your brand peers in AI answers, what sources influence those other recommendations, and where competitors are winning visibility.”
Technical Improvements in LTX 2.5
2:26 to 4:34
Exploration of the enhancements in speed and quality in LTX 2.5 compared to previous versions.
“But you know, like two and a half years later, we're here.”
Challenges in Product Development
4:34 to 7:13
Insight into the challenges of finalizing and releasing the new model.
“Yeah, it's like when you build a product, you want it to be something you can touch and tangible.”
Collaboration with VFX Studios
7:13 to 10:01
Discussion on how LTX is working with VFX studios and the impact of AI in filmmaking.
“So you mentioned big video production companies are using this.”
Cultural Shifts Towards AI in Hollywood
10:01 to 12:00
Exploration of changing perceptions of AI in the film industry and opportunities it creates.
“For us, it's incredible because sometimes we see AIs like this fun thing, funny thing that you can do for friends and stuff.”
Practical Applications of LTX Technology
12:00 to 14:01
Examples of how different industries are integrating LTX's technology into their workflows.
“So let's talk about how people are using it.”
Exploring 2D Keyframing Tools
14:01 to 15:17
Learn about the tools and partnerships used for 2D keyframing in AI animation.
“So if, for example, for the 2D keyframing one, how are they actually, what tool are they using?”
Crafting Effective Prompts for AI Models
15:18 to 17:51
Understand how to create effective prompts to improve AI responses and outputs.
“Yeah, I think that it's all about managing to describe what you want.”
Using Audio and Visuals as Guidance
17:52 to 20:54
Discover how audio and visuals can enhance the guidance provided to AI models.
“So either with a video condition or audio condition or multiple keyframes or anything that can help the model understand your intent in a way that is not textual.”
Creating Training Videos with AI
20:55 to 25:12
Learn how AI can be used to create training videos and onboarding tutorials for businesses.
“There's so many hours of my face on YouTube.”
Show all 22 chapters
Demonstrating LTX Features and Capabilities
25:13 to 28:00
Explore the features of LTX, including prompt structuring and multi-modal generation.
“I'm going to link that in the chat if anyone wants to follow along as we're doing this.”
Exploring Audio to Video Transformation
28:00 to 29:22
Learn about the capabilities of converting audio to video using AI.
“So this way, everything remains the same beside this change.”
Engaging with Audience Questions
29:22 to 32:29
Discover how user inquiries shape the discussion and feature demonstrations.
“Usually we don't add any prompt to audio to video, to be honest, because the audio condition is so strong.”
Innovations in Video Editing
32:29 to 36:58
Understand new features in video editing that enhance content creation.
“We prepared another kind of fun use case and actually a really powerful one.”
Audio Usage in AI Generation
36:58 to 42:00
Explore the role of audio in AI-generated content and its applications.
“obviously seamless there's no cut there um yeah i'll be dang yeah so this is really cool The audio stuff excites me too.”
Leveraging Audio in AI Character Animation
42:00 to 46:00
Learn how audio can guide character performance in AI animation.
“So I think like adding an audio as a condition, even if it's not like full strength and you use it like as a way, weaker guide, I think it really helps.”
Understanding Clip Length for AI Video Generation
46:00 to 48:00
Discover the optimal clip lengths for generating video content with AI.
“And the way that I'm handling the script side is like considering each line in the script as a different shot.”
Splicing Clips to Create Coherent Scenes
48:00 to 51:18
Explore techniques for splicing multiple AI-generated clips into a seamless scene.
“So, I mean, I oversimplified it a little bit.”
Using LTX for Creative Freedom in Video Production
51:18 to 56:00
Learn about the flexibility and options available in LTX for video creation.
“If we generate five shots, for example, of about 30 seconds, can we append them in your app?”
Exploring Community Resources for Video Prompting
56:00 to 57:29
Learn how to leverage community resources for video prompting and model usage.
“with a sample prompt and a sample asset.”
Appreciation for Model Development
57:30 to 57:57
Hear insights on the hard work behind the model and its community support.
“I really appreciate you taking the time to join us.”
Technical Questions on Compatibility
57:58 to 58:22
Discussion regarding the model's compatibility with different hardware.
“You certainly have enough gigabytes there, my friend.”
Transcript
Automatic transcript. May contain errors.0:00everybody what's up welcome humans uh to the neuron live we'll give folks a couple minutes here to to trickle in but we're we're excited to be here and excited you're here uh because we have some cool guests with us today that show off really neat product you're going to want to check out yeah let me let me give a quick intro as people are piling in so we're we're joined here by the team from ltx so if you don't know ltx makes really awesome open weight video models uh with us today is daniel berkovitz who is ltx's chief product officer spent more than a decade turning complex creative tech into products people can actually use appreciate that and alan yar who is ltx's vp of product and he started as an ml researcher before moving into product giving him a strong view of both how the models work and how people use them great to have you guys yeah absolutely thank you thank you for having us yeah really excited to talk about this just just a quick second of housekeeping today's live stream is brought to you by a refs ai search is rewriting the rules of brand discovery.
1:05Aref's brand radar shows how your brand peers in AI answers, what sources influence those other recommendations, and where competitors are winning visibility. Monitor ChatGPT, Google AI overviews, Gemini, Perplexity, Copilot, and more, all from a single dashboard. Go check it out at the link in the chat. Also, as we get started here, if you haven't just yet, please click the like button, hit subscribe up above. It'll help get some legs on this video and get us out there to more people. and we really appreciate it. Now back to video. All right. Real quick, so guys, why don't you tell us a little bit about LTX and LTX 2.5, the new model you just created, for people who's never heard of it, and what does it do?
1:52Yeah, absolutely. So LTX is an open, multi-model, world model. You know, sometimes people are using the term video model. Maybe later on we can talk about like why we're calling it a world model. It's open, like you said, so it's open in the way, it's open source. This is like really something that from the very beginning we thought is a key for success and for people to actually use it. Yeah, it's a 22 billion parameters model that does like multimodal audio and video and also keyframes along the way. Yeah. Very cool, very cool. question how did you build it wow that's a very good question actually i mean um like if you take a step one step back like the company uh because in the company like um our mother company lightfix we did other things before um but when like this revolution kind of started we knew exactly um what it means for our field and we knew that like we either gonna go all at it or it's not gonna work you know it's like in the beginning like we thought maybe we can start building things over other models but very quickly we understood it like in order to be a very significant player in this field and to actually innovate around it we need to have our own I remember it's like not an easy decision to make I actually remember the one meeting where a CEO was like okay we're doing it I was like, are you sure?
3:22But you know, like two and a half years later, we're here. That's awesome. So what's the standout versus 2.3? Curious, what's new? Yeah, I mean, there's a couple of things, actually. I mean, I think like when people say, okay, it's 2.3 to 2.5. I saw some people asking, why did you skip 2.4? I'm going to spare you the I think there are discussions about how the version name is set. But essentially, a few components in the architecture were updated. Because when you're thinking about a new version of the model, we're really trying to cater for multiple needs and requests we hear from different type of customers.
4:05In this release, it's actually special because it pushed two ends of the spectrum. So 2.5 is actually even faster than 2.3. It's based on the same compressed latent space in architecture that that allows you to be so efficient. But because we really improve our distilled model, it actually allows us to run flows that only uses that and really speed up the process. So we're talking about like a 10 second video generation in less than seven seconds. I know I read that and I was like, I was shook by that. That's jarring. Yeah. Yeah, it's like when you build a product, you want it to be something you can touch and tangible.
4:44and waiting three minutes for a video is okay for some users, but it doesn't really allow you to feel like you're controlling it. So like I said, it's faster. It understands the user better because we kind of re-dead the way we do the captioning with a custom enhancer. And on the other hand, which is I think the most exciting thing about it, is that we really pushed pixel quality forward. So we introduced like in two different phases of the generation. In the core diffusion, we essentially create keyframes, like high quality keyframes along the way that allows the video to stay more coherent, like and reduces some temporal issues that we sometimes see like artifacts.
5:33It really pushes like the sharpness and the quality you get at the end. We also replace the decoder with a new one. So basically you can think about it like throughout the process, the different steps of the process, we really push quality up. And I mean, later we can talk, Alon is the one working a lot with big VFX studios and stuff like that. So you can imagine it when you're working for a really big screen. Yeah. You know, every pixel matters. Yeah. When you guys are training a model like this and it's coming together, is it ever hard to hit stop and ship? Is there ever a point where you know that just a few more days?
6:17I feel like you could just a few more days it to death if you're not careful and never have a product. The story of our life. Yeah, it's until the last minute sometimes. You can always push it forward and we have increasing quality and speed. On a daily basis, there are things that are changing. So we set like this release milestone that we all like work towards it, the marketing team, the product team, the partners, tech, of course. And in this release, like 2.5, a bit of the behind the scenes, it was like, I don't know, 15 minutes before the button was pushed. He is not joking. He is not joking.
7:01I believe it. No, no, no. I totally believe it. We pushed it like 15 minutes before the deadline in one hour because the VP of research came with a new thing we wanted to push into it. Oh, man. That is so awesome. That is. So you mentioned big video production companies are using this. How did that come to be and what is it like working with them and serving their needs and all that? So for us currently, my part in the company is like working with partners on the R &D side. So finding like the best solutions for them, making sure they integrate it into their pipeline correctly and just getting the value they should from the model.
7:44And first of all, it's fascinating because we are working. The fact that it's a world model, we are working with such diverse like industries. So from real time avatars to physical AI to VFX to color grading to social media companies. So it's pretty diverse. And for the VFX specifically, so it's like we really had to dive into their pipelines and understand how they work, how they think. We're not trying to replace anything like we are trying to replace some things, but in the pipeline itself and to integrate it into it. So replace like components maybe instead of the entire process. like we're not replacing the artists and it's extremely important it's not just a cliche like we we are actually like the fact that it's an open source model for example it allows them to customize the model to their specific needs so it either on training on their proprietary data and doing the training on their end on-prem without like sharing any ip um or fine tuning for a very specific use case that we're not just shipping as a product company that hey hey, now LTX support X.
8:52It doesn't work like this. So every VFX company and studio just fine-tuned for their specific needs. We learn a lot from it so we can share knowledge in the company and build very specific things for them, such as the quality upgrades that Daniel talked about.
9:15Are you seeing a lot of adoption among professional VFX studios and things like that. I just kind of wondered how the vibe is around that. I think it really changed in the past, I don't know, few months, I would say, both for VFX and to animation as well, actually. Awesome. Yeah, and I think that we are working very closely with Asteria, a production company, film and animation studio that we partnered with. They are like AI hybrid with traditional filmmaking. Great, that's cool. When we visited their studio, we actually saw such great use of the model for real cinema film productions that will end up on the big screen.
10:01So that's amazing. For us, it's incredible because sometimes we see AIs like this fun thing, funny thing that you can do for friends and stuff. Like a toy. Yeah, exactly. But no, it actually goes live to cinema biggest screens out there. So that's amazing. That's really cool. I feel like in Hollywood, or at least the perception from outside Hollywood, is that it's either a threat to jobs or it's slop. And it's really, we've yet to, or at least publicly, we've yet to really have the narrative take hold that, no, this is actually really useful and studios can use it without it being either of those things, which is really cool.
10:41And I'd love to hear a little bit more about how you're approaching that. Yeah, Grant, I think you're absolutely right. And I think it's also what Elan was talking about, because, you know, we've been like going to the West Coast quite a lot in the past two years. And the sentiment has definitely took a big shift, like in the beginning of this year, it was absolute fear in the beginning. And, you know, I think like, from mostly from my opinion, from like from the unknown, you know, it's something new, you're not really sure how it's going to affect you. So you're just keeping some distance from it.
11:10But I think as they saw that, like, you know, beginnings of what they can do with it at the end of the day these people are extremely talented and creative and technical and you know they I think they realize that maybe there's like opportunity for them in that and what we're seeing is that when they actually step into it as really creative people they have the best ideas and they make magic with it you know it's our job to kind of like make sure that they can actually use it for those purposes you know we talked a little bit about like what's new we're we're at least now like the native HDR support.
11:45For VFX Studio, it's a go or no go. They cannot compromise on these things. But it's been a real joy to see that and how they're using it. The best part is that when they surprise us with the thing they're making, like, oh, wow, that's really cool. You made it with LTX. That's so nice. That's awesome. So let's talk about how people are using it. So what tools do they use? How do they work it into their workflow? So like help us wrap our heads around how you actually can use this thing. So yeah, just throw a few examples. So one could be like for a 2D animation, we have like partners working with a model with like their artists are working on like the keyframes and we like LTX is doing the in betweening for them.
12:29So we created like a very specific kind of training on their proprietary data set to have like a model that actually create like very structured non-smeared in-betweens because everything in 2d like the lines are super important so we actually created a solution for them and they are now producing the in-betweens um with the model other partners like are on their way to integrate it as well um so that's on the 2d animation we also have like on the same visit the same day we visited asteria um or a day before we visited like markup robotics um so it's uh a startup from San Francisco that fine-tuned the model to move a robot and do like all sorts of actions like making a sandwich and stuff.
13:18Wow. And that was like maybe the most surprising moment for me at least. Working at Lightrix for like eight years coming from creative background and stuff. So robotics, it's not something that I am familiar with. And suddenly seeing how they just use the same model that we just saw like on the VFX studio that they use it for cinema and they use the same foundational layer the same infrastructure but fine-tune it for their needs that was like pretty insane and other use cases could be like real-time chat like two-way conversational chat with an avatar like in a video and audio that kind of interact with the user.
13:59These are just a few. Awesome. So if, for example, for the 2D keyframing one, how are they actually, what tool are they using? Are they piping the model into an animation software? Are they using a separate tool like Comfy UI? Do they have a proprietary app that they built so that they can use it? How does it, from a mechanical perspective, how does it actually get used? The partners that we are working with currently, working with Adobe Animate, most of them, and then they're building additional plugin or tool that handles all the LTX generation. Then they just work again, hybrid AI, you got the generation, you have the in-betweens, you pick the ones you like, you place it on the timeline, the traditional one.
14:52Awesome. That's cool. And other partners like in the community, as you can imagine, most people are using ConfUI that we also partner with. And we also have an API that support that. Again, for some companies, I think sometimes it's really critical for them to run it on-prem, but others might actually benefit from doing it through our API and kind of let us handle that part. So it's also available there yeah well before we uh before we dive into the demo uh find if we bounce through a couple questions from our viewers real quick uh first question comes from damian rufus and he's asking what actually makes a great first prompt here what context structure constraints examples are framing are we maybe missing that can really improve your first response I mean, I can take it and Alon can add up for sure.
15:53There we go. Yeah, I think that it's all about managing to describe what you want. And I also saw different people doing it in a completely different way. So some people start with something very simple and actually let the output of the model guide them in the right direction. So you can start very simple. We have a custom prompt answer that does the matching of like whatever prompt you put and kind of make sure that the model understand it properly and that it fits the way the model was captioned um and then basically based on the output you're getting and that's a part of it like being really fast you can actually do that they start like adding more and more details like i know yesterday i was working on a prompt of like a car spinning and i started then it was too slow so i started adding like no make it go really fast to really get to what I want.
16:46A completely different approach that you can take is really go and describe the scene in a lot of details. The model can definitely take that. And what that can give you is really how to really control and guide, direct what you want. So starting from what does scene look like, what are the actions that are happening, the order that are happening, what sounds are playing in the background, if there's a dialogue. So I think if you can do it in a very structured way, starting with the style, the camera motion, the setup, and the actions, you're very likely to get what you're asking for. I always tell people to think about their senses, too.
17:23Think about what you're like. I feel like the more information you can give it, the more you can get things dialed in quick and easy. I must say that prompting is hard, and we hear it all the time. Like this black box that you just prompt something, try your best and get something that usually you didn't expect. Like we are constantly working on finding alternative ways to guide the model. So either with a video condition or audio condition or multiple keyframes or anything that can help the model understand your intent in a way that is not textual. We see, not just us, like the people we are working with and the artists at the end they just prefer to talk visuals and not like text at this yeah so yeah what are what are some ways that they could do that like if you're an artist or you think more visually what what are some ways that you could apply your your vision to the the model so first first of all we're gonna see a demo of it but audio is a great condition sometimes you just have like a very strong audio that you want to rely on and this guides the the model like amazingly well it just like drives everything the the how expressive the person is the the environment the people around it um so audio is one thing for video uh kind of references we see a lot of um different uh variations of it so one could be like i don't know for 3d artists or people from the 3d space so having like a blocking animation or something that is very basic from the cg software to just guide like the camera trajectory or the placement of the objects in the scene and having this as the um baseline or like the the driver for the generation or taking a pose of a person like talking with his hands or any gestures and just copying it to the generation instead of prompting, he raises his hand at second two, and then it's impossible sometimes.
19:34Yeah, it's such a more efficient way to communicate if you just have a picture of a guy doing this. It is. We've got one more question here before we dive into the demo, and the rest will hold off until after. But is there an application to creating training videos for businesses? Or have you seen anybody doing that yet? Yeah, so like people using, say, like LTX models to create a training video for a business. Like training in the sense of like an explanatory video? Yeah, onboarding, you know, like tutorials, that sort of thing. Yeah, no, absolutely. So we actually worked in the last two months or so on like an avatars feature that can take like long form.
20:22I don't remember exactly how long we got it to work, but I think it was 30 minutes. Yeah, 30 minutes. So, yeah, and then, like, you can, if you can feed it, like, with your information and data from your favorite LN, probably, then you can absolutely do that. And something else that was really cool that the team was working on is a cut on action. So you can actually do this from multiple angles and keep the timing and everything intact. Yeah, so it really gives you a sense of, like, seeing someone talking to the camera. That's so awesome. I'm now terrified as a podcaster. I'm just kidding. Oh, yeah.
20:59There's so many hours of my face on YouTube. So it could totally fine tune it. It's ready for training. Yeah, exactly. Exactly. Well, go ahead, Grant. Yeah. No, no. You go. You go. How about if we take a look at it? Can you guys show us a tool? Show us around? Maybe see some videos? Sure. So first of all, it's important to note that it's not like anyone can consume LTX in different ways, like via ConfUI, using code through the API. We also have LTX Explorer, which is kind of the platform that we built for everyone to test and evaluate all the different features. So this is the platform we're gonna see today.
21:42It's great and I recommend everyone to just go and test things out. So the first feature that we wanted to show is just, oh my God, can you hear the sound? Yeah, you might want to turn the volume down a little. Wow! Wait, it just played. No, I'm not sure. Go ahead and try it now. Sorry, it's going to be a bit loud. That's okay. It's just a random generation.
22:15So I just prompted like, what was it? A cinematic shot of a rhino charging in New York City Street. So a very basic prompt. I didn't try to create something too directed or like just something to get out of the box. But here I think we have a different example already in the platform of a Dalmatian dog running fast along a sandy beach with some like description of the environment, the lighting, the camera. Where is the van, the red camper van behind it? So I just hit generate to start with something. And here on the platform, we can just choose the model that you like, the aspect ratio, the FPS.
23:03Sometimes we do see that higher FPS gives better results in terms of like, I don't know, for fast motion or very dynamic scenes where you get more kind of better motion blare looking rather than artifacts of like the model sometimes. Yeah, that makes sense. Can we zoom in on the prompt a little bit and maybe you can break down the structure of it? There we go. So it seems like you start with the subject and what they're doing, the action. That is a bit about the aesthetic, the vibe, the tone, also adding some details on the environment like the water, the forest, the vintage grim and red camper van.
23:48then about the camera, how it moves alongside the dog, about its depth of field, the lighting, and then also about the sound, because we are like multi-model models, so you generate audio and video together, and then you can also direct the sound from the prompt. And here we just get like a few versions. Oh yeah, there's a camper van, there's the dog. That's so good, you guys.
24:22Wow. Wow. It's crazy. He's just running and running. In this one, the very easy. In that one, they parked it on the water. Bold. But I think the fun way is if you give it a very specific prompt, like this one, that you have everything in it. So you can then just start iterate on it. So here, for example, I can add like until it runs into the water. Let's generate this one and see. Let's give it eight seconds instead. So the fact that I... What's the URL for LTS Explorer, by the way? I want to share it with people. So LTSIO is like the main website. And from here, you can just hit try now or straight to app.ltx.
25:20Perfect. I'm going to link that in the chat if anyone wants to follow along as we're doing this. Great. Yeah, and maybe something to know about the platform while it's generating is that this is connected directly to our API. So you can just top up as much as you want to try. There's no payment or something. You pay for your generations, but nothing else. You don't have to commit to anything. So free to sign in. Oh, wow. Yeah, that's good context. Do you just like... Oh, Corey, cut out. Sorry. Can you hear me now? Yeah, try it again. Do you just like buy some credit up front and generate away?
25:59Right, yeah. So you're getting like something to begin with, but then you can top it up directly from our console. Awesome. Alon, I believe you were about to say something before I cut you off there. Sorry. Let's see if it runs into the water. Perfect. We're rooting for it. I noticed the van got moved up the hill. Yeah, did it! Something happened to him. He immediately hit it.
26:33Yeah, there you go. That one looks great. You start to direct it, and again, it's not the same generation because we didn't use the older versions as the reference for the next one. We just used the same song, so it doesn't commit to anything from what we saw before, but still you've got the same vibe, and you can just keep on iterating until you get something with your life.
26:57So what if I wanted the Dalmatian to be the exact same spot pattern every time? How can I enforce that? Do you have any good advice there? For these kind of things, I think this kind of moves to the more advanced flows of like using a video reference and start like iterating on it. So we have like, I don't know, maybe I can find it here. Spring shots even maybe. Yeah. Or even start with an image. So if you go like that, these are all text to video. So if you start with an image, you know exactly what's going to be your starting point. Yeah. Yeah. That's probably like the way I would do that. Yeah.
27:33Yeah. That's right. Right. I think for some tasks, if you just want to use the same dog in the same environment, just iterate on it in a more precise way, let's say, or very specific way, then you do have all sorts of flows like relighting or adding things into the scene and stuff like that, or day to night changing the lighting and stuff. So you just use the same video and iterate on it. So this way, everything remains the same beside this change. that's really yeah that's awesome cool we do have a few more cool things that we wanted to share yeah that'd be great let's do it yeah yeah nice so yeah we're just we're just here for the ride the next the next is audio audio to video well which is extremely cool and can be used in very smart ways so for this we've prepared some footage in advance wait let me grab it okay so here's the audio file here's the image we're seeing the one second yeah I mean we're still seeing a switch to a tab yeah yeah oh so there's an option that says share this tab instead.
28:56Share the tab. OK. There we go. Yeah. I knew something's going to. Yes. OK. I love this guy. Wait, who is that? Who is that? OK, and we have this voice sample. Please take a moment, if you haven't yet, to subscribe to the channel so you can check out all of our interviews with the people building and impacting AI every day. All right. Let's just give it a shot. I've begun, Corey. This is amazing. Usually we don't add any prompt to audio to video, to be honest, because the audio condition is so strong. Let's just try and see what happens.
29:40Yeah. Oh, I'm so excited. It's getting personal. It is personal, but it's okay. It's okay. You mentioned the amount of content that you have on YouTube that people can use and abuse. And just how easy it is. We almost did it. We almost trained something on your voices. The instant you could duplicate a voice, my horrible, horrible troll friends went stripping videos off of YouTube. Started just sending you audio files of yourself. Audio files of me. And I was way less cute than I am as a puppet. Oh, actually, speaking of puppets, while this is loading, Amanda Barton in the chat asked, can you use a real-life background with a human-made puppet creature and create variations of that creature with LTX?
30:36If we need to break that down, we can. Yeah, wait, I mean, like... I'm trying to process it. So, like, okay, one way to interpret that is, like, can you film your own footage of yourself, let's say, and then animate that and use that as the starting point and add maybe additional characters into content that you've created yourself. Yeah, so I mean, I don't know if anything I get to show you today, but you can do in-painting with the model, so you can definitely mark areas in the videos and add elements to it. So that's probably the way I would do that. We also have, you can extend the video, so you can start with a certain footage of yourself and extend it.
31:20So it's going to maintain, it's going to be consistent in terms of your voice and the visual and the way you move, but allows you to kind of keep on. So that's probably how I would do it. I don't know, if there's maybe a follow-up chat, we can go into more specific things. For sure. I'm curious to see the video, though. Yeah, let's see this first. I'm going to see what he does and says. please take a moment if you haven't yet to subscribe to the channel so you can check out all of our interviews with the people building and impacting ai every please take a moment if you haven't yet to subscribe it's perfect it's so good can we can we just get this version of you
Read the full transcript
32:05Yeah, I read a few more examples earlier today. Let's see. Please take a moment if you haven't yet to subscribe to the channel so you can check out all of our interviews with the people building and impacting AI. How did you get that footage of Corey? That's so unsettling. When he was younger. uh yeah here you can find yourself it's like uh much like in mathematics recently where where you all discovered i say you all not you specifically discovered an error in a yeah we tried that is so amazing
32:56wow you need to share that you have to share those with us Yeah, you got to send that our way. We've got a full name article. Yeah. Sure. Cool. We prepared another kind of fun use case and actually a really powerful one. I'm a fan of reality TV. That's, if I'm being honest here. Yeah. That's fair. I love it. Y 'all are. We just don't talk about it. Sign me up. Yeah. What's your favorite show since you brought it up? You don't want to get there. Yeah. Fair enough. That's the line, right? The trashiest. Anyway, in reality TV, when you like take footage of like the people talking or any interview or podcast, so you just get like this rough cuts when you need to trim it into like shorter version.
33:45Right. You need to to trim something mid sentence or trim two parts that are not connected. And then you need to either cover it with footage or have like a morph cut or a rough cut or whatever. Yeah. So we've worked on a feature using LTX. The capability behind it is called retake, when you just retake a part of the generation. But let's show you the... One second, sorry. Mm-hmm. Whoops. Okay. So we took two different parts from your show. Welcome, humans, to the Neuron AI Explained. I'm Cory Knowles, and I'm here as always, joined today by the one and only Grant Harvey. How are you, Grant? Today the curveball is that I didn't throw you a curveball.
34:35You've done that curveball before. I remember that episode. Yeah, that's funny. You're just too rude and you don't let Grant talk, but here's the cut. Today by the one and only Grant Harvey. How are you, Grant? Today the curveball... Okay. So now wait, I need to change my... Okay. So the Morph Cut feature is built just exactly for that. I'll remove the example from here. One second. So then I can just upload like the two parts that we just saw, part one and part two that we want to stitch together. We're going to have like a very simple preview here. Welcome humans to the neuron. Grant Harvey, how are you, Grant?
35:29Today the curveball is that I didn't throw you a curveball. Ignore the black part. It was just like the loading. And then we just like regenerating this area here and kind of make it seamless. So he generate. I'm thinking of when someone's audio drops out of recording a podcast, Grant. Yeah, sometimes happens. The ability to not lose a great question or answer, maybe. Absolutely. Absolutely. We have a full, a DAC full of examples, like for audio specifically, because the model can work for audio as well. It doesn't have to be video plus audio, so it can just, all the capabilities, all the editing features and stuff can work on audio alone, and that's great.
36:17I think every editor we've ever talked to or showed this to, they were like, okay, I need this. Because you can see the pain they were having about all the videos they had to edit with these annoying cuts and it made a lot of problems. Yeah. Oh, yeah, it's very much a thing that happens. I actually ran the exact same example before, I think, earlier today. Is it the same? You can just watch this. Welcome. Yeah. So here is the same result. humans to the neuron ai explained i'm cory knolls and i'm here as always joined today by the one and only grant harvey how are you grant today the curveballs that i didn't throw you a curveball wow you've done that curveball before so yeah that's great i mean like that's obviously seamless there's no cut there um yeah i'll be dang yeah so this is really cool The audio stuff excites me too.
37:14Like the ability to maybe, could you use it to clean up audio? Like maybe enhance it? Be like, this guy sounds a little thin. Could it boost your audio or things like that too? Trying to think if I should just pull some examples. Yeah, but like another. Sorry. Go ahead. No, you go first, Stan. No, I was just going to say it. like the um it's it's kind of like uh like clay you can play with i think that like this is why um for us a lot of the time you know people think about these models they think about text to video or image to video or reference to video but when you start actually play with what you can do with it there's so there's so much more like so many different ways uh to do that a lot of those things that along showing you here are coming from conversations with people are using it or trying to do something but think like okay but how can i do that with that or they just think you have these and these problems but it's probably not relevant and you're like well let's give it a try or let's try to fine-tune it for that specific need um yeah and you yeah and it's it's very adaptable in that sense and probably surprising how often trying something you don't think will work does yeah um for sure for sure um you know and i think i think what sometimes people are surprised is like how little bit of data they actually need in order to get it to work we're talking like 15 clips and the model already learns like what you're trying to do.
38:43Yeah. Wait, are you saying that you could use 15 clips minimum to fine tune or is that just in general like for reference? So we saw like, you know, we tried the whole range of like taking, you know, hours and hours of data and thousands of clips. It has definitely advantages, but we also saw like modalities trained on as little as 10 clips so wow i would say it depends on what you're trying to achieve and like how complex the thing you're trying to build but i would i always recommend starting with um little like don't overdo it first try let's start with a small amount of like videos see what you're getting and then build it up actually this is kind of related to a question uh someone someone asked in the chat um if i want to use an actor do i need to upload an image of him or her every prompt and i guess the way i'm interpreting that is you know how how does that actually work with the character reference and maybe what if you're trying to do something where you're doing a whole movie with the same character does it make sense to fine tune that that character and and you know so the video model recognizes them what's your advice there in terms of yeah i mean if there's if there's for sure like a you know a specific character that you're working with the best way to go at it is to train it's a trainer laura um using ltx for that character it's it's it will get you the best results um the other way to do it is again like make sure that you're using um some people do it like train a laura for the an image model and use that in order to keep like character consistency with between different scenes and and then using just some kind of an image to video modality.
40:26Yeah, so you can do both. We actually now, I mean, soon it's going to be available on the platform of Lundi's showing. We're actually integrating a trainer on that. So there's people that are less technical can also play with it. So you'll be able to train your own modalities directly from the platform. To like train yourself. Yes, for example. If I wanted to create a query. Yeah. So we could make Muppet Corey look exactly the same. It could be yours. I think Jim Henson did a fine job. That's awesome. Do you have any other demos for us, or should we take some more questions? I can share some audio stuff, but whatever.
41:10Yeah, we did have a question about audio. I think Dorna Brigade asked, I thought Alain said something about how he doesn't usually add audio in a specific situation, but we wanted to follow up on when you do and don't use audio. Do you remember what you said? And does this bring you a bell? As a condition for the generation? Yeah, I think maybe that's when you were talking about it earlier. Yeah. To me, I think that if you have an audio condition that you need to rely on, just use it. Like, I think this is one of the most powerful features with Altiac. And it really, you saw like the way it moved the hands and everything.
41:57It just makes sense. Like, it really understands the video correlated to the audio and how it should work together. So I think like adding an audio as a condition, even if it's not like full strength and you use it like as a way, weaker guide, I think it really helps. So let me think that through. So an example would be like, I have a certain way that I want my character to perform something and I want to actually speak it out loud and then make sure that it performs it with the same tonality that I'm looking for. I would record myself doing it and then use that as the like reference. Is that a good example?
42:33That could be one example for sure. And also, I don't know, if you take animation, for example, then you rely on an audio track. Like, you don't generate the audio. You have the audio in place, and now you need to create the animation on top. So sometimes you just have an audio that you need to use. But what you said, Grant, is absolutely, like, definitely a use case for that. I have a two-part question. I guess they're two separate questions. I just don't want to forget them. The first one is, will it follow a script pretty cleanly? And the other element I'm wondering around that is, can you...
43:17Let's start with that one. Yeah, so just like, will it follow a script for dialogue? The short answer is yes. I mean, like you usually put it in quotes, like the text that you're supposed to say again, because this is what's in the caption. But even if you don't, like we're gonna do that. If it's properly phrased in the prompt, we're gonna do it anyway. You can imagine that when you have multiple characters, like you have a dialogue, you have like multiple people speaking, it gets trickier. We actually develop a specific modality for that that allows you to dictate what character is saying, what lines over the video.
43:58But I will also say that if the voices are very different, or it's like two different characters, it usually works well. So it kind of makes it easier for the model to understand who's talking. But for example, with Vata Loma sharing with the audio, if you have two voices, you will get two characters. The model is very strict about it. This cannot be the same person talking. It has to be someone else. So yeah, that's probably the the challenging part in that awesome is there a the other part of what i what i was thinking a minute ago was is there a recommended length of a single clip you don't want to pass i know it 10 seconds is pretty long compared to a lot of what we've seen over the you know past few years uh do you go much farther than that or is that kind of where you hit diminished returns or Yeah.
44:47Sorry, I don't go ahead. Go ahead then. No, I was just going to say that the recommended length is around 20 seconds. That's basically where it goes. But like we mentioned earlier, for specific features, we did work in order to make sure it works for longer because for Avatars, 20 seconds is a little low. One thing that I didn't mention when you asked about what's new in this version, we actually also release a component that allows you to not not tell the model how long the clip should be and let the model decide based on the input so the model is now trained because what because again like the model is very uh obedient if you're gonna put a script it's gonna try to push it all into the time you told him so um you have to like well now you can basically just like add your script and based on the length of it and what's happening in the scene the model knows what lens each would generate nice and i think i said 10 minutes but i meant 10 seconds so uh sorry for that at one point you said 10 seconds or i heard 10 seconds translated it i was like i said 10 minutes i didn't mean 10 minutes well you did say you were able to one one time generate a 30 minute clip so it's not outside the realm is probably the answer um yeah i i don't know uh if i want to share this with you guys because I'm nervous, but I created a fork of the LTX desktop that you released when you released LTX 2.3.
46:17And the way that I'm handling the script side is like considering each line in the script as a different shot. And so I've made it so you can actually work with the script and then the tool will actually separate like, okay, each line is a different shot and it sort of like isolates it for you. I'll have to send it your way and get your seal of approval at some point. It sounds really cool. It's actually, it's nice hearing that you're doing that because it's, you know, some of the reasons we're shipping these things is exactly for that. So people can participate, you know, it's like, it's one of the fun things about being open as a community can participate in what we're doing.
46:54And not just like the, you know, the AI community, but researchers all around the world are showing us things that they can do with the money. Like, ah, okay, we didn't know it can do that. Yeah, like training a robot. How do you train a robot? with this model? You're just generating a bunch of footage, and then you use that as the data? How does that actually work? We can't tell you all their secrets, but the idea is actually very cool and very simple. When you think about it like a model like LTX, so a model model, when you show him an image, like an input image, and you ask him to generate a video, it essentially predicts what it's going to look like.
47:34So imagine you have a robot and it needs to pick up something and put it in a box. LTX knows what this looks like. So if you have like, imagine you have like what the robot sees and you fit it into LTX. LTX can tell you what it should look like. And then like if you train your robot to know like how to take the direction from LTX and do it, you essentially have a robot that can do anything you want. That's wild. So, I mean, I oversimplified it a little bit. But those who are doing it, they're fine-tuned LTX on data that looks more like the thing. They're fitting it, and obviously they're working on their own custom modalities to actually train the robot to understand how to move their hands based on what they're seeing in the video.
48:21Yeah, so cool. Yeah, it's really, really cool. Okay, I have a couple of questions from the chat. We can lightning round through some of these. So somebody, there was two questions. Steve Hartman said, how does a person splice several clips together to create a coherent scene? Would they use that morph tool or is this just like putting it on a timeline? What's your advice there?
48:50I think the best way is to use images from the same scene. I think that's the safest way, I would say, just to place like a storyboard or whatever, like first frames for each shot that works together. I think like image model. The last frame of the other or something? Yeah. So yeah, you would use an image model or your own reference that you hand draw or you shoot yourself and then use that as a frame and then... Yeah, I think that's the easiest way and like you have a lot of control working like this. I think it makes a lot of sense. You can definitely use like, I don't know, the extend feature sometimes to just continue a shot in a different way or taking it from there and then trim the first part or something.
49:40Yeah, Daniel, do you have? Yeah, I mean, I guess it depends a little bit on what you're trying to do. You can also prompt the model to do this like multi-scene, like a multi-shot clip. It's a little bit trickier to control but um like we actually i think that like i'm not sure if it was posted but we actually have like a specific guidance on how to prompt for that that we're gonna um that we're gonna post soon um yeah it's a little bit trickier to control it i'm not gonna lie um but it's but it's fun that's yeah i think like alan's suggestion was the most controllable the most steerable and and that multi-shot one is the coolest but still you know you gotta be careful with it yeah we had another oh go ahead no no i was gonna say it depends on like how much surprise you want in your life basically yeah that's fair that's fair and then uh it looks like so we got another question from ricardo just in case it got lost in the shuffle i asked a question about gore vfx and censorship i think maybe this is a good question like if i was creating a horror movie let's say um could i use ltx to do this or you know what's the what's the rating limit on what that can do.
50:54Yeah, I mean, so you can do that. Like, it's limited, essentially, by the way we're prompting the model. We definitely don't want to get anyone to do anything too crazy. Yeah, which is totally reasonable. Yeah, exactly. But we're trying to be as flexible and open as possible. Awesome. Yeah, that's fair. Maybe, Ricardo, you filmed the gory parts yourself. but you use LTX for the rest. Okay, so then another question here. If we generate five shots, for example, of about 30 seconds, can we append them in your app? This is sort of like what we were just talking about, but what tool do people use? I mean, you can just do it in Premiere, right?
51:36Or how are people using it? Yeah, I think I would go to an added developer. Unless you want something specific to happen in this transition between them, And then you can definitely use LTX to make it something different. And again, there I would also say, like Alain was saying, you don't say you need to prompt something very specific. Sometimes letting the model kind of like take it where the model thinks it needs to go is a good place to start. That's cool. That's good advice. Next up, oh, I clicked the wrong one. Do these videos come with open captions? I have no idea the answer to that question around AI video in general.
52:19Yeah, that's a good question. Can you ask the model to caption the video? You can. I'm not sure. I'm not sure it's going to work. It wasn't, I would say, it wasn't trained for this specific. That's for sure. It will probably add something. Yes. What is that? Random characters here or there. That's a cool feature for someone to build out, even as an add-on or something at some point. Next up, can people try it before buying credits? Is there any short free trial functionality? I think on Explorer, there's a small amount of credit just to try. You can try a video or two from one of the features.
53:06I will say that, as Daniel said, Explorer is not a product that you now need to commit to a tier, an yearly plan, a monthly plan, whatever, that just renewed after a month or a year and you're not noticing it. You can just pay on the go how much you want just to try another generation or two. It's very flexible and easy to use and test things out. Yeah, and it depends on what hardware you have at home, obviously. like you can also run it like locally. Yeah, let's actually address that. So how big of a graphics card do you need to be able to do this? How many gigabytes? It's a good question. It depends for what exactly.
53:49I will say I'm probably not the hardware expert. Our recommendation, like our official one, is to have like at least 32 gigabytes. But we saw people doing it with much less. So if you have, I don't know, 4090, 5090, I have an RTX 4000 and I think it's 24 so you should be able to run it on that, obviously the more you have the less offloading you need to do and it's more flexible to work with it but we also saw people running it on their MacBook Pro with M5 chips, it's pretty fast also so if you have I mean, a recent MacBook Pro, you can run it there as well. Yeah. This is cool. This is cool. I've been using LTX with my students at Tufts, where I teach an AI class for film and media studies.
54:47Is there someone on your team who is working with educational institutions? Good growth area. Yeah, that's a great question, actually. At some point, we were really trying to push for that. and we did some activities around this area. I would just, like my advice, we'd be like to send us an email. Let's talk. We love seeing the way the community is using it and we'd love to support it in any way possible. Excellent, excellent. I got one more. I don't know if we missed anything else, Corey, but I want to kind of combine two thoughts that people shared. What do you got? Yeah, go ahead, go ahead.
55:24So somebody, Damien, right at the beginning, was asking if you have any advice for prompting 2.5 specifically beyond what's in the official docs, like any hot tips. I think you've shared a good amount with us. But then there was someone else in the chat who also asked if you have a few sample prompts or how do you recommend people start? We walked through what makes a really good prompt outline, but if you could maybe address any additional tips there, that would be helpful if you have anything that comes to mind. If you go to the platform, so every feature you try is actually already pre-populated with a sample prompt and a sample asset.
56:04So you don't need to kind of like, okay, where do I get like an audio sample now to check like audio to video. So you can just use whatever is there. And it's also a good example of what works well. So if you go to the platform, you will see all of that right there. It's exactly for that because we saw people just like, you know, struggling with where to begin. Yeah, and I will add maybe that there are so many people in the community working with the model, so you can find like tons of resources, prompts, and stuff on Reddit, on Discord, on X. You should definitely look for this there as well.
56:40Do you have any channels you recommend, or like any, like this Reddit, subreddit is a good one for these tips? So on Reddit, it's the stable diffusion one, I guess. Most of the discussion is there. On Discord, the Banadoko channel account. It's really good for those who are trying to run it on their own. Very cool. Yeah, we'll share the link to that in the newsletter that we do with this. The gentleman with the educational question just asked about posting email addresses. If you're not comfortable throwing them out here, we could always... Is there like a team at LTX? I mean, like, you can email me.
57:29It's like a debugger reads at LTX.io. Okay. Got it. Cool, cool, cool. Well, guys, this has been great. I really appreciate you taking the time to join us. The model is sick. Thank you. Really, really impressed. Thank you so much. It's a lot of hard work for a lot of people, really. Chisies for, like, our amazing research team for building this. That's awesome. Yeah, and I just want to thank you again for making this open weights. And it seems like your company is really committed to that. And I think it pays off because people want to support you when you do that. Absolutely. And thanks for muppeting me.
58:06Oh, wait, we got one last word. Will it work on a DGX? Do you know what a DGX is? Yeah, on DGSpark? Yeah. Yeah. Yeah, yeah. Cool. You certainly have enough gigabytes there, my friend. Yes. alright well everyone thank you so much for coming and watching today we really appreciate you being here please take just a moment to like and subscribe and go check out the newsletter and join 730-ish thousand other people who read it every morning we'd love for you to be one of them I'd also like to give a quick shout out to A-Refs thanks again for sponsoring today's episode go check them out and on that note that's all we have for today so farewell for now humans thank you so much thanks again guys Let's go.
58:50Thank you.
From the publisher
Want to make better AI videos but have absolutely no idea what you’re doing?
Perfect. Neither did most of us at first. 😹
Join The Neuron LIVE this Thursday at 10 AM PT for a total beginner’s guide to AI video generation and video prompting, featuring the team behind the newly announced LTX-2.5.
We’re going hands-on with the new model while learning the fundamentals of actually prompting video models effectively.
What we’ll cover:
🎬 How video prompting is different from image prompting
✍️ How to write your first good AI video prompt
🎥 How to control shots, motion, subjects, and scenes
🧠 What video models actually understand when you describe a scene
🛠️ A live demo of the brand-new LTX-2.5
❓ Live Q&A with members of the LTX team
No previous AI video experience required.
If you can describe the video you want to make, we’ll show you how to turn that idea into a much better prompt.
And because this is live, bring your weird prompts, difficult questions, and video ideas. We’ll put the model through its paces together.
📅 Thursday, August 13
⏰ 10 AM PT
Learn more about LTX-2.5
🔗 Explore LTX-2.5: https://ltx.io/model/ltx-2-5
📰 Read the announcement: http://ltx.io/newsroom/introducing-ltx-2-5
🤗 LTX-2.5 on Hugging Face: https://huggingface.co/Lightricks/LTX-2.5
