206 | The Ultimate AI Showdown: Comparing Video Generation Tools with Tianyu Xu

15 Jul 2025 · 27 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Leveraging AI - Episode 206: The Ultimate AI Showdown: Comparing Video Generation Tools

Overview In this episode of *Leveraging AI*, host Isar Meitis welcomes back AI video creator and educator Tianyu Xu to discuss the rapidly evolving landscape of AI video generation tools. The episode emphasizes the importance of understanding the capabilities and limitations of various tools to make informed decisions for business video content creation.

Key Concepts Discussed

  • AI Video Generation: The tools are becoming production-grade, offering new capabilities while still facing limitations related to consistency and realism.
  • Core Methods of AI Video Creation:
  • Text-to-Video: Creating videos from written descriptions.
  • Image-to-Video: Generating videos based on static image inputs.
  • Video-to-Video: Modifying existing videos to produce new content.

Notable Tools Reviewed

  1. Veo 3:
  2. Best for realism but lacks flexibility and is cost-prohibitive.
  3. Integrates sound, dialogue, and ambient noise into video creation.
  1. Kling 2.1:
  2. Offers consistent results without breaking the bank.
  3. Good for both text-to-video and image-to-video applications.
  1. Sora:
  2. Initially promising but currently has limitations in performance.
  3. Best utilized for image generation rather than video.
  1. OpenArt:
  2. An aggregator platform that allows users to access multiple AI models.
  3. Provides flexibility and convenience but may come at a higher cost due to its aggregating nature.
  1. Minimax and Runway:
  2. Emerging tools with unique functionalities; discussed within the context of evolving capabilities.

Insights from Tianyu Xu

  • Generative AI Empowerment: AI video creation is accessible even to those without filmmaking backgrounds.
  • Character Consistency: Maintaining the same character across multiple videos remains a challenge, especially in text-to-video methods.
  • Prompt Structuring: Properly constructed prompts are key to achieving desired video content results, including character movements and camera angles.

Practical Takeaways

  • Understanding Tool Limitations: Knowing the strengths and weaknesses of each tool is essential for effective video generation.
  • Cost vs. Features: Higher quality often comes at a premium; assess your budget against your video production needs.
  • Use Cases: Consider when to use native tools (direct from AI companies) versus aggregator platforms for optimal results.

Conclusion This episode of *Leveraging AI* provides valuable insights into AI video generation tools, making it a vital resource for business professionals looking to enhance their video content. Tianyu Xu’s expertise helps demystify the complexities of video generation and assists listeners in making informed choices to leverage AI effectively in their business strategies.

---

Resources

  • [AI Business Transformation Course](http://multiplai.ai/ai-course/)
  • [YouTube Full Episodes](https://www.youtube.com/@Multiplai_AI/)
  • [Connect with Isar Meitis on LinkedIn](https://www.linkedin.com/in/isarmeitis/)
  • [Join Live Sessions and Newsletter](https://services.multiplai.ai/events)

Call to Action If you found this episode helpful, consider leaving a five-star review on your favorite podcast platform!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to part four of the Ultimate AI Showdown. In this episode, we're going to cover AI video generation tools with the help of the amazing Tian Yu-Hu. But if you've missed episodes one, two, and three of the Ultimate AI Showdown, go back to the previous episodes and listen to these as well. Because in them, we've covered comparing tools for data analysis, comparing tools for image generation, and comparing tools for vibe coding. So you can know which one to pick for those aspects as well. But now to the AI video generation tools comparison with Tian Yu. In the next few years, AI technology will change our world dramatically.

0:36Whether you are a business executive trying to catapult your business forward, or just somebody who refuses to be left behind and want to advance your career, this is the show for you. I'm your host, Isar Maitis, a serial entrepreneur and an AI enthusiast. You'll hear invaluable practical tips from innovative business leaders, AI practitioners, and some of the brightest AI minds in our world today on how you can leverage AI in ethical ways to advance your career and grow your business.

1:11Our next AI showdown topic is maybe the one that has the most progress in the past few months. So while since the end or from the Q4 of 2024, we have seen incredible significant progress across the board with AI, with introduction of reasoning models and then mixed models and better agent capabilities from multiple providers around the world, one area has really exploded in capabilities, mostly going from something that was cool and maybe a geeky thing to do at the end of 2024, and now is at production level on many different aspects in 2025. And that area is AI video. So before that, still in 2023 and 2024, a lot of people started creating AI video like our guests today.

1:57But the problem was it was almost impossible to get consistent results, to get the same characters, to get the same feel, to get the same look, to get different angles, to control the camera, to have the same product in the featured video from different angles. And when you try to create videos for business, that is a necessity, nothing short of that. And all of that has changed in 2025. It started with the introduction of Sora at the end of 2024 in the, whatever it was, 12 days of Christmas from OpenAI. But there's all the other amazing tools that made amazing progress in the past six months, including Minimax and Runway and Kling, and obviously the VO family with VO2 and now VO3 from Google.

2:38And if you haven't watched any VO3 videos, you probably weren't on YouTube in the past couple of weeks. but all these tools are amazing and they all have pros and cons and things they do great and things they don't do so well and so learning which use cases you should use different tools for is a necessity at this point just like more or less everything else with AI. Now we're going to explore this with the help of Tian Yu Xu who grew up in actually data and research fields with highly structured approach to processes and building success and switch to becoming one of the leading AI video creators and educators, at least on LinkedIn.

3:14I follow everything he does. He shares incredible insights on different video tools and how he's using them and what's the process in each one. And his combination of structured approach to success combined with his amazing creativity makes him the perfect person to learn AI video from. And hence, I'm very excited to have him as the person to share this with us. He was also a guest of the show back before. So Tianyu, Welcome back to the Leveraging AI show. Thank you so much, Isa. It's a pleasure to be back. I'm really excited about this because I double with video, but it's not the core thing that I do.

3:47I mostly work with businesses on data and infrastructure and setting up processes. And so I'm personally very excited to learn everything that you have to share. Sure. Where should we start? Maybe I can give an overview of the key methods of AI videos, and then we can compare the different top models and just use a few examples. Sounds like a good plan. Okay, sure. So when you talk about AI videos, you know AI has been there in the videos for decades, but what makes the real difference is that it's generative AI because right now you can just create a video, generate a video with just talking to an AI model.

4:25You just need to give the prompt or give an image or give an existing video to generate a new video. And this is something that does not require a lot of technical knowledge or you don't even have to be a filmmaker to make these videos. So that is something that truly empowers almost everyone who wants to express themselves creatively using AI. So I think video is actually the most exciting technology in AI this year. And the foundational method, the key method for AI video is text-to-video. It means that all the video models are based on prompts. So you need to give a text description, describe the movement of camera, describe the subject, the setting, the context, and describe the movement, the actions, and then you will have the video.

5:11So that is text to video. Then there's also another method called image to video, which means your input is one or two images. So you upload an image and then give it a small, give it a short prompt or give it a camera setting. You can make the image move. Just like animation, you just animate the image. This is image to video. Then the last method is video to video. So basically, your input is a video. For example, you can upload a black and white film made 100 years ago, and then convert it to a colored film. So that is video to video. You can also upload my talking video and change my face to a cat.

5:49That's also video to video. So these are the key methods. Of course, there are others like AI avatar, or a lot of special effects, special VFS effects that you can also do with generative AI. But today, I just want to focus on two key methods. One is text to video and the other is image to video. Fantastic. Yeah. One more thing to add to image to video just to enhance, and I'm sure there's going to be examples, but just to let people know, when it comes to image to video, we started with all the models with a starting image as the starting point of the video. And over time, we got two additional additions.

6:20One is ending frame, so you can tell it where to end. So the image you load is the final frame of the video instead of the initial frame of the video, which has its benefit. And now you can actually do that. And now in some of the tools, there's even keyframes. You can add more images along the way to have more of an artistic control over the flow of how the whole video is going to work because you give it hints where it needs to go in different timeframes. But that's all I want to add. And now really let's dive into examples and talk on how it's done and best practices and the differences between the tools.

6:47Sure. Let me share my screen. And for those of you who are listening to this as a podcast after, first of all, this is going to be available on our YouTube channel. and there's a link to that on the show notes. You can jump there right now if you're not driving or something like that. But we're also going to share everything that's on the screen so you can understand what we're looking at. Fantastic. So I'm sharing my screen now. What you are seeing is Gemini. I'm inside the Gemini app. So basically, I want to first show you the different results from Bale 3 in image to video. And then I'll move on to the same, I'll use the exact same prompt across different video models.

7:21And then we can compare the results, right? So this was inspired by the Yeti on TikTok and on YouTube. I don't know who actually, I don't know who started this trend, but now it's so popular. If you go to YouTube shots or TikTok, you'll see Yeti vlogs everywhere. So inspired by this, I added my cat character inside the same Yeti setting. So here within Gemini 2.5 Pro, and then if you click on video, you can generate video directly on Gemini. So this requires at least a Google AI Pro account, which is$20 a month, or Ultra account, that is$20 a month. So I'm using Pro in this case. You can generate about three videos per day using Vero 3, but that is a fast mode.

8:05So these videos are Vero 3 fast mode. Vero 3 also have a full mode, which is the most professional version, which I will show you after this. so the fast mode is five times cheaper than 80 % cheaper than the full mode and then here this is a prompt a key benefit of Vail 3 compared to the other video model is that it integrates everything so the other video models gives you moving the motion the image the the movements but Vail 3 gives you the sound, the dialogue, the ambient noise. So all the things you can include in the prompt. So I start with the setting shot on a professional camera in a reality food show.

8:51So this is the style of the film, which is a professional camera in a reality food show. And then I talk about the character, a white cat with blue eyes dressed in yellow apron. and the action of the character is massaging a soft round-shaped dough next to a real yeti dressing black apron shedding cheese on the cheese grater. So this is the scene. And then you can also add in camera movements. For example, the camera pushes in and focuses on cat's face. Then the cat says, welcome to the kitty yeti show. Then the camera then pans to focus on the yeti. The yeti says, let's cook pizza together. Yo-hoo.

9:29Okay, so this is something you can only do this in Vail 3. And let me show you the result. Welcome to the Kitty Yeti Show. Let's cook pizza together. Welcome to the Kitty Yeti Show. Let's cook pizza together. Welcome to the... Okay, it's quite impressive, right? Yeah, again, for those of you who are not seeing, like the only thing I got a little different from the prompt and what's the actual reality so the reality looks like a reality show you see a cat playing with a doe and you see the yeti grating cheese and the yeti is the one saying everything other than the meow in the middle I thought the cat is going to say the first thing and the yeti is going to say the second thing but other than that it's pretty incredible yeah that's the point I want to get into because we are using the fast mode the fast mode which is a cheaper version of model so the cheaper version of the model takes less compute than the full model.

10:27Yeah, that's why not everything has been fully rendered. It's still amazing, by the way. But it's still okay. It's still okay. And then I want to show you the same prompt on the full model. Unfortunately, I already exhausted all my credits on Gemini for the full model. But I have some credits on OpenArt, which is actually one of the best aggregators of all the AI models. So within OpenArt, you can also choose Vail 3. and then this is the video. Welcome to the Kitty Yeti Show! Let's cook pizza together! Yoo-hoo! Welcome to the Kitty Yeti Show! Let's cook pizza together! Yoo-hoo! Yeah, you see the difference?

11:10Yes. Even the voices are, like the cat voice is incredible, right? It's actually making it sound like a cat. Welcome to the Kitty Yeti Show! Yes. It sounds like a real cat. Yeah, yeah, yeah. Incredible. Yeah, and it also follows my prompt closely, right? And then we talk about in the prompt, we have the camera pushes in and focuses on the cat's face. Then the cat says, welcome to the Kitty Yeti Show, exactly as I have shown in the prompt. Then the camera tends to focus on the Yeti. Then the Yeti talks. So that is something quite amazing I found. Yeah. So following this, I want to show you the limitations.

11:51Although I think VEL3 is the best in text to video at the moment as of today. Maybe tomorrow it'll not be the best. But there are still a lot of limitations of text to video. So it means that it's quite difficult to keep a consistent character. So the next time when you use the same prompt, you might generate a different cat, a different yeti. So for example, I continued this story. and in the next prompt within the same Gemini I said so everything keeping the everything maintained the same except the actions so the yeti comes to the frame and then drops a pile of sliced pineapple onto the pizza and then the yeti says don't forget the pineapple and the cat says yes that's the real pizza okay and then we have the video here oh don't forget the pineapple Yes, that's the real pizza.

12:46Oh, don't forget the pineapple. Okay. So the characters are not that, are not really the same. Yeah. That's the limitation. The cat is very close, but the Yeti looks different and their voice is different too. Yes. Then the next one I continued. Every day you can create three videos. So I just generate 10. The next one is the final piece. So they are eating. By the way, again, for those of you who are not watching the screen, the beginning two thirds of the prompt is exactly the same, right? It's a shot of a professional camera in a reality show, setting, scene up, close up, a real white cat, like all of that stayed the same like the original prompt.

13:18And then just the last section changes what's actually happening in that particular scene. Yes. In this way, you can at least maintain a high level of consistency in the style and in the setting, but character consistency is still not that fantastic if you are using text to video method. So this is the last video. In the last video, I have both the cat and yeti eating happily, eating the slices of pineapple pizza. And then the yeti turns to the cat and says, this is the best pizza I've ever had. Then the cat finishes the bite and says, haha, you should try durian pizza next time. All right. So then we have the video here.

13:58This is the best pizza I've ever had. You should try durian pizza next time. This is the best pizza I've ever had. You should try durian pizza next time. I like the sound. I like the ambient noise, the sound of them eating. That's the best, better than the voice. Yeah. So again, you can see that the character has changed. The Yeti becomes more like a CGI instead of a realistic Yeti. Yeah. Yeah. And, but if you use the, Again, I generated the same videos using the same prompt with the full Vail 3 mode, OpenArt. And here we have the results. This is the second prompt. Oh, don't forget the pineapple!

14:42Yes! That's the real pizza. Oh, don't forget the pineapple! Yes! That's the real pizza. Okay, so this is more realistic, right? More realistic than the fast mode. All right. So in the interest of time, I will move on to image to video. But before that, let me show you the legendary Sora because I don't think it's anywhere near. Yeah. So this is a text video on Sora. As you can see, it almost never followed my prompt. Yes. And the video is not consistent and the stuff is turning around and the dough turning into grated cheese. Like it's not. And the cat is not dressed. It just looks like a cat.

15:28Yeah, I wouldn't use Sora for most use cases now because there's no point. Yeah. And again, then, yeah. But I would use Sora for image generation. So essentially, this is the text tool. This is GPT-4.0 image model inside Sora. The same with ChatGPT. I used it a lot. It's my default image model because it has the best prompt adherence, means that it listens to you almost everything you can generate almost anything you can in your prompt and also it works naturally with within chat gpt and within within sora and and you can even prompt you can even prompt with your natural language with complex grammar complex structure yeah so here i generated these images for image to video because i want to maintain a consistent character in all the videos so we have one image of cat and yeti in the first scene and i also generated the same images for the second for the other scenes yeah and the next step is to the next step is to upload these images into ai video models and then compare the results so if you can keep if you can maintain the consistency of your of the characters in your storyboard or across all the images then when it comes to videos it's quite easy to maintain the same level of consistency.

16:53So that's how most of the AI videos in the market on social media are made. So just to connect the dots for people while you're opening the example, the VO3 still does not have image to video. VO2 does have image to video and all the other tools have image to video, but VO3 still does not have that functionality. VO3 were just released two weeks ago. And so I'm sure they will add that functionality sometime soon. Like Jan, And you said maybe by the time you're listening to the podcast, it will be added. But right now it does not have that capability. Maybe it's available. But actually it is available on Google Flow.

17:28There's a frame. There are$250 a month creative. But even if you can use it, it doesn't support sound. So the sound. So you pick one or the other. Yeah, sound. No, the sound is the best part of Vail 3, but no. Yeah, okay. Let's come back to image to video. For image to video, across all the tools, I think if you have no time at all to learn about this, if you have no time to try all the AI video tools or all the AI video models, the only model you should use is Clean 2.1 because it just works for almost everything. Okay, so I'm going to show you this. I'm here inside clean. I just want to make sure that my screen is clean.

18:13I'm sharing the okay, so I'm sharing this screen and already generated the images the videos from the images I generated on Chanchi PT on Sora So as you can see here, there's also sound but not no dialogue

18:31And the sound is not as realistic, but again for those who are not watching the video is amazing Like it follows the same like character from the images that were created and the cat is doing the dough and the Yeti is grating cheese and it zooms in on the grated cheese. It actually looks really realistic. Yeah. And it's amazing. Listen to your prompt maybe 80 % of the time and then just move. And the next one is adding pineapple to the pizza. Okay. So we have it here.

19:04and some extra pineapple lots of pineapple

19:10in the prompt although they can't talk i wrote in the prompt the yeti talks and the cat talks they're talking to each other so this is something i think we can we just need to to manually add some conversations manually add the voice and then it will look perfect Kling for people has a lip sync function, right? Yes, yes. Can you add it to these characters as well, or it only knows how to do humans? Most tools only work. Most tools have the lip sync for humans, but very few can do animals. So the only tool outside of Baird 3 that does good lip syncing for animals is Higgs field. Interesting. Higgs field also works for cats.

19:54Yes. Okay, so now we know Vio generates amazing voice and sound. If you use the more advanced mode, it's going to cost you more, but it generates even better consistency. And we see that for image to video, then cling works great. I must admit that still the movement and realism of the video on Vio 3 looks better, but from a consistency perspective, you still have the problem because you cannot use the image to start with. Yes, that's right. Image to video is almost the only way if you want to create a story with consistent characters. And then I want to show you a few others. Another interesting tool is Hylor.

20:35This is text to video. So it's Minimax Hylor and the text to video, I use the exact same prompt on the here. And you can see that they are, the videos are actually pretty precise. Although it's not a realistic video, there are 3D characters of the cat and yeti doing exactly as I told them to do. Yeah, and the camera movement focuses on the face of the one and then the other. So from a control over the scene, it's very accurate. Yeah. And then again, it also works with text or with image to video. So in the prompt, you can set up the camera movement. And yeah. And again, we can see here from a realism perspective, even though the characters look closer to what you started, the physics of the world is not accurate, which was one of the problems in the older models, where in this case, the pineapple shows up out of thin air from the Yeti's hand and so on.

21:35So it's still not in the level of Kling and or Sora like we've seen before. That looks a lot more realistic. Yeah, that's right. And I think Halo does the 2D or 3D animation quite well. but not the realistic scenes. Got it. Yeah. Another model that I want to show you is Luma Ray 2. I think Luma Ray 2 is the best model for many subjects. And even for the cat scenes, it's comparable to clean 2.1. It's comparable to clean 2.1. And you can see the results of the same image to video. Yeah, it looks amazing. Yes, hyper realistic. take but sometimes you get some random stuff yeah basically they're random stuff so again i think from what we're seeing so far from a realistic physics perspective like the universe behaves like the universe we know vo is still number one maintaining consistent characters is a limitation which again i think is going to be solved and then cling is number one when it comes to creating videos from images that if you want to keep consistent characters and then the other ones are not bad but has morphing faces morphing pizzas weird things showing up in the middle of the frame and issues with the physics of them so again but there's other use cases other than realistic videos where these tools might actually shine yeah that's right and if you compare with the price obviously veil three is way more expensive than the others it's five times more expensive than any other top model.

23:10But I think the direction is that the models are getting more powerful, but more affordable. If you compare CLEAN 2.1 with CLEAN 2.0, they even reduce the price. So the new 2.1 Pro for CLEAN only costs a fraction of the 2.0. So that's a good direction. And I only showed you the results from 2.1. I haven't even shown you the results from the most advanced clean, the high compute version of clean model yet. But I think they are already pretty impressive. Yeah. Okay. Any other examples? I think that, yeah, that's all for tonight. Yeah, awesome. So I think my final question is, you touched a little bit on the structure of the prompt, right?

23:55You talked about how it needs to be done in order to have the scene, the character, the action, the camera movement, like all these kind of things described. but then you showed a tool that is an aggregator. And I know a lot of people don't know these tools. So let's talk about that in two seconds and then I'll let you go. So what kind of aggregator exists today? What are the benefits of using them versus using the individual tools? And because I think that's going to help people who want to get into this, that makes it a lot easier, I think. Yeah, that's right. So you can think about the tools developed by AI companies like ChatGPT and Gemini.

24:26So they're owned, they're developed by the AI companies. And then the aggregators are like perplexity. So they have all the models, all the best models optimized for some specific features. But in the AI video space, a lot of aggregators are actually quite useful because they integrate not only the best AI video models, but also the image models so that you can have everything in one place. And the biggest benefit is convenience. And you can even build your own workflow or train your own custom model within these tools. tools then but then the benefit of the the original tools owned by the ai companies like clean or gemini or google flow is that the the price of generating videos are slightly cheaper because they are from the direct they are the wholesaler they're not the retailer and also they these companies often launch new features the best features on their own apps first before even releasing the API.

25:22So if you want to learn the best features, the latest features and test on the latest on your projects, I think you should go for the tools owned by the AI companies. But if you want something to scale your production or something, or if you just want convenience, then you can opt for the aggregators. Awesome. Some name of the aggregators? So you mentioned one that you use. So the one I've shown is OpenArt. And there are others like Freepeak. Freepeak, Korea. There's three that I know. And also, I forgot to mention, now there are also a lot of online workflow tools that makes it much easier for you to build your own workflow from a prompt, from an idea to a complete video.

26:02For example, you can use Flora AI or Weave AI. Weave. Weave, yes. Yeah. And they also allow you to connect to different large language models and then image models and then video models and create your workflows. yeah they are quite pumped now yeah tianyu thank you this was a great walkthrough i think it gives a lot of people a lot of a examples of b food for thought on how they can develop their own stuff i really appreciate taking the time and being with us and sharing your knowledge with all of us you're welcome and thank you everyone for the time

From the publisher

👉 Fill out the listener survey - https://services.multiplai.ai/lai-survey
👉 Learn more about the AI Business Transformation Course starting August 11 — spots are limited - http://multiplai.ai/ai-course/

Is your business ready to trust AI with its video content?

AI video tools are evolving fast — but not all of them are ready for prime time. Some promise cinematic magic. Others... produce glitchy fever dreams. So how do you separate the cutting-edge from the cutting-room floor?

In this episode of Leveraging AI, we welcome back AI video creator and educator Tianyu Xu for an expert deep-dive into the Ultimate AI Video Showdown. If you’re a business leader aiming to elevate your video game with AI — this is your go-to resource.

From cat-and-yeti reality shows (yes, really) to real-world production use cases, we break down what tools like Veo 3, Sora, Kling, Minimax, and Runway are doing well, where they fall short, and how to pick the right one for your business goals.

📌 In this session, you'll discover:

  • Why AI video is now production-grade — and where it still falls apart
  • The 3 core methods of AI video creation: Text-to-Video, Image-to-Video, Video-to-Video
  • Which tools win at consistency, sound, and prompt accuracy
  • Why Veo 3 leads in realism (but not in price or flexibility)
  • How tools like Kling 2.1 can deliver serious results without breaking the bank
  • A breakdown of aggregator platforms (like OpenArt) vs. native tools — and when to use each
  • How to structure your AI video prompts like a pro for business-grade results
  • Real-world advice on creating consistent characters and camera movements using generative tools

Tianyu Xu is a leading voice in the world of AI video creation. With a background in data, research, and structured thinking, he now brings a methodical — yet wildly creative — approach to generative video. His hands-on insights help business leaders cut through the noise and produce high-quality AI content fast.

About Leveraging AI

If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!

More from Leveraging AI

All 330 episodes
206 | The Ultimate AI Showdown: Comparing Video Generation Tools with Tianyu XuLeveraging AI · 27 min
Listen in VO