Midjourney, ByteDance, and Hailuo’s New Video Models [This Week in AI]

21 Jun 2025 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Denoised - Midjourney, ByteDance, and Hailuo’s New Video Models [This Week in AI]

Overview In this episode of *Denoised*, hosted by Addy Ghani and Joey Daoud, the discussion revolves around recent advancements in AI video generation technologies. Key topics include Midjourney's new video model, ongoing IP lawsuits, and innovative tools like Topaz's Astra and FLUX's Kontext model. The hosts also explore the implications of AI character ownership and its potential to become a legal frontier.

Key Topics Discussed

  1. Midjourney's New Video Model
  2. Midjourney, known for its image generation capabilities, has launched a video model.
  3. Key Features:
  4. Outputs at 480p resolution.
  5. 5-second duration clips with an extend feature for up to 20 seconds.
  6. Limited motion settings: high and low motion.
  1. Ongoing IP Lawsuits
  2. Midjourney faces lawsuits from Disney and Universal regarding IP infringement, specifically for generating outputs of characters like Yoda.
  3. Discussion Points:
  4. The focus on output quality and the challenges of regulating IP in AI-generated content.
  5. The potential for stifling innovation in the media and entertainment industry.
  6. The competitive landscape with Chinese models that do not prioritize IP regulations.
  1. New Tools and Technologies
  2. Topaz's Astra Cloud-Based Upscaler:
  3. A creative video upscaling tool allowing users to blend traditional upscaling with more artistic interpretations.
  4. FLUX's Kontext Model:
  5. Enhances contextual awareness in video generation, allowing for more consistent character appearances across shots.
  1. ByteDance's Seedance Model
  2. Introduced as a competitor in the AI video generation space, with notable features:
  3. High output quality (1080p).
  4. Enhanced instruction-following capabilities.
  5. Strong performance in dynamic scenes, such as water motion and multi-angle views.
  1. AI Character Ownership
  2. The rise of AI-generated personas raises questions about intellectual property rights.
  3. Discussion on potential business models for talent agencies representing AI characters.
  4. Legal implications of using AI-generated likenesses in advertisements.

Key Takeaways

  • AI Video Generation: The pace of innovation in AI video generation is rapid, with new models emerging frequently, each with unique features.
  • IP Regulations: As AI tools evolve, so do the challenges in regulating intellectual property rights, especially as tools become more accessible and powerful.
  • Future of Content Creation: The evolving landscape of AI-generated content suggests a shift in how creators and consumers interact with media, potentially leading to new business models and legal frameworks.

Conclusion The episode underscores the ongoing developments in AI technology within the media and entertainment industry, emphasizing both the potential and challenges that come with these advancements. The hosts encourage listeners to stay informed as the industry continues to evolve.

*For more insights and discussions, tune in to the next episode of Denoised.*

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00sometimes like you know i will go to a bunch of different tools to get the output that i'm looking for. But then the problem is like, they all have like slightly different qualities. And then it just feels a little less consistent in the output. You know, if I'm combining shots from Luma and combining shots from VO, you know, maybe each shot I'm getting what I want. There's not a consistency across the whole sequence. Yeah, for sure. Exactly. And it's similar to, you know, sensors or film stocks. And that's the reason they do. Yeah, you can't go for a cut from a Blackmagic camera to a red camera, unless you have a really good colorist that can match the two.

0:28Yes, right. But then the AI generation is so baked in, it's hard to modify that.

0:36All right, welcome back to Denoised. In this episode, we're going to try something a little bit different than our usual format. We're going to try more of a This Week in AI kind of roundup format. So we're going to try that on Fridays for a little bit. So let us know how it goes. You know, a little bit AI-centric. That's for you, Matt Workman. Not changing to AI land yet. Taking a very broad approach of what the definition of virtual means. This is all virtual. Yeah. I think it's all virtual. We're in the simulation. VP land. Yeah, it's all virtual. We're not talking about LED walls here. So yeah, we got a handful of stories, mostly that came out this week.

1:06Some that came out last week, but we weren't around with summer schedules. So yeah, let's jump into it. The first one that's kind of going crazy on my feed today and yesterday. Midjourney now has a video model. Yeah, very exciting. Midjourney was one of the OGs of image generation, what, two, three years ago? Yeah, when you could only use it on Discord. Yes. For a long time. I mean, they just launched their web portal, I don't know, maybe end of last year. Yeah, the OGs also in a lawsuit. yeah so the mid-journey lawsuit that's really interesting because this is the first time look we all knew that you could generate ip like uh you know iron man or spider-man yoda on a lot of models not just mid-journey yeah and this is the first time the studios are actively going after them yeah for the outputs is universal and disney yeah uh for like specific outputs of like yoda and some of the comic book characters yeah i did see some commentary that they were going after mid-journey because they are the least funded out of all the major generators and might be the easiest target so they they they want to kind of obliterate them or or set a precedent or acquire them okay i don't know i'm guessing either well i don't know but i mean they're all doing their own internal ai stuff as well i don't know i mean i think this goes back to what we talked about before where i still think the you know the focus is on the outputs and it's like sure put blockers and these things where you can't generate yeah ip protected characters it's already been trained right yeah like it's stuff that pandora box is open and even if like maybe these modest specific models were shut down because they were trained on ip data now the models are big enough where you like and we know like looking at deep seek and other models where it's like we can train models on less material and still get better outputs or as good outputs sure you know i think what outputs you create with that and putting protectors in for that to prevent infringing on IP, I think that'll probably be the outcome.

2:53I think so. Yeah. Look, one thing the studios have to be careful of is they're relying on AI technology to take them over the hump that we're experiencing right now into another era of prosperity. Not just media entertainment. I think a lot of other industries like FinTech, MarTech are all relying and looking at AI to go through the hump of efficiency and sort of the next generation. And one thing they can't really do is stifle innovation and kind of crush the promise before it even has a chance to hatch and become something that's meaningful and provides a lot of value. And not to bring up global policy and stuff, but also, I mean, China has competing models that are just as good and they don't care about IP at all.

3:38You cannot regulate, you know, Tencent's models or Alibaba's models. Yeah, so, you know, if these get slowed down, it's not like it's going to stop. Having said that, I do see, I mean, of course, intellectual property is a studio's main source of revenue it is worth its weight in gold that's you know that's what they really are about right so if they have ip like if they bought lucasfilm for four billion dollars then yoda is a very real asset with a real price tag on it you can't just go around generating baby yoda or yoda anywhere yeah yeah i mean protect the outputs like you know the training whatever but the outputs add more controls on that absolutely yeah but now going back to the video model the journey from innovating uh so the video and video model so this isn't a huge surprise because there's been a lot of talk for a few weeks about the video model coming the main headlines around it 480p out of the gate um so pretty low compared to the other ones i feel like every time there's a sub hd output topaz is like yeah dude i mean which we will talk about in a second yeah a lot of image generation models like at its core are like 760 by 760 or 1024 1024 but then they could with enough upscalers and down samplers on either side you don't even notice it yeah you know because it's able like stuff like topaz it's able to generatively fill in the details right not just upscale image right you don't need as much compute yeah to get to that finish line uh durations i believe are five seconds but you that has a built-in extend feature and you could extend it up to four times so you can kind of get a 20 second ish clip yeah out of a total and then it's kind of got two motion settings um so compared to the other models i mean pretty limited settings for now but they've got a whole pipeline of i'm really excited about the high motion setting because all of the ai stuff that i look at all the ai films that i look at they all feel either static camera or just like a really yeah like i feel the opposite i feel i'm excited about the low motion because i feel like everything i see is like some epic sweeping shot of something moving.

5:42And it's like, what if you just want a nice, chill, like background play kind of shot? Oh, okay. So I was thinking like in terms of the one-er in the studio. Yeah. Like, how would you do that with AI? That would be a high motion thing because you're moving, traversing through space, 3D space, going up the stairs. And the entire time you have like a handheld shake. Yeah. I know I didn't publish it, but I was trying to replicate a studio-ish shot. Oh, come on. How could you not show that to me? for the inside the AI studio trailer, which also would never, uh, wasn't finished yet, but I'll do a trailer after the series is done.

6:13But I was trying to replicate that using VO3. And so it was a lot of just text to image prompting and, uh, some success of like following a character running through like sound stages and stuff. Yeah. But yeah, it wasn't exactly what I wanted. Yeah. I gotta say like, I don't think text to video is going to be the way we do things. No, no. Like it has to be something more like a spline or something more controllable where you're clearly defining the camera arc and stable virtual camera was like a good attempt at that like no define the path through this 3d environment and then yeah that control the pace and the ins and out points and all that stuff yeah i'm seeing something like the the the stable camera is i mean i haven't seen anyone else do it ish like that because it's sort of any of the tools that do offer camera controls it's either like a preset motion or runway is gen three i mean they haven't really brought a lot of the features to gen 4 yet but gen 3 had a pretty cool one where you can kind of click and drag to define what kind of camera motion you wanted yeah um oh speaking of runway cristobal if you're out there listening we'd love you on the show so hit us up all right yeah good yeah just retweeted one of our tweets so uh yes please come on the show yeah the camera control yes is lacking in a lot of the tools text the video yeah i mean i feel like that's just like the first stop when they launch the model and then they roll out the other features like even vo2 was originally text to video and then now they've got image inputs and faster models and so it's just like yeah the in between frames is a good attempt at putting in granular control into what is video generation could be anything yeah yeah and then also even at the other extreme end of video to video and you know i've just been seeing crazy outputs with luma's new modify video feature which is outside of runway pretty much the only other true video to video we can give it an input video give it some input images and um well pica does a good job with that stuff too they do video to video isn't that what john finger does he does a lot the stuff he's been posting lately is all luma okay and i think he's been like consulting or working with luma okay pica was pica what you could modify the video you give it the video and you'd be like you know that was the one painting model it was like an in painting got it yeah i mean i guess maybe you could style transfer the entire video but yeah uh luma's has been wild with what it can do and what I've been seeing online.

8:30Yeah, no, Luma's pretty powerful, for sure. Yeah, I'm excited about, look, the more tools we have out there, the better. And what I'm realizing with a lot of the video models and the image models is every one of them is a little bit better at something than the other one. It's like in the car. They're all flavors. Yeah, it's like BMWs are going to be better at cornering. Mercedes is going to be better at soft luxury. I mean, for a more appropriate metaphor, I just think of film stocks. Go for it. Different film stocks, you know, they were known about different qualities, different looks, better for low light, better for outdoors better for different skin tones yeah yeah the more you mess around the different ai tools and you find it's like okay this one's better like vo uh is very good at photorealistic looking outputs i found it's not that great when it's like a little bit more stylized you're trying to do something a bit weirder you know the runway or or um luma or you know tend to be better with that uh if you do animation and stuff yeah or just the type of material trying to do like an animation if you like if your images are for you that you're trying to use for image to video are animation images and some other models work better for that um Midjourney, I would guess, would probably be a lot better at that because so much of the material on Midjourney is very surreal, very stylistic.

9:35And the outputs that I've been seeing online are in that realm of that surreal land. And now seeing it come to life with video, it looks very good. Realistic stuff might not be its strongest point. But also, the outputs are 480, so you're still kind of battling with. Yeah. I still want to see like a Spider-Verse style, very stylized and also really aesthetically pleasing and dreamy style of storytelling. That's somewhere in between pure 3D Pixar style animation and, you know, photoreal live camera. Like it's somewhere. Somebody, please. The tools are out there for you. I had another meeting with my Dolphin Level collaborator.

10:15And we have another project we've been talking about that has a stylized look like that. So experimenting with that. we'll keep you posted awesome okay joey you're not telling me enough man all the side hustles outside of this pod i need to know a lot of projects going on okay all right all right topaz topaz yeah coming back to topaz big fan of topaz from way back when i've used topaz for years uh they had their standalone they've recently it's called topaz video but it used to be called something else more technical but they had a video upscaler yeah a photo upscaler they were like the original upscalers yeah using ai it wasn't even previously upscalers use probably machine learning machine learning denoisers and edge detection denoisers were even you could denoise without ai okay yeah you would know more than this yeah there is mathematical algorithms for that stuff like turning noise into color information okay like that i used to have a video and so there are like a bunch of models or methods that you would pick and it was sort of based on what your source material was and what you're trying to do.

11:16Like if you're trying to fix interlacing issues, if you're trying to do a turn a video into slow motion, or if you're just trying to up res archival material, which is what I mainly got it for, because we were doing a lot of documentary work. And you would get like, it was good, but you get, you know, mixed results sometimes. And also, if the source material was like pretty blurry, there was not a lot you could do with it. Now, a while ago, they rolled out Starlight, which was built into Topaz Video. and that was their sort of first AI diffusion upscaler. Right. They're on their seventh version now, something like that?

11:45Yeah, Topaz Video 7, which is a standalone desktop app that you pay for. But then the other week they just rolled out Astra, which is a separate, it's a website, it's still in beta, but it's a separate product, website, you use credits, and they're calling it a creative video upscaler. So you can give it your source footage and then there's a slider, you can either upscale it and be like, hey, honor. Just upscale. Just upscale, maintain the consistency as much as possible. or you can slide it to take a more creative interpretation. Love that. So that, yeah, that's where the generative part probably comes in.

12:15It's probably attached to some kind of image diffusion or video diffusion model. And that has, they're probably doing some compositing where the upscale is still happening and then it's weighing what the creative upscaling is and then mixing between the two. Yeah, it's a good guess. I mean, it's all software architecture at the end of the day, but you got to interview Manji at AI on the Lot. I did, yes. It's that full interview is up. So you can check that out. And yeah, we talked a bit about that. And yeah, I was curious too, of like, because I feel like so many of the, you know, sort of generative AI video workflows are, you know, make your image, put your image in the video generator.

12:49And then most of these generators are still 720, maybe 1080, or in mid-journey's case, 480. And then run that through Topaz to upscale it. Or even not necessarily if just the resolution number is fine, just to give it more detail and to sharpen it. Depending on what your shot is on the AI outputs, a lot of them just lack a lot of detail. Yeah, I think Asteria or rather Moon Valley mentioned that their model is not only ethical, but it's 1080p. It's like one of the first HD models. Yeah. And also when I talked to Jenny, which they're doing work with kind of documentary recrees for like the TV documentary genre where they're like, usually don't have a lot of budget to do like, you know, big historical productions.

13:26They had an issue where their outputs were getting flagged by QC of not being sharp enough. And they ran it through Topaz and were able to add the detail, the sharpness back into it. And then it was able to pass a QC. Yeah. I mean, in the camera world, we have all of these tests that can vet a camera, right? Like, what is the dynamic range of this camera? What is the circle of confusion performance of this camera and so on? And that's what a couple of people inside of Netflix do. And that's why they have the Netflix approved camera list, right? Imagine that for AI models, video generation models.

14:00I mean, I think that's what we're going to be doing. I mean, if you're doing a production that involves AI, like you're probably going to want to do test shoots. I mean, also another project that we talked about, like, I'm testing out different methodologies, because we got to figure out like what's the workflow going to be and you know we do it on a small scale and figure out what that workflow is going to be so that we can a dial it in and kind of have to be consistent but also like if we bring on other people so we all know like this is the workflow so we're doing the same thing because you know everyone sort of has their the more you talk to like people that work with gen i they all have their kind of tools they lean to or kind of favorite methodologies they work with but i've been finding finding as well like sometimes like you know i will go to a bunch of different tools to get the output that i'm looking for yeah but then the problem is like They all have slightly different qualities, and then it just feels a little less consistent in the output.

14:42If I'm combining shots from Luma and combining shots from Veo, maybe each shot I'm getting what I want. Yeah, there's not a consistency across the whole sequence. Yeah, for sure. And it's similar to sensors or film stocks, and that's the reason they do... Yeah, you can't cut from a Blackmagic camera to a red camera unless you have a really good colorist that can match the two. Yes, right. But then the AI generation is so baked in, it's hard to modify that. Yeah, you're not working with a lot of data there. Yeah, unless, you know, AI video generation models open up like some kind of raw format, right?

15:15Where you can dial the photorealism up and down. You can dial the dynamic range up and down and stuff like that. Hey, dude, that's a million dollar idea I just came up with. Probably get there one day. How much compute is that going to take? Right, on top of the video generation. Yeah. You want 10-bit EXR files? Yeah, but anyways, act to Astra. Yeah, no, we're big fans of Topaz, man. You guys keep doing your thing. yeah so astra yeah i think it's beta so that and that that's kind of in the credit system so i mean the one nice thing about the topaz software you buy the software you run it on a computer and it runs locally um starlight they do have a local starlight mini or starlight light that can run locally on an nvidia card there are other starlight model it has to even though you're in the app it has to go up upload your footage to the cloud and process it in the cloud and then it brings the footage back down yeah astra is completely cloud-based yeah cloud-based is the way to go So it's really like, yeah, even if you have like a 5090 GPU, you can't compete with the cloud because it's elastic.

16:12Yeah, it'll spin up as many things as you do. Yeah, it'll spin up as many H100s or some really high end GPUs up there. Yeah. Also cost you more. Yes. Yeah. And for everyone that was saying, oh, subscription model, no thanks. A lot of times there's no other way around this. Yeah. And so when OpenAI charges 200 a month, like that money is going to electricity and people and stuff that runs the cloud. Yeah. All right. Other updates. Technically, this came out like a week ago, but we didn't really talk about it or two weeks ago. Flux's new context model. Yeah. So Flux is also another big OG in the image generation game.

16:47You know, you have Mid Journey, Stable Diffusion, and Flux were the three that was very early. And I believe at the beginning, Mid Journey was based on Stable Diffusion. and I think the history of it is Flux was founded by some folks that left stability and went to build their own company. I could be wrong but to start Black Forest Labs which makes Flux. Okay so that those and those things happen like within months so you know fast forward to two three years later we now have these three OG companies that have done image generation. So Flux had Dev and Schnell which were the two big image generation models.

17:24Dev, Schnell and SDXL are probably the three most popular DIY hobbyist, even professional grade tools for comfy UI. Yeah. So if you want to download something, run it locally and modify the crap out of it, put attachments on it, IP adapters on it, you know, everything. Those are the three models to go to. Yeah. And how they look good. Yeah. Those are like the good local models. Yeah. So now that, you know, stable diffusion has moved on to SD 3.5 large, which is obviously a much more complex, heavier model, able to do more photorealistic stuff, able to do anatomy well, text well, all the things that were challenging with SDXL.

18:02Flux has also moved on to a bigger, better, more expensive model. Expensive meaning computationally expensive, which is context. One of the cool things that I thought about context is it eliminates the need for small, lightweight fine-tuning, what LoRa's are typically. So when you want character consistency, the way to do it is just six months ago the way to do it a million years ago yeah so you go on a site like replicate where you can build a laura and uh you know i want cartoon joey in every shot and i want the same cartoon joey to be this high you know this color this t-shirt whatever and you generate that laura then you attach it to flux schnell or sdxl and then in comfy ui you're generating consistent shots yeah but you to generate it you need like a good 20 to 50 images and you have to label it really well sometimes you have to remove the background uh it's the data grooming that takes longer than the actual model training yeah and if you don't groom the data right then you've screwed up yeah you can have weird outputs yeah like joey can have three eyes right yeah yeah i did train a laura myself but it was like i guess all the photos were kind kind of similar-ish.

19:17So then every output, it was always like, no matter what I would say. It's like, make Joey sad. So now Flux Context had kind of built that in. So you generate, let's say, you know, Cartoon Joey and Flux One Context, and you can call that Cartoon Joey. And then on the next prompt, say, Cartoon Joey is now going to a bar. And then it remembers it, it has contextual awareness. It'll pull that in. It's similar to what we're seeing now and getting used to it like chat gpt or one way where you just give it a single image of us or the person or whatever and be like hey put this person you know in a western city yeah and it just takes the image and has a pretty good understanding of that person and puts them somewhere else yeah to me it feels like um as the models are getting more and more complex all of the inner working bits that you and i talk about today are getting swallowed up in it and i think in a year or two nobody will even talk about fine-tuning or building custom models because the big models will have the ability to intrinsically do that so well yeah you just do it yeah okay here's this thing just you know do it or even like maybe at some point just give it a picture of someone and it knows them and then later on it's like oh hey you know put addy in the scene yeah it's like uh you know you buy a new iphone and ios already has the calendar app the mail app the browser in it you're not oh which browser should i plug into my iphone which calendar i mean i mean i use a calendar a different different one yeah okay but i'm saying like it's it's sort of built in and you don't even think about it you're like yeah my phone can do calendar mail and browser why do i need new ones yeah yeah yeah yeah except if you're a pro like joey then you want unless you want to use fantastical or spark for your email please sponsor us so yeah flux context yeah i mean it's cool because i've uh the demos are interesting because you just give it the image it has like a lot better spatial understanding i've seen demos you know it's like oh you know you give a picture of the person it's like show me the back of their head and it like makes an image and i'm like oh it looks like yeah if i move the camera behind them looks like yeah like it looks like as if i move the camera behind them i just you know or give it an image and it's like hey put them somewhere else and you know just single shot image yeah my guess is there is some sort of a 3d model that's taking the 2d image yeah building a crude 3d environment model and then that's being used to influence the generation for the next one right so there's there's like a history and a memory associated with this model and that's probably why you need a cloud-based service because you need all that stuff to get registered to a user yeah that's the other point with this with the context model because i was thinking like oh i can get back into comfy ui mess around with this thing but there is no local model there's no open weight model you can download and run this locally uh you have to run it through their api through their cloud as of now they did say that they were going to work on a version that yeah you could run locally but right now you have to use their api yeah unless you plan on having like 10 high performance hot you know electricity heavy machines at home like our features the cloud man with all this ai generation stuff all right next up to date yeah so uh video models are constantly improving themselves so obviously we saw vo2 go to vo3 and that was a major bump you know it's as big of a bump as when unreal goes from 5.4 to 5.5 like it seems like a tiny little decimal number like man does it do a lot more so there's in that same vein of thought minimax which is built by chinese company hilo is now on a second version i think it's called o2 minimax o2 is supposedly according to some influencers better than vo3 the better than the rating stuff i always find that so weird and subjective because it's also it's like better at what what at what i need to do what i'm trying to do but anyways yeah it's good it's really i think most people look at it and like most people are not you joey they're not making commercial grade stuff and you know making a living off of it most people are just prompting and looking at it online and seeing oh that looks good and that doesn't look good so i think the judgment is either from photorealism prompt adherence and perhaps detail and contextual awareness and yeah those are highlights of like what's new and improved in this model yeah 1080p output so pretty big output yep compared to everything else uh sota instruction following so uh sota is state of the art yeah which is weird yeah it's nothing fancy it's just the next so presumably better at following the written instructions and prompts and extreme physics mastery, which goes into everything else we're talking about of just understanding the physical world.

23:53Yeah. So for photorealism and the demo shots are showing, you know, clowns blowing fire and circus stuff, which I think is also intentional to demo that because there have been so many, there's a bear doing backflips and I've seen a lot of tests. This was gymnastics was like one of the weak spots with a lot of these models. I think even like VO2, there are a lot of tests of just like, you know, Hey, have someone do a back flip or you know somersault and it was just like the weirdest distorting yeah because body moves and they just weren't our range of motion is so finicky like you know we have bones that move in all different directions but don't move in one right and it's hard to train like your fingers can bend backwards this much right and it can bend side to side but yeah yeah it can't bend all the way back yeah like our knees can only bend one way how would an ai know how to I mean, I think it depends on what it's trained on.

24:46But yeah, and I think that's why they're showing, they're demoing the circus stuff. The other thing that I think a lot of video generation models fail on, and now I'm noticing a lot of demo reels have it, is knife cutting vegetables. Oh. Yeah, so the interaction between the knife and the cucumber or the tomato. Still, because I feel like that's gotten a lot better. It's gotten better. It's gotten as good as spaghetti eating. Spaghetti is, Will Smith and spaghetti has gotten very impressive. But it's still not there. I'm sorry, Mr. Will. Yeah. yeah i mean i yeah i think that's there's um yeah i think that's why this demo they've got a guy a guy on a high wire bouncing around and a look and then jumping off and landing on someone at all like the physics feel you and i should just do uh a of generative video benchmark yeah standard yeah yeah we should i think we we have a good idea of what to do and what not to yeah i debated this because of like figuring out what the baseline prompts yeah would be but again i mean i've also gone back and forth through that too because it's like i guess there's yeah i guess it's useful to have a couple of like baseline yeah like if you generate how it does if you generate a sunrise in a dark environment that would be a good indicator of dynamic range you take the brightness level of the sun and the darkest level at anywhere else in the frame that's your real world contrast ratio value yeah whatever the output image is and we could do a spaghetti thing that could be like a physics test right yeah uh yeah i've been one of my tests was pushing a button because i had this issue a year ago and i was just trying to have some one like i just wanted a finger pressing a big red button and i had huge issues where it would like rub the button or like massage i think i've seen that one yeah i think i posted a tweet about this like all the issues where it was just like would not press the button other people were like yeah i've had the same issue and i did do i tried it again with possibly vo3 one of the newer models and it did mostly do it um but that was like that's been my own benchmark as well like can i tell it like finger press this button yeah and it does it correctly yeah for me it's a knife cutting cucumber knife cutting tomato yeah i've seen there's another genre now of um with vo3 of knife cutting like glass fruit and it's like a asmr kind of genre um but like it looks pretty good oh and knife cutting like fake fruit that like is like glass fruit i like that that's pretty creative yeah all right so yeah halo to hate hail high luo high luo to new model yep and so out of out of china you have tencent with hanyan you have alibaba with wan w-a-n and high luo with minimax those are the three major video models and another new one from bite dance oh which is tiktok's parent company has gone under the radar seed dance 1.0 and it has a little speaking of ranking it has the number one spot in the artificial analysis ranking for text to video and image to video yeah i don't understand those ranking metrics yeah i'm not quite sure how they measure it yeah but impressive stuff multi-shot views of the same scene so being able to create different angles but being the same physical space it has done really well at that prompt adherence ranks really high high motion there's shots of a surfer in a wave which looks freakish freakishly accurate really yeah nice diversity of styles reasoning's outside of distribution one of the issues with water with video generation is the water starts to look like gel it looks more viscous than it is okay yeah like i noticed that with i think i'm gonna say wonder dynamics that looks like water man yeah that way that's pretty damn good yeah yeah maybe a little too foamy though i don't know i can always nitpick i know yeah i mean like his movement looks a little like offish but i mean this looks pretty good yeah so uh there's a mini version of the model or a light version of the model you can run on foul um i don't know if the full model's out yet or i think i don't think it's out out yet i think this was like a early preview that they released Do you think Disney's going to sue by dance or see dance?

28:50It depends what output you can make. Like, can you make Mickey Mouse outputs out of it? When I tried, I had weird, people have been saying like, oh, you can make, you know, like people are making the Stormtrooper videos and those other things out of VO3. That's so good. I sent you one. Oh yeah, I sent you one. Yeah. It was so funny. Yeah. I don't fully know how they're doing it because when I was messing around with VO3, I was trying, even I was trying to even mention like 100 year old movies. Like I was trying to create some scenes from like, you know, Lumiere brothers type stuff and it would block it.

29:21And I was, it didn't really say why, but I was guessing it was because I was mentioning some copyrighted like name that was in their system that would block it. It's like a blacklist. So I don't know how people are making, I mean, how many people are making the stormtroopers or the Mickey mouse videos or other things that they're saying VO3 is popping out. Maybe they're leveraging the API and the API has less set of rules. Maybe that, or they're just like have a very clever way of prompting it to make it look like it. um also you could give it any specific words does vo take image reference vo2 does maybe they're doing it that way maybe it goes back to like if you if you give it an image of a character and then it makes it into a video is the model still really spot like rely responsible for generating the the video output with the character if the input was yeah so in my opinion like this is just me thinking out loud a video generation model should be able to accept the image uh classified as what it is whether it's spider-man or stormtrooper or mickey mouse stranger things ip you know and check it against the database of known ip and then tell the user that hey you don't have the permission to generate this video yeah like that's the least that it can do to protect itself from lawsuits this is clearly possible because uh youtube does this all the time and we'll detect what you upload and compare it against anything that already has uh the copyright id on it and flag it yes flag it is either you can't monetize it or and we're gonna take the but you could leave it up or just flat out this video is blocked because yeah and you don't even need ai for that that's like old school computer vision yeah yeah yeah so yeah i think that i think protecting the inputs going back to the original lawsuit we talked about it makes the most sense I'm sure these things will clamp up over time.

31:07Yeah. Okay, other update. Kria now has their own image model called Kria 1. Yeah, remind me why Kria was on our hearts and minds. Kria was always interesting because they were the ones that would do the most sort of real-time instant generation. They were the ones that you could either connect your webcam phone or load up any sort of like camera input and film stuff. And then in real time, it would just start rapidly generating very low-res. Yeah, like 10 frames a second. Blurry images, but like very realistic looking images. I also had another tool where you could roughly kind of block out your scene and move shapes around and it would instantly update the image based on your style and your text prompts.

31:48It was like very high, rapidly AI image generation. I remember that. That was like their big claim to fame. And then they also had a platform where you can, you know, like all most of these platforms, connect each other models and bring in other tools and sort of be a one-stop shop for generation. Leonardo had that for a while too. Leonardo, yeah, Flora. I mean, a lot of these tools all have Hydra. They've all started trying to be the one shop stop platform where you just, all they have to do is connect to the APIs of everything else, and then you can just keep generating. But Create did not have their own image model.

32:19Now they do, and some of the things they're saying that it does is help make it more realistic and get rid of the, quote, AI look. I love that. Yeah, so I mean, they've got a bunch of varieties on their website. It feels very, like a lot of stylistic images. It feels kind of similar-ish to MidJourney in the different styles and stuff that you could do. But yeah, I mean, the stuff looks great. They've demoed it out. Obviously, when they demo stuff, it's cherry picking the best stuff. You know, under the hood, the getting rid of the AI look, I think under the hood, it's really just adding a fine-tuned model that's fine-tuned on photorealistic content.

32:54So like stuff that's out of a camera raw or EXR or something that's just like really high fidelity, very sharp with lots of dynamic range. And then feeding your AI output into this fine tune model and then spitting out something that's quote unquote more real. Yeah, there's a demo too from also from Justine Moore that was showing a comparison of outputs of different women from different models. and Kria sort of had the most kind of realistic, you know, not like perfect smooth skin output, you know, just kind of highlighting like it can do, keep things more, you know. Imperfect. Imperfect, you know, that, you know, mimics real life.

33:38Example there, the mid-journey one is just like clearly AI. The journey has very smooth, perfect skin. I would say the GPT-401 looks pretty four-wheel to me. Yeah. But then Kria is taking into account specular highlights and direction of lighting really, really well. Skin blemishes. Skin blemishes. So, yeah. I think that may be like also a pass with just facial data because as humans, like I mentioned this many times on the pod, we are highly susceptible to imperfect and artificial faces. Genetically, you know, since we were cavemen, we have been accustomed to reading faces. We look at faces every day.

34:16That's how we build trust. That's how we detect enemy from friend. and so like all those things are built into our brain so when we detect a face that's like even one percent off it's uncanny valley it doesn't look good yeah really something's weird about this yeah yeah i don't know what the pricing of the krea compares to other models but new model on the block yeah i'm glad it's all there the competition is good and you know as we predicted like there is a new model every week yeah and that that's probably why we're changing the format of the show a little bit today right just just so much to cover yeah just because but sometimes when we're doing sort of three stories, it's like, well, you know, there's a lot of stuff happening.

34:50More than three stories. Yeah. Especially with each week with AI updates. And then, yeah, last one, speaking of other just new apps, Arcads. Arcads AI. So this one's focused on AI actors. And yeah, similar to maybe Hagen. Yeah, or Hydra. I believe that they have their own model, which makes it interesting for yeah, the main one is sort of like, they're a model that's targeting a talking actor so i think their use case is probably ugc ads uh stuff like that and but more control over editing the performance and creating the characters i i heard something really interesting the other day i forget where but it was essentially if we create digital talent synthetic talent you know um can a talent agency represent them and then can their likeness be used for product photography or you know uh advertisement like who owns the light who owns the likeness so like let's say you and i come up with you know mr charlie or whatever like some model male model and then we have like a talent agency represent charlie and then we go out and build a commercial for pepsi or something and it's charlie drinking pepsi right and we have his imagery like from because we built him from scratch we have his likeness we even build in you know uh skin blemishes and all those things that make humans humans that could be a whole new business model for talent agencies and what's the current model with like the existing uh they're not ai but they're artifact they're um synthetic the cg uh yeah which i would imagine would be a similar yeah so model like i would think like like a company owner yeah broad who owns little mikaela they are making deals with you know sketchers or whoever like these big companies and they're producing content for them and it's licensed to sketchers to use little mikaela yeah so i would assume it would be a similar playbook like yeah instead of it being cg generated it was ai generated but if you're the person or company that created it and no one's uh yeah i mean obviously the i'm trying to think too because there was the copyright office ruling on like what ai generated work can be copyright protected and so i guess that could come into yeah but like too because if you're just like oh if you like go to chat you're like hey give me an image of like a cartoon character you know that looks like uh you know a male cartoon character and it's been something out i believe the copyright ruling was like that was not there was not enough human you can easily creation involved in the loop you can skirt around that by just doing an image reference or a sketch.

37:35And that is human interference into the creation. So if Spider-Man is a... Spider-Man doesn't exist. Sorry, folks. If Spider-Man as an IP is able to be licensed and used on products used on video games, why not a AI avatar that is built by humans? Yeah, I mean, I agree. I'm sure we're going to see this. I've not seen this already. I mean, I feel like the AI characters that are generated now are more of like, Like, hey, let me just mass produce AI characters that look like real people to just quickly generate UGC looking ads to spin up on TikTok and try to sell a product. Yeah. But we will probably see, if not already, I mean, I'm less familiar with this space, AI personalities and characters.

38:17Once we get past the IP lawsuits, this will be the next legal frontier, I think. Yeah. If someone creates an AI persona, AI character, and then someone else takes that character and starts running their own stuff. And then, you know, the originator of that character tries to sue them. And it's like, hey, like, this is our character. you can't use it, is that going to be enforceable? Yeah. And can that digital character join SAG? Right? Like, it's gonna get crazy. Also, what does this do to the influencer economy? Like, look, I think we can all agree we don't need any more influencers. There's quite a bit out there.

38:48What happens to that industry? Does it get decimated? Or do they pivot into this new economy of? I mean, I think and I've been thinking about this too, just from a sense of even the stuff I'm building with like VP land and like why I've sort of have, there are other newsletters and other publications out there. And I feel like you don't really know who's behind them. And I've sort of been more of the face of it, partly for thinking of this reason of like, well, if everything's sort of commoditized and it's easy to spin up a newsletter, it's easy to spin up a YouTube channel, like a faceless YouTube channel, it's easy to spin up AI avatars.

39:20Then the counter to that would be like the human connection, the human person behind it that you can connect to. Yes, this. And, you know, so I mean, maybe there will be overall less influencers, but I still think human influencers where you can run into them, you can see them will still survive and will still have that appeal because it is a real person that you could see outside of the TikTok world. Also, it'll force them to be better at what they do. And just I think we have a lot of influencer slop right now uh-huh we'll see a lot of that go away there's gonna be i think it's gonna be way more of that well there's gonna be ai influencer slop i was thinking about yeah a lot more ai influencer slop or just ai with vo3 there's just so much there's like the yeti cavemen selling cameras or something there's just like so much ai generated um ads and and channels now that are spinning up that yeah it would fall in the slop bucket but for sure people watch them they have millions of views like people i think going back to stuff we talked about before like The general public doesn't really care where it came from.

40:19It's like, is it cutesy and entertaining and engaging for me? Is it able to move products? Is it able to move products or just keep people watching so these people spinning these up can collect a couple pennies for the ad revenue, which if you're getting millions and millions of views, that turns into real money. Yeah, I think the litmus test would be when you see an AI-generated synthetic model selling something on the Super Bowl, like at that level, million dollars a minute, then the brand see value in it. And clearly the consumers are buying that product. We didn't even talk about the Kashi ad that ran during the NBA that was all generated with VO3 from PJ Ace.

Read the full transcript

40:55PJ Ace, I think that's his name. Another kind of big X creator. And it was all it's leveling up. Text to VO3 prompts. Really good spot. Very, it was, it's Florida man inspired. So it's a bunch of crazy, like Florida people shots, like man on the street interviews and like crazy scenarios. Okay. VO3 worked really well for that because it's all just single shots. So it's like, you know, text to video. But it was an ad that was AI generated and ran on a broadcast TV. We're getting there. Yeah. Wrap it up. Yeah. Thank you to 1NDR for another five-star review. We thank you, sir. Thank you for updating your three-star review to a five-star review.

41:31Yes. Can you not give us three-star reviews, please? If your review says great show, it should be five stars, not three. It's, I'll say it's really easy on the podcast app to. You can edit your review. Yeah. Thank you. We appreciate it. Yeah. We talked about a lot more stuff than usual. So thanks for everything that we talked about as usual will be at denoisedpodcast.com. Thanks again for watching. We'll catch you next week.

From the publisher

The AI video generation race heats up with Midjourney's new V1 video model despite their ongoing IP lawsuit with Disney and Universal. In this episode, Joey and Addy break down the latest developments in AI creative technology, including Topaz's new Astra cloud-based upscaler, FLUX's Kontext model, and impressive new video models from China including Minimax Hailuo 02 and ByteDance's Seedance. Plus, they discuss how AI character ownership might become the next legal frontier as tools like Arcads AI emerge.


#############

The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.

More from Denoised

All 101 episodes
Midjourney, ByteDance, and Hailuo’s New Video Models [This Week in AI]Denoised · 42 min
Listen in VO