In short
Podcast Summary: Denoised Episode - Open Source Image Models Flood In, Nuke Goes All-In on AI, Google's Lyria 3 Music Surprise
Overview
In this episode of Denoised, hosted by Addy Ghani and Joey Daoud, the hosts discuss the latest developments in AI image models, VFX compositing, and music generation technology. Key topics include new open-source models, Foundry's acquisition of Griptape, and Google’s innovative music generation tool, Lyria 3.
Key Topics
- Open Source AI Image Models
- Importance of Open Source Models:
- Hosts emphasize that open source models are vital for community growth and innovation, allowing users to test, build, and modify tools.
- Current limitations include the ability to run models locally due to hardware constraints.
- Highlighted Models:
- FireRed:
- An open source model specifically designed for high-quality image editing.
- Compared to existing models like NanoBanano and Quen, it offers specialized editing capabilities.
- Quality assessment suggests it is between ZImage Turbo and Quen 2511.
- Offers the ability to make specific modifications to existing images.
- Recraft V4:
- Noted for its visually appealing output, particularly suited for social media and enterprise-grade applications.
- Unique feature: Provides SVG output, a rarity among current models.
- ByteDance's New Open-Source Model:
- Positioned as an alternative to Quen and ZImage, with claims of better quality.
- While the quality varies, the hosts critique its text rendering as artificial.
- Foundry Acquires Griptape
- Foundry has acquired Griptape, a node-based AI platform, indicating a strategic move towards integrating AI capabilities into their flagship compositing software, Nuke.
- The acquisition aims to enhance user experience and expand AI functionalities within VFX workflows.
- Google's Lyria 3 Music Generation Tool
- Lyria 3 is introduced as a quick and effective music generation tool that can produce tracks based on various non-audio inputs (e.g., slide decks).
- Hosts experiment with generating music based on their podcast context, showcasing its versatility and intuitive integration within the Google ecosystem.
- Importance of legal clarity regarding ownership of generated music tracks is highlighted, suggesting potential for widespread use in business presentations and video projects.
Discussion Highlights
- User Experience and Practicality:
- Hosts compare the practicality and performance of various models, discussing how certain models are better suited for specific tasks (e.g., image editing vs. text generation).
- Future of AI in Media:
- The conversation touches on the evolving relationship between filmmakers and AI, predicting a growth in AI-driven production tools and the diversification of production methodologies (hybrid, synthetic, and animation approaches).
- Cultural Integration:
- The episode reflects on the cultural implications of AI in creative fields, with humor and anecdotal references to their personal experiences, enhancing relatability for listeners.
Conclusion
The hosts conclude the episode by inviting listeners to explore the discussed models and tools while encouraging feedback on their experiences. They also highlight the importance of community engagement through reviews and comments to help grow the podcast's reach.
---
Key Takeaways
- Open source models are crucial for industry innovation.
- Foundry’s acquisition of Griptape signals a shift towards AI in traditional VFX workflows.
- Google’s Lyria 3 presents exciting possibilities for AI-generated music tailored to diverse content inputs.
- Continuous exploration and experimentation with these tools will shape the future of storytelling and media production.
Call to Action
- Listeners are encouraged to try out the new models discussed and share their feedback.
- Support the podcast by leaving a review on platforms like Apple Podcasts and Spotify.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Importance of Open Source Models
0:46 to 2:22
Discussion on why open source models matter for innovation and access.
“So it's really important for everybody to play around with open source models.”
Introducing FireRed: A New Open Source Image Model
2:23 to 4:25
Exploration of the FireRed model for image editing and its capabilities.
“Yeah, so FireRed is an open source model specifically built for image editing.”
Evaluating FireRed's Performance
4:26 to 6:38
Review of testing results and comparisons with other models like NanoBanana and ZImage Turbo.
“And that is something that is open source.”
Beyond Image Generation: Logic in Model Outputs
6:39 to 9:15
Discussion on how models are incorporating logic and world knowledge into image editing.
“because image models are getting better and better and they're becoming more and more available as an open source model.”
The Specialization of Image Models
9:16 to 10:08
Conversation about how different models are suited for specific tasks and styles.
“Yeah, and also, I don't know if the researchers build it this way or not.”
ReCraft V4: The Instagrammable Model
10:09 to 12:42
Introduction and review of the ReCraft V4 model and its aesthetics.
“Next up, what you had Recraft flagged before.”
ByteDance's New Open Source Model
12:43 to 14:00
Discussion on ByteDance's latest open source image model and its claims.
“has been Photoshop and Nano Banana integration.”
Exploring ByteDance's New Image Models
14:00 to 17:09
Discussion on ByteDance's latest open-source image models and their performance.
“Yes, so ByteDance drops the Quen models, right?”
Nuke's AI Strategy with Griptape Acquisition
17:10 to 19:35
Analysis of Foundry's acquisition of Griptape and its implications for Nuke.
“So Foundry, the maker of Nuke, one of the industry standard compositing software, just acquired Griptape, an AI node-based platform to integrate with their software and sort of this AI-centric strategy.”
AI in Film Production: The Future of VFX
19:36 to 21:47
Exploration of how AI is influencing film production and VFX workflows.
“I think we both had that prediction, or the Comfy thing, and then I was like, Foundry.”
Show all 16 chapters
Google Gemini and Lyria 3: Music Generation
21:48 to 24:35
Overview of Google's Lyria 3 music generation capabilities and its unique features.
“Google Gemini has got into the music generation game with Lyria 3.”
Testing Gemini's Music Creation
24:36 to 28:00
Live examples of music generated by Google Gemini and insights on its capabilities.
“Let's play this next one because this one was like, what did you say?”
Starbucks Name Pronunciation and AI
28:00 to 28:39
Explore how AI accurately captures cultural nuances in name pronunciations.
“This is, because you've said this before and I listened to it, it reminded me of this issue that would happen at Starbucks if I went to, like, Hialeah Starbucks or Doral.”
AI in Music and Creative Tools
28:40 to 29:41
Learn about the integration of AI in music creation and the visual understanding of AI processes.
“You had to give it the most boring thing, didn't you?”
Gemini Music and Legal Implications
29:42 to 31:07
Discuss the potential of Gemini music and the legal questions surrounding AI-generated content.
“I don't want anyone to make the world a better place than we do.”
User Experience with AI Tools
31:08 to 31:50
Consider how AI tools may seamlessly integrate into applications like Google Slides.
“in the future, because it's all built into Google Slides, Google Docs, and so on.”
Transcript
Automatic transcript. May contain errors.0:28What was your input for that? new model, open source models, your favorite thing that I think you've been playing around with, right? Yeah, I do love me some open source models. Look, not because the quality is the best. I think in a lot of cases, in most cases, I think Nana Banana Pro today will probably be the quality. Why I think open source models are so important for the community and the industry as a whole is because these are the models that move and inch each one of us forward in terms of testing, learning, building other things and sort of modifying, working in comfy. So it's really important for everybody to play around with open source models.
1:12And I think there's limitations on how much you can run locally. And that's probably one of the biggest reasons why open source models are just not able to compete with API models, which tend to be much, much larger. I have been getting a new appreciation for open source stuff, though, messing around with OpenClaw because token costs are a real issue. And if there was an equivalent that could run on local hardware, then that's a huge plus. Like Kimi, but then the challenge is to run the big enough models that can have the same performance as OpenAI or Clawed. You need a really beefy machine. Yeah.
1:52And I think in the future, we're going to see two things. We're going to see the model sort of come down in size or at least be the same size and just be much, much better, right? Because things just get more efficient over time. The other thing is hopefully once all the data centers are being built out, like we're going to get access to Blackwell GPUs and Rubin GPUs and stuff in the future. So then our own desktops and Mac minis and things like that could be much, much more powerful. All right. So first one that you saw, this one's from a new company we haven't heard of, FireRed, from the FireRed team.
2:23This is new out of left field. So what have you seen with this one? Yeah, so FireRed is an open source model specifically built for image editing. So this is, I think, the first wave of image models was text to image, right? Make the generation happen fully synthetic. And now that we have been playing around with image editing for a few years now, now the challenge is, okay, I've already generated the thing or I have a thing already. I took a photograph or I have this graphic. How do I make small changes? And then how do I have high quality with that and consistency with that? So this is, I think, specifically built for image editing, the same way NanoBanano is really good at editing images.
3:05The change here is, of course, it's open source, so you can download the weights and then throw it into Comfy. Or you can build your own sort of open claw-like agent that just runs this in the background as an image generation. Yeah. And sort of the other equivalent to this open source at the time was Quen, right? So sort of the next other option that would be like Quen level. I believe this is also a Chinese companies. So not much info about the folks that actually made the FireRed model, other than the fact that it's made by a team called FireRed. We don't know nationality, team size, business justification and all of that stuff.
3:46But so far, some of the images that I've been seeing, I would put it as good as ZImage Turbo, perhaps somewhere between ZImage Turbo and Quen 2511. In their acknowledgments page, they said they would like to thank the developers of other amazing open source projects, including Quen Image. Oh, interesting. So you think maybe related teams, split off teams. I'm just saying a lot of these models come from people from teams from other companies that split off and then built their own model. Yeah, the other thing that Alibaba is really good at is developing frameworks for their models. So, for example, the Alibaba models like Juan, all of the Juan video models use an architecture called VACE, V-A-C-E.
4:28And that is something that is open source. So you can build other video models that leverage that same type of framework. So perhaps the Quen models have a framework that is shareable and editable, and then the FireRed team is just using it. It's sort of like building a kit car, you know? So most of the chassis and the engine are already there. You're just building on top of that. So you got a chance to mess with this, right? Yeah. What have you found? And I've got some sample images up from their demo page. So FireRed, did I send you a FireRed image? Because I thought I sent you the other one.
5:04I don't remember. Actually, yes, I did mess with FireRed. Yeah. So this is an image that I made with FireRed. This is something I use over and over as my litmus test. 1970s busy street in New York. And you're just doing text to image. Text to image, like novel generation. Yeah. So some of the things that I look for is do the cars look messed up? Obviously, they don't look like any of the models. Like that's maybe a Pinto in the front. Doesn't look like an American car. Yeah. Looks like a London taxi. Yeah, right. So like these are obviously errors. and then having a like a busy street with a bunch of people is really difficult so then you go into and see if the people look real and of course that New York City skyline in the back what what does that look like and of course the overall vibe and mood is is it gritty is it dirty is it urban so I use the same script on I use the same prompt on Nano Banana 1 as well as Nano Banana Pro and now I'm using it on all of my image testing.
6:00So here it feels like it's Z image-esque quality. Like it's like maybe 75 % there, but certainly not Nanabinana Pro quality. No, I mean, I see like none of the text is rendered normally. All the cars here on the side are kind of blending together. So yeah, it's got the vibe of it, but nothing, you look at the traffic, like a lot of the details, nothing about this is New York specific. Yeah, and I think we did this test with C-Dent. the latest C-Dense, was it 4? That came out recently? 4.5. So this is about the same quality as C-Dense 4.5, I would say. It's still nothing to shake a stick at because image models are getting better and better and they're becoming more and more available as an open source model.
6:47So it's all good. It's interesting on their demo, on their image of demographics, because it seems like it also has some world knowledge and logic because some of the examples were, please correct the errors in this image and it's a blue colored pencil with a red line and then it changes it to a blue line oh interesting uh and then same with this other one that's a tricycle with triangle wheels and the prompt was just please correct the errors in this image and then it changes the triangle wheels to round wheels so that's interesting that it has more that it seems like it's beyond just a diffusion model to like make modifications that you specify absolutely like world understanding yeah some Some of the newer models are using two things.
7:27They're using a VLM to a visual language model to read the image to figure out what the contents are and how it's placed and all that. And then once that is understanding and seeing it, then that logic goes into an LLM to then correct it, to then mitigate it. And then that instruction out LMLLM goes back into the image model to. Can you spell that? L-L-L-L-L-M. um yeah my tongue is not working today so it's just too cold in la yeah we're warm we're cold-blooded animals we like freeze in place with dude under 60 degrees in la it's too cold man yeah we're in the 40s right now and it's raining again there i used uh the weather as an excuse what were you saying yeah so there's there's like a loop where and this is how i'm guessing it's architected obviously we don't know we don't have the actual architecture we just have the weights but um a vlm will read and see the image that logic will go into an llm to take action the action will then go back to the diffusion model to generate or regenerate oh cool i mean this is anything else you want to add about this or no uh we got a couple more models yeah this is the we got a couple more models but this is a good way honestly i mean it's always good to have options too because also like as we're talking about with everything with these benchmarks and stuff and there was like you know the benchmark to the best but so many other times it's like what are you trying to do and what model is the best for that so like i think this whole like just general like it's the best model at like everything is never the case it's always just like what are you trying to do and what's the best model for that so having more options especially in the open source area like besides quen uh is always good and the other one.
9:17Yeah, and also, I don't know if the researchers build it this way or not. They probably don't. They want every model to be general purpose and just good at everything. But in real world use, you find that this model is better at cities and landscapes. This model is better at people. This model is better at text. And it just happens to be that there are specialized use cases that the community sort of just kind of finds out about. Yeah, or style transfer. They did call that out specifically that it scores you behind style transfer. So yeah, it depends what you're trying to do. Yeah, specifically on the style transfer, I'll say that the architecture has to be quite different from a general diffusion model.
9:56You have to know the structure of the image different from the texture and the fine detail of the image, and those two have to kind of ride their separate lane in order for you to create styles convincingly. So I think this is architected from the ground up to have better style transfers. All right, cool. Next up, what you had Recraft flagged before. Yeah, so Comfy, I'm on Comfy Cloud, and they send weekly newsletter. Shout out to Comfy for doing that. And they just sent out this newsletter saying, hey, Recraft V4 is now on Comfy Cloud. So I'm like, I've never heard of it. So I go check it out, and I generate a bunch of stuff.
10:33and it is the most Instagrammable model to date. It's like everything just comes out cool, man. Yeah, I mean, look at this image and like the text is sharp. Like it looks like an Instagram pose and she's holding a sign that says, recraft before, it's now in comfy UI. And it's all super sharp and super bright and colorful. And yeah, I call it the hype beast model. Like everybody's just super cool. Yeah, everything here looks, I mean, my only weird issue is like, it looks like New York, but the license plates look like Europe. But besides that, compared to what we just saw with Red Fire or FireRed, way more coherent and consistent.
11:12Yeah, and I think this model is particularly going to be good at humans, like single-subject humans. So the stuff that you generally see on social media marketing and things like that, like somebody really cool wearing some sneakers that aren't meant to be$300. So I think for those kinds of things, this is probably a purpose-built model. And they specifically state that this is an enterprise or professional-grade model, which means probably that it's very customizable. I was going to say, does that mean it's more expensive for a generation? Yes, that too. But I will call out, because I just noticed this, SVG output, so you could do vector output, which they said they're the late ones, and I believe that because I've not seen anyone.
11:56It's really hard to find a model that will just do a transparent PNG output. let alone I have not seen anything that can do a vector output. So that's good, too. As far as I know, the way that this is done is you take the output from pixel space, and then you're doing a conversion from pixel space to vector space. However, if these guys have figured out natively from latent space into vector space, then that would be something novel and new. See, look, I'm learning stuff from you, Addy. I'm taking a note because I'm working on a project, and I've needed transparent assets. and I've been trying to find models that can make that.
12:32The only one I could find was ChatGPT image 1.5, but now it's okay. All right, now I'm gonna try this one. See, look, I learned something too. I took a note down for this. I'm gonna try it after we're done. Dude, my secret weapon with stuff like that has been Photoshop and Nano Banana integration. It's been solid. To make something transparent? Just, yeah, whatever edits or modifications you wanna do that's like a Photoshop-esque task, you could just bring it into Photoshop and just have Nana Banana mask it out and give you a transparency. Yeah, that's good for your one-offs. I need something that has API to do it at scale.
13:06All right, Mr. Scale. I only do a thousand at a time. Thank you. None of this manual hand process work. Yeah. API is... I'm running 15... My agents are running 15 companies on OpenCloud. Yeah. My OpenCloud doesn't know how to use Photoshop. It only knows APIs. Hey, quick segue. way did you see some of the open claw updates where it's like i built a company and uh elon is the ceo and warren buffett's the cfo and here are their person they're like building a whole org no i didn't see that one i didn't see i didn't see like what okay a quick update of the guy i can't remember his name right now but who developed open claw open ai hired him i don't know if you saw that one yes yeah i hired him i mean open source they got him which is also ironic since he initially called it Claudebot and then Anthropic was like, change your name.
13:57And then he changes it to OpenClaw and OpenAI is like, you're hired. I like the name. Come on in. Okay, ReCraft. Cool. One other model. What else we got? Oh, another one from ByteDance. Yes, so ByteDance drops the Quen models, right? No. Alibaba. Sorry, ByteDance. ByteDance's lab seed drops seed dance and seed dream models however this is not either of them yes also this is bit dance dance and dream are not open source but i think this new one bit dance is yes you're right the the seed models are not open so yeah seed models are not bit by dance drops new open source image model bit dance i love it just says better than quinn and z image like Like, that's their marketing.
14:47I mean, this is from an AI search Twitter account. So, you know, I would take the Twitter hype with a grain of a lot of salt. I mean, I'm guilty of this, too. Sometimes I say this model's better than that. But in reality, better at what? Exactly. And going back to what we said, it's like, better at what you're trying to do. But, yeah, some of these demos, you know, quality-wise, the humans can look a bit AIE. but text rendering and a lot of this stuff looks pretty good. I disagree. I think the text stuff is just a little too texty. It looks like it conforms to the font too well. That farmer's market board should be a chalk font.
15:29My level of good was, can I read it and understand what it's saying? No, no, dude. We're way beyond that now. Compared to Red Fire or whatever. At least this text comes out legible. It's too legible. Yeah, I know what you're saying. It looks like this board looks like I took an image of a whiteboard and went into Word and then typed in a marker font, team meeting at 3 p.m. Don't be late. It doesn't look like someone actually wrote that with a whiteboard. And this, I can already say, this farmer's market chalkboard sign doesn't look like someone actually drew it with chalk. It looks like a chalk font.
16:08And these are the dead giveaways to what is AI generated or not, right? I'd like to be normal. I can pick these up. They're like, yeah, something's off there. Yeah. Like you could never arrange flower to be this perfect, right? Like just conform to the lines of that bit dance flower font thing. This is maybe my new favorite image. Oh, is that a Shiba? I think a Corgi with a ferocious Corgi with a Mortal Kombat warrior. Yeah. Look, it's an open source model. I think it's their first. Who does Zimage? They don't do Zimage. Zimage them? No. Yeah, so I think this is outside of the seed lab. Seed lab is like their R &D lab.
16:50So Bitdance model would be maybe another team, another lab, another place somewhere. But Bitdance is a huge company. Good on the model updates. Yeah, good on the image models. So folks, keep trying out new image models. Send us some feedback and we'll try to generate some more here. So this was kind of on our prediction card-ish, but not the company we thought. So Foundry, the maker of Nuke, one of the industry standard compositing software, just acquired Griptape, an AI node-based platform to integrate with their software and sort of this AI-centric strategy. I had not heard of Griptape before, but looking into it, it seems like it's a very node-based, initially targeting enterprise clients.
17:36but from my quick glance of what it did, very similar node-based ideas to stuff that we've talked about. Yeah, this is giving me straight up Invoke vibes or Weedy vibes where it's like a very professional-grade node-based system for AI work. However, it's not comfy UI. It's not to the level of that granularity. It's still very much user-friendly. I don't know. I mean, looking at their GripTapes website, I mean, even on their header, start creating visually. script with Python when you want. So, I mean, the fact that they're offering coding and scripting in the node-based app leads me to believe it does offer a lot of granularity, but I've never used it, so I can't say for sure.
18:19Yeah, I mean, if it's going to integrate with a product like Nuke, it would have to be of the highest caliber of tinkering and control. Yeah. Right? So in that case, I would think absolutely there's scripting involved and custom nodes involved. It's interesting. I mean, it makes total sense, right? Nuke is probably the most well-known node-based tool in the VFX arsenal. And people have been integrating some AI stuff with Nuke. It's happening already. I've seen demos of it. But for them to go all in on an AI-native node-based tool and then maybe retrofit that into existing Nuke, I think that would be the plan.
18:58Yeah, I guess you have Nuke nodes and some of those nodes are AI that can generate background elements or just other elements that you need. There's already integration with Bebel when you're doing relighting and getting their various passes and then bringing that into Nuke. So something in that realm that kind of just keeps everything in one spot that you can just call up different models and nodes and generate things. Yeah, it makes sense. It's a good move. Yeah. I think our prediction or my prediction was Comfy UI gets acquired and I was like, maybe by Foundry. Dude, that was my prediction.
19:33Was that your prediction? No, Comfy getting acquired was my prediction. You said Foundry. I disagreed. I think we both had that prediction, or the Comfy thing, and then I was like, Foundry. And then you're like, no. Okay. I think it's going to be someone like AWS or some non-VFX player. Or they just, I mean, they raise some money. Maybe they just stay independent, open source. Hey, Comfy Anonymous, if you want to come talk to us, drop some hints, come on the show. Comfy Satoshi. come be satoshi did they're making a satoshi movie with ai did you see that i saw i saw bits about that yeah i've seen more and more you don't seem too excited out about oh yeah i've seen more on the overall thing i keep seeing is more like movie announcements that involve some part of ai in the workflow as part of the announcement yeah which i mean we've known it's going to go this way people kind of like get some weird freak out but it's like uh yeah there was a link i had it saved i think it's in the newsletter but it was um yeah basically one film's going into production that's gonna you know rely heavily on just sort of like gray or blue screen box and generate the environments around the people which you know it's like you still have actors and it's just generated environments which we've talked about this for a long time there's gonna be three lanes of movie production with ai at least in the near future that i see first lane is exactly what you just said.
21:01It's like performance-driven hybrid. So people are still people using real cameras, but then that goes into background replacement, relighting, and all that stuff. Second lane is fully synthetic, and we're seeing a lot of testing with that. It's like the stuff that Dave Clark puts out, right? Like the people are synthetic, the background, everything is art-directed with AI. Third lane is animation, but I think keyframes are important. So the keyframes and some of the pose control is still manual and hand done, but then obviously no rendering, right? It's all AI-generated stuff. It's the AI. Yep.
21:33So that's like the Cucco thing that Paul Trillo did. Yeah. That's a good example. Yeah, that's a good example. Yeah. Going back to the Foundry and Nuke, I think it's a good play. I'm curious to see how they integrate AI into Foundry, into Nuke. Yeah, big fans of the Foundry. Yeah. Shout out to our homies there. All right, third story. Google Gemini has got into the music generation game with Lyria 3. You've been having fun with it. It's fun. It's quick. Much quicker than like Suno or maybe even Eleven Labs. It generates in a few seconds. And the quality, to me, a non-music person is pretty damn good.
22:15So I generated a 3 just to kind of show you guys what we got here. Do you have a Q-Dub, Joey? Yeah, let me play it. But covering what it does, one thing that stood off, again, just quickly kind of covering what I thought was interesting and unique about it, because it's not just text to music or music restyling. You could also give it other inputs like slide decks and videos and images and just use that as sort of the inspiration and kind of see what happens. right right it's very much a compliment to the google suite of products like vo and nano banana and google drive and everything else like you need a music companion for your deck to kind of help present the thing that you want to do what i'm saying is interesting is the inputs what you give it is different than i've seen like you could just it's like a non-audio input yeah you give it a slide deck and be like make make a song about this uh and see how it interprets that i'm giving it to noise.
23:16Yeah, and that goes directly with the Google ecosystem play, right? Like a lot of times, like if you ever make slides on... Yeah, like if we're doing a notebook LM and it's like, hey, give me some background... Well, it does something... It doesn't just do background music. It's like, all right, let's go with... Let me show your examples and I'll show what I'm talking about. But first off, to use this, it's basically right now you just go into Gemini and you could hit on the bottom that you want to... in the tools and that you want to make music. And so then that'll kind of shift it into music mode.
23:43all right so did you just throw our cover art in there yeah i'm gonna see what it does this golden ring feels heavy stone i rule this kingdom yet i rule alone a whispered thread on distant wind the crushing weight of other men's sins my blood for the realm and the northern sky a thousand hearts that cannot die i bear the crown of shadow and light to hold back the endless night okay you remember ed sheeran in game of thrones yes so this was my inspiration i was like okay make like a like a very uh analog sounding i don't know if you were watching the new game of thrones and you're feeling inspired oh no i i heard terrible things so i'm staying out of it wait for me i just caught up with it it's great you're wrong you heard wrong okay okay okay maybe the first two episodes maybe it takes a little bit to get into but the last like two episodes have been wild oh shit okay game of thrones season one vibes oh okay yeah just like crazy stuff happening it's good anyways you know what you were right about alien earth and the studio so uh okay you're also the whole season the episodes were like 35 40 minutes and the whole season ends next week it's like six episodes so it's pretty short okay okay it's not a big commitment to get into it okay i have a recommendation for you yeah predator badlands okay i gotta watch that yeah i know it's that's on streaming now i miss in theaters okay good yeah For me, that one was like, whatever, that's okay.
25:10Let's play this next one because this one was like, what did you say? Whoa, whoa, whoa, wait, Joey. What? You're taking the technology for granted. It's doing voice. You know how hard it is to do voice like that? Yes, sir.
25:23Okay, all right, let's go. All right, next one. Joey and Addy on the airwaves. Talking about the future that tech saves. AI's making all the movies now. trying to figure it all out somehow it's a digital world going faster we're all just hoping it's not a disaster yeah it's just a noise in my headphones the noise in my headphones it's at least as good as blink 182 yeah maybe it's not sure what was your input for that dude this is crazy that it interpolated what we do on the podcast because i just said uh make a song about joey and addy doing a podcast i didn't say anything about filmmaking or ai it's literally nothing else i don't know how i figured it out do you use gemini a lot or no because i'm wondering if gemini's memory no you know just had some other previous context saved uh no no not no that was a brand new session yeah that's it's bizarre but yeah kind of creepy wow all right the song itself was whatever but now knowing that you didn't really give it much context yikes it is sentient after all world models that anthropics he was right man i mean like what he said like we don't know if cloud is sentient or not that's kind of a scary thing to say what if it just sings the song is like help me i'm trying to escape what if it just pretends to be dumb right like that's what we all fear like like asi or it's already here the models just pretend to be dumb until we all connect our open claws to the internet and then it's like strike yeah it's sophisticated it knows when to attack of course yeah it has strategy all right you got one last one that you made for me oh yeah this is a nod to joey's miami heritage
27:41Do you speak Spanish? I was going to say, if I spoke Spanish, I could maybe get better assessment. I bet you got bizarre details of your life spot on if you translate that. I'll try to translate it later. But the thing that I thought was really funny was in one part of the song it went, Yo-i and Miami. This is, because you've said this before and I listened to it, it reminded me of this issue that would happen at Starbucks if I went to, like, Hialeah Starbucks or Doral. They call you Yo-i? this is how they spell my dude the fact that this model got that right that's creepy so if you're just listening to context well first off in spanish j could sound like a y sound so if it's like a very spanish area i'll get my name pronounced yoey and the starbucks cup spelled it um y-o-w-i the w is is just the best yeah yoey yoey so um yeah i i did really enjoy that from the song that it got that part awesome hey it sounded like uh like daddy yankee ish kind of it seems like it keeps failing when i try to give it just the thumbnail of denoised but i had one song so i gave it a slide deck from an old notebook lm session when i was researching for like our comfy ui episode and just sort of a denoising and images in general so i kind of just gave it the explainer video of how ai makes images and then i didn't give it i just gave it the video of the deck and then that was the prompt.
29:17So let me play what that made. You had to give it the most boring thing, didn't you? This is exciting stuff. We've covered all these different components, diffusion, units, VAEs, text encoders, and it can feel a little abstract, but this isn't just theory. There are tools that let you see this entire process visually. What if you could actually connect these pieces yourself like an old school switchboard for your own creativity
29:53it sounds like a disneyland ride it sounds like a disneyland ride or it sounded like the intro of a song uh or the intro of a keynote speech that would have been in uh like silicon valley yeah yeah Gavin Belson's company, yes. I don't want anyone to make the world a better place than we do. Better than whatever the line. I wish that show would come back. I know, they really need to, yeah. Hopefully, I mean, Netflix is so good at buying other IP and then just bringing it back to life. Resurrecting it, yeah. That one just feels like, especially with everything going on now, there's just so much material.
30:34now is the time alright well that was a fun ending yeah I mean for me this feels like the Gemini music thing kind of feels like Sora 2 where it's just like fun meme songs nothing that I would put in production right now the big question looms is it commercially safe if you generate the track do you own it some of the enterprise stuff that needs to be legally worked out. And knowing Google, I think they're already thinking along those lines. So I wouldn't be surprised to see a lot of Lyria generated music on your deck or on your video on your pitch in the future, because it's all built into Google Slides, Google Docs, and so on.
31:17Yeah, I mean, we're using this now in Gemini. But like, to your point, I could see this easily being something where you like make a slide and Google Slides. And then it's like, hey, you want some like background music for this? And then you don't even know it's called Lyria. Yeah, you're just like, yeah, make music. And then it's like, okay, cool. Right. And it's just like going to be another hidden product in their product line. All right, cool. Good place to wrap it up. We haven't gotten a five-star review on Apple Podcasts in a long time. It's been months. Please give us one and it'll help us more than you know.
31:43Yeah, I think we're doing pretty, YouTube's probably our biggest platform. And so we do get a lot of comments there and we appreciate all the comments. But if you're feeling generous, going over to Apple or Spotify and leaving a review there is super helpful for it to get some more reach over on those platforms. Thanks for everything we talked about at denoisedpodcast.com. Thanks again for watching. We'll catch you in the next episode.
From the publisher
Addy and Joey break down the latest batch of open-source AI image models: FireRed's specialized editing capabilities, Recraft V4's enterprise-grade output with SVG support, and ByteDance's newest open-source offering. They also cover Foundry's acquisition of Griptape, an AI node-based platform that signals where VFX compositing is headed, and test out Google Gemini's new music generation feature Lyria 3, which creates songs from unexpected inputs like slide decks and video thumbnails.
--
The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.




