In short
Podcast Episode Summary: Denoised - The Topaz AI Workflow That Turns 8-bit AI Into 16-bit Cinema Quality
Episode Overview In this episode of Denoised, hosts Addy Ghani and Joey Daoud discuss new developments in AI and virtual production, focusing on insights from the Infinity Fest panel regarding Gaussian splats and AI-driven workflows in filmmaking. Joey elaborates on the topics he moderated at the festival, highlighting innovative technologies and the potential they hold for the media and entertainment landscape.
---
Key Topics Discussed
Infinity Fest Highlights
- Event Overview: Joey recaps the Infinity Fest, where he moderated a panel on AI innovations in virtual production.
- Panelists: Featured industry experts included Paul DeBevick (Netflix), Jason Shugart (NVIDIA), and others.
- Focus on Gaussian Splatting: The panel discussed the application of Gaussian splatting in creating lighter, faster 3D environments in filmmaking.
Gaussian Splatting
- Definition:
- A method for reconstructing 3D objects using floating blobs of data rather than traditional polygons.
- It allows for faster processing and requires less computational power.
- Real-Time Rendering:
- Gaussian splats enable real-time rendering by placing color blobs in 3D space that blend seamlessly.
- 4D Gaussian Splatting:
- Captures motion in 3D environments, allowing for dynamic camera movements in reconstructed scenes.
Comparison with Other Technologies
- NERFs (Neural Radiance Fields): Outdated in comparison to Gaussian splats, which have been rediscovered for modern applications.
- Limitations: Current tools for editing Gaussian splats are still limited compared to traditional polygon-based workflows.
Practical Applications
- Virtual Production Use Cases:
- Tools like Valinga and PortalCam facilitate the incorporation of Gaussian splats into Unreal Engine, facilitating more efficient production workflows.
- Examples include scanning significant historical sites (like Auschwitz) for digital reproduction without physical disruption.
Upscaling AI Outputs
- Topaz AI Workflow:
- A pipeline that enhances 8-bit outputs to 16-bit EXR files for cinema-quality results.
- Involves multiple Topaz models to denoise and enhance footage, achieving higher detail and reduced banding.
Bit Depth and Dynamic Range
- Key Insights:
- The importance of bit depth in capturing a wider dynamic range in darker scenes.
- The role of noise profiles in digital cinematography, with techniques employed to replicate the natural imperfections found in traditional film.
---
Discussions on AI Tools
- Nano Banana vs. Seedream:
- Nano Banana excels in photorealistic outputs and precise image editing, while Seedream is better for creative, novel outputs.
- Workflow Enhancements:
- Joey discusses his use of various AI tools for website design and image generation, emphasizing the importance of clear prompting strategies to maximize output quality.
Call to Action
- Audience Engagement: The hosts invite listeners to submit questions for upcoming episodes, enhancing interactive discussion on AI and media production.
---
Conclusion This episode of Denoised provides a deep dive into the cutting-edge innovations in AI technologies that are shaping the film industry. Through engaging discussions and expert insights, Joey and Addy offer a compelling overview of how advancements like Gaussian splatting and AI upscaling are transforming storytelling in media.
For detailed show notes and resources discussed in this episode, visit [denoisedpodcast.com](https://denoisedpodcast.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00All right, welcome back to another episode of Denoised. We are going to focus on a bit of a recap of Infinity Fest. I hosted a panel there on 4D Gaussian splats and AI and virtual production. Addy could not make it, so I'm going to fill him in and fill all of you in while we're doing it. And then we're going to talk about some of the other things, the updates we've been seeing in AI filmmaking workflows. You ready, Addy? Sound good? Awesome.
0:24All right. So, Infinity Fest. It was last Thursday and Friday. Super cool event. A lot of really good panels. a lot of good talks focused on the panel that I've moderated. We had a panel on AI innovations in virtual production, but that was the real thing that we were talking about a lot was Gaussian splatting and 3D, 4D Gaussian splatting. We had Paul. Look at you moderating big panels. Big panels, yeah. Boy Cavalier himself, Joey. I was curious as well. I was like, I'm just going to ask you questions. I want to know as well. And there were like some good teasers at the end too of things that I wish we had more time to get into because I'm also still curious about them, like triangle Gaussian splats.
1:08Yeah, and I guess the audience, I mean, remind all of us what the AI connection to Gaussian splats are, because we tend to think of that as traditional photogrammetry, even though it's clearly not. It's different, but it's also not AI. So on the panel was Paul DeBevick from Netflix's Eyeline Studios, and then also Jason Shugart, who works at NVIDIA now, but has an extensive history in VFX. Yeah, I know Jason. And they made clear that Gaussian splats is not directly AI, but it is related in that field. So I think the correlation there is Gaussian splats are not generative AI. They're not creating something novel from scratch.
1:50But the solve itself, how the point splatters, the Gaussians, are correlated in three-dimensional space is using an AI engine. And yeah, also just like a short explanation of what this is and how it's useful. Basically, if we kind of compare how we would rebuild things in 3D, the original kind of way that is still pretty standard is polygons, a bunch of little lines connecting everything and you're building out your shapes, but that can use millions or billions of little dots and polygons to build out whatever 3D object or space you're building out. Gaussian splats, a splat is a floating blob with a variety of information inside it.
2:32And then if you have millions of these blobs with information inside it, that can rebuild your 3D space. And the advantage of this, or 3D object, the advantage of this is it can run much faster, much lighter weight than if you're rebuilding the same scene with polygons. Correct. Yeah, I think Gaussian splats are still using a sort of quote unquote, real time rendering, like the fact that when you are moving the camera, it's calculating that novel view in real time. But instead of rendering per se, like with ray trace, or, you know, bouncing light off of surfaces, the way traditional computer graphics does, what it's doing is it's placing these blobs of color in the three space that corresponds to that object.
3:16And because they have a Gaussian splat, you can actually put them right next to each other and they blend in like a seamless object. Yeah. And in between this two, there were NERFs, which kind of had a little bit of a moment, neural radiance fields. But the thing with NERFs was you had to process the data and sort of got baked in. Is that kind of correct? Where like Gaussian splats can kind of be run in real time because the data is in the splat? again that's beyond my understanding but i i think you're right in that nerfs are an outdated technology and most of the development work that's happening in this realm is now happening in 3dgs and the interesting thing was that gaussian splats are not new it was like a based on a paper i think 15 20 years old but they sort of just rediscovered that it has this application in 3d space and 4d space which we also talked on a bit too so i mean the idea of 4d gaussian splats is you are capturing not just a 3D environment, but motion in a 3D environment, so people dancing or moving, and you're able to reconstruct that scene in four dimensions of 3D space over a period of time and reposition your camera, move your camera in that 3D space.
4:27Yeah, so when you say 4D, you're not talking about Interstellar, like the end scene with... I'm talking about 4D as in Muppet Vision 4D. it's three dimension over time yeah and so right and there was that demo clip i think we talked about on the podcast like months and months ago of remember that clip of the guy sitting down i think it was like a chinese youtuber and he was sort of sitting down and then it was like a 4d gaussian splat and the camera was sort of moving around him and then it kind of went a little bit viral in our nerdy sphere because people were saying oh it was generated from like one video angle and then um it came out where i was like no they needed 20 cameras to pull that off but it was still you know it's still cool but they needed 20 cameras yeah we we never caught the bts and i think i was um like my friend jim gadol did kind of pointed that out like just so you know this is as involved as traditional photogrammetry with like 50 cameras when it was the company behind that technology uh they were demoing at infinity fest as well i think with a 4d view or something and they had the rig and that this demo clip set up but they've been doing a really clever solve solution for building out a rig with a variety of pcz cameras where they can capture 4d gaussian splats uh and that like so they're not building the tesseract per se what was that that's the thing from the interstellar i mean i haven't seen that in a long time what was the tesseract i don't know if this is carl sagan i don't know if there's some real carl sagan described four dimensions using an object called the tesseract oh that's right that I was in that Black Matter, Dark Matter book as well.
5:59Yeah. So like, you know, like a cube shadow is a square and a tesseract shadow is a cube. Okay. Does that break your brain? I'm trying to. You can't picture it. I'm trying to wrap. Our brain can't picture it. No. Yeah. Okay. But yeah, when we talk four dimensions, that's where it goes. So loop us back into why Gaussian splats are important for modern filmmaking. Okay, so bringing this back into tangible benefits and also some of the other panelists and what this had to do with virtual production. Also on the panel, we have companies or people like Fernando Rivas, who's the CEO of Valinga. I did a big interview with him at NAB, so there's extensive stuff on that.
6:38But he basically built Valinga, which is an Unreal Engine plug-in. You can take a Gaussian splat scan that you do, and you can make these through a variety of techniques. but we also talked about PortalCam from XGrids, which is sort of the latest device that kind of does this all in one and makes it easy to do. You can scan a room. We covered that. Yeah, we covered that. Yeah, we're talking about, yeah, PortalCam launch about a month or two ago, you can take your scan of a room, bring it into Unreal Engine with the Linga, they can load Gaussian splats and load it into Unreal. And they also have ACES color management.
7:08And then you can push that into a wall. And basically, you can have your environment based on a real life environment, scan a room, scan a set. They've also been using that a lot as backup. So you scan a set from production, if you have to do pickups later, or if like season two gets picked up and they weren't quite counting on a season two or variety of reasons, you can take a scan of whatever space you want and put it on a wall. And so that's like a very practical way. Also, they had a case study, two of them, that Valinga did where they scanned Auschwitz, the concentration camp. And that one was they just did not allow anyone to film at Auschwitz if they're trying to do any type of production.
7:48They did a scan and then now making that scan available to productions if they're doing a project around Auschwitz or that wants to use the location, they can put it on a wall and have that space digitally so they can film there digitally, but not disturb the actual site with production. There's so many places in the world where bringing that place into a volume or into VFX would be super beneficial. I mean, Auschwitz is one of them. Even something like the pyramids where there's all types of permits and time limitations. Tourists. Or tourist inconvenience. Chernobyl, you know, like the show. Yeah.
8:22Like if you were to recreate that in a volume, that's probably the best way to do it. But I think the advantage over traditional photogrammetry where you take hours of video or thousands of images is that there's a big solve process. And a lot of it is hand done. Traditionally uses reality capture, which is now acquired by Unreal Engine. And you're talking about not just running it on a super fast computer with a lot of GPU power, but an artist manually going there, cleaning stuff up, aligning things. And it's tedious work. versus a Gaussian splat, you upload the images. And if you shot them correctly, it should more or less, quote unquote, solve for itself.
9:02And then you get this sort of super lightweight file back, which is like 100 megabytes, not like 10 terabytes or whatever. Yeah. And that could then in real time kind of sort of render it as your camera moves. Yeah. I mean, other practical use, you know, we've seen with Lightcraft Jet Set, that their value prop is you can use your phone as a camera tracker, and you can load in your 3D scene into your phone. So when you're filming, you can roughly comp in and see your 3D space. You could bring in Blender and Unreal scenes, but they're going to be kind of heavy. But you can also bring in Gaussian splats into Lightcraft jet set on your phone and see this in real time on your phone.
9:40And because the splats are much more lighter weight, you get a better experience when you're filming. I'm curious to know what Paul DeVebbik said at the panel, because as you know, he's one of the most expected folks in VFX today. So I'm just curious what he said. So yeah, I mean, Paul, he kind of talked about a lot of the basics that we sort of just covered now of like what is Gaussian splatting, both 3D and 4D. He also sort of alluded to a paper that they've got coming out at SIGGRAPH Asia on super resolution techniques to achieve extreme close-ups from large volume capture. So we didn't get into more detail about that, but that's a paper that they've got coming out soon.
10:17And they also sort of talked about some of the limitations of Gaussian splats. So the tools to edit a Gaussian splat, like a scan that you make, they're still limited, you know, definitely not as much as the existing 20, 30 year pipeline of working with polygons. Yeah, the other thing, the other limitation with Gaussian splats is they are ratings field representation. So basically the light that you captured during your scan is sort of baked into it. And so also relighting Gaussian splats is trickier. That's something where currently right now, something like a combo of Gaussian splats, LiDAR technology, and traditional photogrammetry kind of comes into rebuilding a space if you need to have that flexibility of relighting.
10:54That's something that Global Objects specializes in, where they sort of combine all of these techniques if you need to have a space where it's an Unreal Engine scene, 3D, but has the flexibility of relighting. That's one of the limitations right now with Gaussian splats. For sure. but the advantage to having baked lighting is a lot of the reflections are super accurate and it just looks like photogrammetry a lot of times totally messes up like glass or metal shiny objects whereas gaussian splats are really good at that stuff but then again it's baked in and the other thing he sort of alluded to but we didn't get into details and so maybe we'll try to break this down in the future was um you know we've gone from polygons we've gone to nerfs we've gone now we're looking at gaussian splats what's after that and triangle gaussian splats are what was alluded to but i didn't really get into details about that and i have not looked it up but i know some papers out about that uh i'm guessing it is a combo of polygons and gaussian splats yeah i would imagine tries would be more computationally efficient for gpus because gpus are really built for tries and quads, like the math behind.
12:04Just computing a million or a billion tries at a time in parallel versus a Gaussian splat, I think is a sphere. So geometrically, it's more complex. Okay, that makes sense. So moving to tries would give you infinitely better resolution. So that would be like sort of the data stored instead of like a sphery blob, more of a triangly blob? Triangly blob, but that triangle is a flat triangle, which is a natural unit for modern computer graphics. All right. Okay. Okay. Interesting. That's my guess. We'll have to dig into this in the future. Okay. And then also just to round out who else was on the panel and what they covered, we had Ben Abergel who works with ETC and is a creative technologist involved in ETC.
12:47And we'll talk about the Bens in a bit, but he worked on Pathways, another project that ETC is producing. and they were initially trying to do some form of 4D Gaussian splats with a few iPhones. I believe they recorded a scene with like four iPhones, but it just was not enough data needed to create a Gaussian splat that you could like move around and still have good quality. So they shifted and they kind of shifted to a workflow that was more something we've talked about before of like iPhones, Lightcraft jet set, people, real actors, green screen space, and then restyling the background with like a kid bashed kind of Unreal Engine scene, and then restyling that with video to video with AI later on.
13:33The thing, going back to the original 4D Gaussian splats that they're trying to do with the iPhone, I did ask like how many iPhones would you actually need to successfully pull that off? And he thinks probably around 20 to 30, the same number that the other Gaussian splat videos are using. And then lastly, Ella Roberts. She was the AI supervisor for the Wizard of Oz upscaling at the Sphere, or upscaling's not the best word. Magnotis, I believe, did all the work. Yeah, she works with them. She's the creative AI supervisor. And yeah, in broad sense, we kind of talked about the workflow and pipeline of what they did, but obviously couldn't get into too many details.
14:12But yeah, I mean, And that was just sort of a whole different project in taking existing formats and sort of breaking them down from the original film and then rebuilding out each scene for the 16K massive landscape of the sphere. Yeah, it's a nice cross section of the industry on that panel. I mean, you have everybody from like tenured, you know, one of the foremost experts like Paul all the way up to like you on the AI frontier side, Gabriel on the, you know, sphere and location entertainment side. And then you also have NVIDIA. You have Jason from NVIDIA. No, it was really good. It was a really good.
14:50It's a really cool panel. I'm bummed that I missed it. Yeah. Yeah. And shout out to Corinne for assembling that panel. But yeah, I also want to talk about one other thing. while we're on a FinFest because there's another really interesting presentation about upscaling regular AI outputs, which are usually 8-bit RGB 720, maybe 1080 output to 4K cinema quality 16-bit EXR files, which we've talked about how like, oh, RAV3 from Luma, their cool thing is they can do that in the output and the generation. But what if you don't want to use RAV3 or you have other outputs what is this workflow and so what if you just want to upscale you already have the generation done exactly and so this is a really interesting workflow where the uh that etc's other project the bends which is sort of this kind of uh cute animated short with like a fish uh deep sea fish underwater kind of traveling around and making its way to the surface they built out a pipeline using a variety of Topaz's models.
15:54The local model has nothing to do with Starlight. And so basically they would take their source footage, they would run it through Nick's denoise, kind of strip out the noise from it, then run it through Gaia Operates, which is one of their other models. And then there were a couple other steps in the process, which I don't have in the screenshot that I took. But basically it was like three different Topaz models. And then out of that, they were able to get much more detail, much higher quality EXR file, 16-bit, that in the color grade and everything else, they just had a lot more latitude to adjust the shots, match the shots, something that you would be used to if you were shooting with a cinema quality camera.
16:35Yeah, it's really interesting work. The thing that fascinates me more about the resolution upscale is the bit depth upscale. So, you know, when you talk about typically 8-bit is 0 to 255, right? So if you have, you know, your sun in the frame, the highest value that sun will have is 255 versus real world dynamic range. Like what our eyes see, you know, is it won't be 255. It'll be like 1 million units, right? Yeah. Like how do you get that crazy of a dynamic range and how do you create it artificially interpolated just from that? right from having this like kind of crappy source frame that you're able to like pull and generate all this data yeah it was really impressive and i wish i had a shot of the waveform monitors because they had before and after the waveform and that's where you could really see the latitude you get and also the other advantage is like it reduces banding a lot you know it's a banding it's like just because it's limited if you're an 8-bit it's limited amounts of color space and so right especially because this short film took place in like the deep sea so there's a lot of dark shots And so even with the kind of gradients and some of the shadows and the dark, you would see the like banding lines and this upscale process.
17:47And you could see it in the waveform monitors where it went from like very rigid steps in the waveform monitor to a much more gradual curve, which is what you would want to see or normally see where you get more of a nice gradient with color space. And I think a lot of the sort of assumptions about having darker scenes with less light, you know, I'm thinking Game of Thrones, like all the scenes, right? It's just like you think that you don't need that big of a bit depth because things are dark. When in reality, our eyes are actually much more sensitive to shadows and differences in the darker lights and the regions that if you throw more bit depth at the darker regions, you actually get more detail out of it.
18:32So it still benefits the scene, even though there is not a sun in it, for example. It's not a daylight scene. If it's a dark scene, you still need more bandwidth in the bit-depth area. Yeah, and also just when you're in the color grade, too, and you have that control and flexibility. And what the experiment they're doing, too, is they're trying to really build out an AI generative workflow, but in a professional Hollywood filmmaking pipeline. So this entire process was supervised by an ASC cinematographer, Roberto Schaefer. and it uh you know he's using his eye to analyze the scenes and like try to get what he's used to with shooting with an ari alexa or sony venice uh out of these ai outputs and then the other thing they did was they did a grain film pass with photochem and they showed some kind of comparisons of like with grain and without grain and just you know adding the grain just makes it way better way better yeah for lack of a better word well one thing grain really helps with is banding like uh you know there's a process called dithering which reduces banding so it's essentially adding noise into the frame and grain is kind of noise so it does help overcome a lot of the limitations with ai having you know 8-bit and low dynamic range yeah and i thought he said something really interesting too because like when i thought about adding grain to an ai image you know or even just a digital image part of it has felt like oh is this cheating or is this you know is this doing something like you know faking it because like it you know wasn't actually shot on film or adding film grain to it but he said something interesting where he's just like it doesn't matter what you shoot with even if you're shooting this with digital cameras like every sensor has a noise profile like they all like nothing is shooting nothing is surgical clinical yeah yeah like and if you're doing you know even an animated cg film if you're doing outputs and it's just like too clean because like that's the only way you could get generate something that like had absolutely zero noise or grain and it just looks too weird and too clean and adding the grain in that process just you know helps sell the illusion and it's not like uh it's not like oh you're adding fake film effect it's just you're adding something that replicates even if you shot it with a digital cinema camera yeah you know the the grain part of cinema is so important that a lot of cameras now have built-in grain adjustment.
20:51I don't know if the RED cameras have it just yet. Perhaps the V-Raptor might, but the Arri Alexa's for sure, you can not only pick the level of grain, like a 0 to 100, but also the texture of the grain. Okay, and this is separate from like picking your ISO setting or whatever? Yep, exactly. So this is something that the camera body will add after acquisition. So it'll like kind of overlay it. And I believe if you're in Arri RAW format, you can tweak it down the road as well that that would make sense all right so yeah those were the highlights for me from uh infinity fest i mean a lot of the great panels but uh i know we've talked about this for a bit but those were the two standouts for for me but yeah it was a good good conference a lot of a lot of really interesting talks let's talk about some of the tools people are using all right what are you using what have you been your go-to tools lately so i tend to focus more on image generation stuff over video generation so i'm like hyper aware of all of the nuances and the limitations of image generation.
21:54For example, I use Nano Banana, like, I don't know, I make maybe 50 generations a day on Nano Banana just to kind of figure out what it can and can't do. Are you doing, when you're using Nano Banana, are you doing like just straight text to image or are you giving it any input stuff? I'm doing mostly image editing so i'm going through a bunch of so i'm on the personal side i'm building out a website and for the website i'm going to take a bunch of images of real world places for example just if you think of the huntington library right like beautiful place if you just take photos of it with no people in it how do you bring it into nano banana and then add the people that you want like populated with the crowd or hero characters and things like that so nanobanana has been really good at that you could give it a reference of each of the people that you wanted to be filled and then with like the markup tools you can actually dial in where they'll be sitting or standing and it preserves the backplate really well like it doesn't distort any of the actual photo and then when it puts on the new person or the object sometimes i'll put a car in the car will have the reflections of the environment, like an HDR almost.
23:12That's cool. Yeah. Yeah, so it's doing a lot under the hood. And I really have not seen an image model that capable at doing that task yet. When you said markup tools, are you using the markup tools built into FreePick or are you doing something else? FreePick, yeah. Yeah, okay. And that works really well for precise notes and what you're looking to create. For sure. And then the other thing that I'm doing is obviously with website design and build, And I'm not like a web builder. Like I'm just doing this because I need it. It just needs to get done. And I figured just this world of AI that we always talk about, why not put it to the test and actually use it?
23:50So I've been using Ideagram 3 for a lot of the logo and text. Oh, okay. And even motifs and stuff like design elements of the website. You know, like for example, if you have like a tractor or something and you want a hand-illustrated version of that tractor, you can put it through Nano Banana, get a tractor back, and then put it through Ideogram and make a font in that style and have the text reflect the logo and live in the same sort of ecosystem. That's cool. Have you done anything where you've had an AI mock-up of the website and then give it to Claude or something to make the website based on the mock-up?
24:32So right now I'm playing around with Framer AI. Framer is like a known tool like Wix or Squarespace where you can build websites. So they have an AI engine built into it that's quite good. It'll give you a skeleton of a website and then you can go in and you can add animation, add text, add another page and so on. There's just so much, so much AI tools at your disposal. Have you done a comparison with Banana Banana and C-Dream? Because I feel like C-Dream is probably the closest. C-Dream 4, where you can give it 16 or 20 input images. All the time. Yeah. How have you found any pros, cons? I would say Banana Banana is more built for, quote unquote, production.
25:18Like, it has a more photorealistic output. It tends to understand compositing better. It tends to understand what I just mentioned, like image-based lighting and some of that stuff better. Seedream, I would say, is better for creative, completely novel outputs. So if you're creating the world from scratch, I think go to Seedream and get a weird, wacky view of the world that you never could have imagined. Versus if you already have a photo and you're just trying to do essentially Photoshop on it, then use Nada Banana. I will say I've also found this tip a week or two ago from Henry Dobrez, 4nano banana, where, you know, sometimes I've had issues when I'm trying to give it an image and then change something in it.
26:05And then the prompting, sometimes you're like, change this or, you know, replace this. He says that the best prompt is show me. So show me an aerial shot. Show me a side shot of that person. You know, show me blah, blah, blah, blah, blah. So I've been using that more and I have found that that does work pretty well. I love that. As a prompting trick. I'll keep that trick in mind. Yeah. And the, I mean, the other thing with like C Dream and Nano Banana both is like you can prompt it like a thousand sentences, no problem. It'll like swallow it up, you know? Whereas I think there was like a 500 character or 500 line limit on the text encoders before.
26:43So if you're in comfy UI land, if you use like T5 or XXL for the clip encoder, I think those things have a 500 character not 500 character 500 token limit whatever that means oh okay yeah versus like nano banana it doesn't really tend to have a limit you could just keep going do you have you found like longer prompts work better so a lot of times i'll go into chat gpt and i'll i'll take a small prompt and i'll i'll be like hey can you boost the prompt that's literally what i say and it'll come back with like a really verbose prompt and then i'll go back in and I'll adjust some keywords. And then I'll have this giant paragraph that I'll then cut and paste into Nana Banana.
27:24Yeah, same. I use that for creating prompts. I'm not going to write a long prompt. Yeah, the other thing I've been seeing pop up a little bit more on X. And I mean, now this could take us for granted because it's coming from, this is from M, who's the co-founder of Scenario, who that is a website that makes custom, helps you make custom models, training models. So he is selling the pickaxe. But I've been seeing some demos and talk of LORAs and if LORAs are still relevant and where they still come into play. We've talked about this before of just like with Nada Banana, Seedream, do LORAs still make sense to train some custom model to get the specific output you want consistently when you could just prompt it in something like Flux Context or Nada Banana or et cetera?
28:07And so this was a demo from him of turning something into an isometric view. and basically his argument still for doing a laura and specifically a laura float for flex context is you still just get much more consistent outputs at scale and so this is sort of a demo uh isometric view of a building compared to i'm guessing this is probably just a text prompt with an input image where you get you know kind of looks like it but if you're trying to do something consistent over a series of shots it they'll you know you won't get that consistency yeah so the argument is that you still need lauras for the highest level of control yeah and for i mean so the thing i've also been trying to wonder too because like to train a flex context laura you you don't just give it a group of images of what you want you have to give it a pairing of images you have to give it like input image output no no no you give it a small group of images but you'll have to caption them no i think you have to give it like input output pairs you have to of pairs to train it like this is what the input this is the output this is what like you have to train it with a pair of images how do you know what the output is if you haven't generated it yet that's the thing you got to make the output like so you have to modify you know 20 or so images to the output that you want in order to train it and so that's where i've been wondering it's like so do you just do that you know with flux context or not a banana yourself and just kind of refine it so it looks really good so you have that pairing or do you modify this elsewhere that's the part i've been a little fuzzy on because with with some of these models you have to give it the pairing to train you can't just be like here's you know not you can't be like um mid-journey mood board where you just make a mood board of like a bunch of images that you like and then that sort of trains the model of what type of output you're looking to get oh that's interesting i'm looking at Flux Context Laura training on Replicate right now.
30:07According to Google AI overview, Flux Context Laura training involves preparing paired images, start and frames of a subject, uploading them to a training service with a unique trigger word and then running the training process. Oh, that's interesting. Yeah. So you have to like modify like 20 or so images manually to what you want to train it to do that. Flux is like one of the OGs of image models, right? It was Flux and Stable Diffusion, and I think Flux is built off of Stable Diffusion, if I'm not mistaken. And so the way to train SDXL and some of the older models was to just give it 50 or so or 30 or so images with the caption, and then it'll train a low-rank adaptation, which then latches on to the foundational model.
30:55right but you're saying that now you need to have an output to attach to an input no caption needed and that's your laura training yeah like start point end point repeat 20 or so times and that's the data needs get a style like let's say you're trying to recreate like a very particular style of animation right like black and white animation so you give it a colored image as an input and a black and white image as an output? Yeah, I would guess so. But it's like, my question was just like, to make that black and white image, do you just do that by hand, using whatever process you want? Yeah, I would go to Photoshop and just desaturate it or something.
31:33I don't know. Right. But I mean, if you've wanted something like more stylized, so there's a project we're working on where it is like supposed to be an animation look. And what I've been doing was I found that if you try to create the frame, because it has it has characters, it has like, you know, we need to have specific locations with specific people in a specific animated style. I found that if I gave something like Not a Banana or Sea Dream, all of those commands, like people, place, in a certain style, it would melt its brain. And it just wouldn't give me the thing in the style. But I found if I had to make a photorealistic-looking image with the character and the place that looked the way we wanted to look, And then I've just been running it through Flux Context with a consistent text prompt of like change this into a painterly, you know, oil based painting style.
32:27And then the output image looks pretty good and looks pretty consistent. I haven't found a need to train a lore on that, but I'm thinking I could because I have the start frame and I have the end frame. If I wanted to simplify the process or if other people on the team were doing it to make sure it was like a more consistent process, I probably could just give it the start frame of the photorealistic looking image, the end frame of the stylized image and train a lore on that. It seems like a pain in the butt. Yeah. As long as you've been like, well, the text prompt works. It's like as long as everyone just copies and pastes the paragraph of the prompt and for flux context, like it works pretty well.
32:59So why do I need to train a lore? Yeah. I'm still on the fence. On the flip side, some of the closed down models that are not open source, you know, Nana Banana, Seed Dance, maybe. definitely you know chat gpt there is no lower training right so the only thing it does have is reference image insertion yeah or just trying to text prompt it if you need to change the style into that's where chat gpt or cloud comes in handy where you have it write a very detailed text prompt and the prompt is very detailed where if you keep running it consistently you get pretty consistent outputs yeah the jury is still out on which image model is the best and i don't think we'll ever get to like the most comprehensive image model ever uh i think i think it's always going to be which image model is the best for what you need to do and now you're kind of seeing those lanes being defined like i mentioned ideogram right like ideogram has got the logos and text and font stuff down like that's their lane uh i think nano banana the reason it gets so much attention is because it's just really good at photorealism but that doesn't mean it's the best model it's just really good at photorealism and then you have c dance which is really good for creative work and sort of animated looks and stylized stuff you know i down the road six months from now we won't be talking about what model is the best it'll be more like what model are you using for x and y and z maybe in a future episode we could revisit because uh i keep seeing stuff from x xai whatever their video generator is called yeah it's on version 0.9 grok yeah Yeah, Grok.9.
Read the full transcript
34:35I saw the Will Smith eating pasta done on that one. It looked pretty good. Yeah, and I think it's probably getting popular again because I think with the back and forth, the boomerang of Sora, like Sora 2 comes out, do whatever you want, FIP, and then a week later, a complete U-turn of like, you can't do anything. And now people are like, I wanted to do whatever I want. And then Grok is like, you can do whatever you want here. The door's wide open. We don't care. yeah and you gotta remember the whole sam altman elon musk beef right like they're they're gonna one-up each other at every step of the game as as we go on yeah yeah if only there was a an elon cameo on grok so you could have sam altman and elon duke it not on grok i'm sorry on sora so you could have sam and elon duke it out with logan paul here's an idea you could do it with traditional vfx you could do that too actually no i take it back you could do it with pretty much any open weight model you just take reference images of both of them it's pretty easy yeah you could do it with a one 2.2 animate i mean it's a pretty good demo problem recreating images of that fight because today i just made somebody holding james cameron like this and as i was uploading james cameron's image i was like are they gonna flag that this is a famous person and it was fine it didn't do anything hey denoisers so joey and i had this idea why don't we start answering our audience's questions.
36:01If you have a question about anything that we've covered in the past, certain AI models or what have you, or even stuff we haven't gotten to yet, submit a question at the following email, denoised at vp-land.com. Again, it's going to be on the text here if you're watching it on YouTube, denoised at vp-land.com. And we'll get those. It is vp-land because someone else is just sitting on vp-land and wants to pay it charge thousands of dollars for it. So until the show grows way, way bigger. Yeah. We'd love to hear your questions. Uh, we could do like a mail mail room episode in the future. Uh, questions about, yeah, just stuff that's happening or workflow questions, um, send them over and, uh, links as usual for everything we talked about at denoisedpodcast.com.
36:50Thanks for watching. We'll catch you in the next episode.
From the publisher
Joey takes you inside Infinity Fest’s panel on Gaussian Splats and AI in virtual production. Learn how this technology enables lighter, faster 3D environments and how ETC upscales AI outputs to 16-bit EXR files. Plus, Addy and Joey compare AI tools like Nano Banana and Seedream for consistent image generation.
--
The views and opinions expressed in this podcast are the personal views of the hosts and do not necessarily reflect the views or positions of their respective employers or organizations. This show is independently produced by VP Land without the use of any outside company resources, confidential information, or affiliations.




