In short
Codex (OpenAI’s desktop coding agent) using Blender MCP to generate a 3D Blender scene from phone photos of the Culver Theater; discussion of how AI workflows will keep 3D in the loop for camera control, plus related news on SIGGRAPH, ComfyUI/MCP, Nuke SmartRoto/Grip Tape updates, and AI video generation quality/costs.
Guests
No guests. Hosts are Addy and Joey (remote).
Key claims
Codex can install Blender MCP via prompt and build a usable “gray box” 3D model in ~10 minutes from imperfect photos, outputting editable geometry in Blender and even a render; GenAI will become the “finishing” layer while 3D handles precise camera/geometry; local inference matters for unreleased IP; AI video still struggles with lensing/perspective and consistent character realism.
Notable examples
Culver Theater model (rounded corners, art-deco approximations, separated geometry); NVIDIA static Gaussian splats; ComfyCode.ai (agentic ComfyUI MCP sidecar); Nuke SmartRoto; “AI Odyssey” trailer critique; ByteDance CDance 2.5 football teaser; Thinky Machines “Inking” open model.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAnticipation for SIGGRAPH
0:16 to 3:08
The hosts share their excitement about attending SIGGRAPH and discuss its significance.
“I know people enjoyed that we were in person, but sorry, we're back remote.”
Upcoming AI Workflow Summit
3:08 to 4:42
Discussion about an upcoming AI Workflow Summit and the presentations expected there.
“And then outside of the official conference, there's a bunch of satellite events, one of which, if you are in LA and want to come by, you don't need a SIGGRAPH pass for this, and it's free.”
Exploring Blender with Codex
4:42 to 9:38
The hosts discuss using Codex to create a 3D model in Blender from photos.
“So I'm excited to see what he's cooked up.”
The Future of 3D and Gen AI
9:38 to 14:00
A conversation about the role of 3D modeling in the era of Generative AI and the evolution of workflows.
“And look, you have all of your geometry separated out in the outliner.”
The Future of 3D Modeling with AI
14:00 to 18:10
Explore how advancements in AI and GPUs are transforming 3D modeling workflows.
“And it's still like, let me look at the computer.”
Innovations in Interactive Video Formats
18:10 to 22:00
Discuss new interactive video formats and their implications for digital content creation.
“Like, if you have a use, you need software, you need something, these like tools or solutions that spin up based on need because it's just so easy to spin something up from scratch that does exactly what you need to do.”
The Role of Open Standards in Software
22:00 to 24:20
Understand the importance of open standards in the evolving landscape of software and formats.
“All right, you want to talk about older software now?”
AI Integration in Established Software
24:20 to 28:01
Learn about the integration of AI features in established software like Nuke and its impacts.
“Roto stuff and matte stuff doesn't work well with it as far as I've heard.”
AI in Film and Performance
28:01 to 29:41
Explores the integration of AI in filmmaking and the importance of human performance.
“Yeah, it's like the same Tilly Norwood type headline, you know?”
The Potential of AI Models in Production
29:41 to 32:25
Discusses the challenges and advancements in AI model generation for media.
“I do think you need a film, like a cinema camera with real glass and real film back because the perspective, the distortion, the sort of geometry of the real world, AI just can't get right.”
Show all 13 chapters
Open Source AI Developments
32:25 to 38:29
Covers recent advancements in open source AI models and their implications.
“If you give it a wide shot and then close-ups of things, will it understand, like, okay, this is how everything relates, but the action should be here and things should happen here.”
Future of AI in Creative Workflows
38:29 to 41:03
Examines the impacts of AI on creative production workflows and local solutions.
“Yeah, I wonder why Thinking Machines launched their thing open source.”
Wrap-Up and Farewell
42:00 to 42:12
Concluding remarks and gratitude to the audience.
“A lot of it is like the big keynotes and stuff, which go on.”
Transcript
Automatic transcript. May contain errors.0:00Can you set up Blender MCP? I gave it these photos and I said, can you build the Culver Theater and the surrounding area in Blender using the Blender MCP? So I'm like, okay, go do the thing. And then here's the thing. Oh, wow.
0:16Welcome back to Denoised. Addy, how you doing? We're not in person anymore this time. I know people enjoyed that we were in person, but sorry, we're back remote. It's too hot to drive over to that side of town. Your tires will melt. All right, so we got SIGGRAPH coming up next week. Yep. I'm excited about this. This is my first SIGGRAPH. It's very convenient that it's here in LA this year. Are you going to hit it up? Yeah, I'll be there Monday morning. If any of our viewers are there, love to say hi. Yeah, let us know. Just look for me. Yeah. Let me see. There's the Academy Open Source Day on Sunday.
0:49So I'll be attending that. And then, yeah, most likely Monday, Tuesday. Okay. I might join you for the Sunday one, actually. Yeah, come by the Sunday one. Have you been to SIGGRAPH before? Oh, so many of them. Yeah. Yeah, I mean, it's such a mainstay conference for those in, especially on the creative technology side, the bleeding edge side of media and entertainment, because a lot of the newer research papers get announced there from universities, research labs, and then the new tools or the upcoming tools from the companies get announced there. It's like NAB for the Uber nerds, if you will. Yeah.
1:28So like, NEB has like the big releases from like the big companies. But then that like one thing they're really not sure about, they're really experimenting and betting big on, that's going to be at SIGGRAPH. Is it more of a talk kind of conference or a expo floor conference or like both? Yeah, I think the bulk of the floor space is actually the research entities presenting their findings. So like you'll have university graduate students and you'll see stuff that you're like, oh, my God, this is like five years out, 10 years out. Like, for example, the early days of VR, that was the first place I saw haptic touch stuff.
2:07Was that like SIGGRAPH 2016 or 2015? Yeah. So like some stuff will fizzle out and never really make it into a business, but then some stuff will stick. It's a glimpse into the future of media and entertainment, in my opinion. Yeah, I'm excited to check it out. There was like an LA SIGGRAPH event a few weeks ago, and Netflix did a teaser of some of the stuff they're going to present. And some very cool papers and some very cool research that they're doing that I'm excited to see get out there. I can imagine Netflix research announcing new stuff there. It's also very strategic on their part because you have a high density of really amazing AI researchers.
2:49So if you want to hire them, if you want to woo them, that's the place to do it. You have recruiters out there at the same time. If you're looking for work, I think that that's a good place to connect. So all in all, the fact that it's in LA and it's primarily a media and entertainment-based thing, those two things will play well together. Yeah, yeah. I am excited for it. And then outside of the official conference, there's a bunch of satellite events, one of which, if you are in LA and want to come by, you don't need a SIGGRAPH pass for this, and it's free. We are at the, it's not the sexiest name, but the AI Workflow Summit.
3:24Yeah. Under the hood. There's going to be a lot of cool stuff here. It's at, not officially KinoFlow, but it's at their satellite. Yeah, they have giant warehouse space in their bank. RD warehouse. It's a cool stuff. Yeah. But there's going to be some cool stuff. Some presentations on static Gaussian splats from NVIDIA, real-time volumetric workflows. And one that I'm very excited and curious about is Ramiro is going to be, He's been doing a lot of work with ComfyUI and agentic ComfyUI workflows. And so he's rolling out ComfyCode.ai, which, from what he kind of briefly told me, but we'll get a lot more info there, is kind of like a MCP sidecar for Comfy that sounds crazier than the official Comfy MCP.
4:11This one could work local. This one has much more of the agentic backend that can just do some wild ComfyUI workflows and integrate with Sempti 2110 and all sorts of other higher level production pipeline stuff. Yeah, it seems more hardware oriented and more ready for tangible real world use in a video production world rather than comfy being like a software running on the cloud kind of thing. I've known Romero for a few years now, and since the time I met him to now, he's always been on the fringe of the extreme innovation side. So I'm excited to see what he's cooked up. He's a brilliant mind.
4:48Ramiro has been pretty wild. I mean, I see his like to post of, he was like an early OpenClaw user. And he's like, I got OpenClaw spinning and coding like five different agents at once. And so I'm very excited to see what he's got shipping. And then at the very end of that, we've got a panel that I'm moderating. You may be on it. I don't think I'll be on it. I'm bummed to miss this one. But yeah, yeah, half of the noise representing. I'm sad that you won't be there. But yeah, we'll be there. or yeah half of us will be there um so yeah this uh we'll post a link it's a free event you don't have to have a sigraph pass it is tuesday evening there most likely will be food and drinks but i can't promise that it's not my event as long as there's something cold to drink there should be something there let's talk about blender okay so open ai released gpt 5.6 there's three different version.
5:40Soul is like the beefiest one. And it's sort of the equivalent of Anthropics Fable 5. If you spin up Codex, one of the things I had seen people do online was like, I just gave it images and I'm like, Hey, go build this thing in 3D, go build this thing in Blender. And it did it. And I was like, wow, all right, let's, let's try that. Yeah. Elaborate on it. It did it. Yeah. I mean, this is going to be the done, the quickest tutorial. Cause literally all All they did was you download Codex, which is... So Codex is their competitor to Cloud Code. It's like a desktop application. Codex is their competitor to Cloud Code.
6:15It was a separate app, but I'm stopping myself because now I'm looking. I believe they actually got rid of it and they just rolled it all up into the ChatGPT app. So basically it's a copy of the Cloud app where Cloud app has regular Cloud side and then has Cloud Code side. ChatGPT has ChatGPT side for normal chats. And then it has Codex, which is their coding agent. and also it just can do more things on your computer so literally all I did was I started a new project and then I said can you set up Blender MCP and then it anybody anybody can type that in it's it's a third-party tool it found it and installed it and set it up yes it did wow and then for this test so I had some pictures from the Culver Theater from uh on the lot and uh I I gave it some of these photos.
7:04Now, these aren't the best photos because they're one-sided. I didn't get a good angle of the corner and the other side. But I gave it these photos, and I said, can you build the Culver Theater and the surrounding area in Blender using the Blender MCP? Stop there. I just want to say how difficult that would be for a modeler if you go to the picture. It's a lot of art deco, intricate architecture. It's not everything. It's not right angles, right? um like if you look at the four things holding up the sphere up on top like that that's a pretty difficult thing to model at least for me yeah well took your heads up with that part but i also didn't give it great images i would be curious i haven't tried to give it a reference video even if you couldn't you could extract the frames if you had like a drone shot a bunch of frames um yeah i mean this is a very preliminary test i am curious how far how much better it would get the more images you get or if there's like a cap of like it just doesn't i'm just curious about the first quality based on what you have from your cell phone looks like so right so i gave it to this i'm like okay go do the thing and then here's the thing oh wow that's almost there yeah it's pretty yeah so because i didn't have a great angle of the corner it kind of this is actually a second pass i gave it a revision and i said the corner should be more rounded and uh it's a corner door the first pass was a bit flatter but overall it got the lamps it got the overall shape of the it's okay yeah like the theater sign is the biggest egregious error i mean look it's giving you rounded corners some spheres it's building it with like basic shapes um look at the look at the floor can you go back to your image let's see the floor by the door yeah like that that's such a terrible picture for the the art deco yeah you know concrete stamping thing on the floor but But it did an interpretation of what it thinks it could do.
8:58And then, like, how many windows are there? One, two, three, four, maybe? Yeah, it didn't get all the door windows. It got a solid thing, some things out. But, man, for a first pass, geez. Yeah. This is good. And, look, it didn't really have a good image on this side, but it sort of made up these art deco. Like, there are movie posters on this side, and it's not really in the image, but it did a pretty good job of just sort of making something that made sense for the design. Look, for first pass, this was a refined... This is technically the second pass, and I said to change a few things, and it did that.
9:37But look, overall... This is such a great starting point. And look, you have all of your geometry separated out in the outliner. This is all ready to modify. I mean, yeah, I'm thoroughly impressed, Joey. Look, it got your top thing pretty well. I mean, it's not what it actually looks like. No, not even close. I didn't give it a grip, but it got the ball and it got some of the stuff here. This is not a great, you know, if I had a drone and gave it some better angles. So that's what I'm thinking. Like a 3D scan company and people that do digital joins of sets and stuff, they have the equipment to then, instead of feeding it into reality capture and making a mesh, now they can feed it into codecs.
10:22I'm thinking, too, like for me, it's like if a lot of the workflows I am messing with, like I don't need super accurate photo reel 3D models. I just need like gray box. Get me roughly there so that I have the geometry and we can do some camera moves and then just use these as base inputs into like video to video workflows. So like, I'm like, if it gets me 70 % of the way there of like what it should look like, and has roughly everything in place, we can clean all the other stuff up with like Gen AI passes. 100%. I can see that. Also, I'm seeing this for traditional VFX, where you have to send a team out to scan, put it all back together, right?
11:02Now that is just a bunch of phones, really, and maybe some ladders to take higher enough pictures. Oh yeah, or a small drone. I have the DJI Pocket Drone, which they built it so it's technically under the weight. And then what? It looks like it took you like six minutes in codecs to go through all of this. It's insane. Yeah, once it spun up the server that built this in maybe 10 minutes. And then you can also light and modify that. Oh yeah, also I didn't ask for it, but it gave me a video. Like a render. Render movement. Yeah. And then I said, oh, hey, can you relight it so it's like the lamps are illuminated and the building's illuminated and it's like a moody night scene.
11:44And then it did this. Obviously, you could tweak it and refine it, but I just said, hey, turn these into practical sources. Yeah, I would just take over in Blender because it's assigning emissive values to some of the trim, which is not accurate. But the light, like the little round light posts, that looks nice. I mean, look, here's what I'm trying to say is, you know, a year or two ago, I thought that the Gen AI revolution would really take us from a 3D world back to a 2D world a little bit. Because why bother with 3D when you can just get to final pixel, right? And this was my incorrect assumption back then.
12:22Since then, I've noticed and I learned the hard way is that you still need 3D because Gen AI just doesn't give you the control, right? If you want the camera move, if you want, you know, like that sphere to be held up by four things, like these very nitpicky things, then you got to do the thing in 3D. And then Jenny and I will hopefully fill that photorealism gap in 2D, which is what we're seeing a lot now. I think in a year or two from now, my prediction is we're going to be in an entirely different world that is so deeply in 3D, but without the complexity of traditional 3D. Like if you think about.
13:02Yeah, it's simplified. Yeah, it's just like all the things you don't have to do. You don't have to model. You don't have to do shading. You don't have to do lighting. You don't have to do geometry optimization and like rigging and skinning and all the things that is associated with a 3D complexity bubble. Right. And you're talking about people's job is to do very specific things, very specific lanes. And now you're going to have just an operator, just an AI 3D operator that can do all of these complex 3D functions in Blender, in Unreal, hopefully in Maya, and still have, quote unquote, a traditional workflow with the simplicity of using ChatGPT or Claude, just like prompting and typing.
13:49Yeah, especially as these things get faster and more real time. The thing that was also really impressive with this was it took control of the computer, but it felt much more responsive than when Claude takes control of my computer. And it's still like, let me look at the computer. Okay, let me assess the screen. Let me figure what I'm looking at. Okay, let me take an action. And it's still this very slow process where the codex felt a lot smoother. I mean, it was also talking to it via MCP under the hood. But as this keeps getting faster and faster and voice models and all that stuff gets better.
14:20and then you're just literally in a gray box room or just in whatever space you want and you're just talking to the thing and you're like, oh, hey, like, okay, I moved the camera here. Okay, let's add a tree over here. All right, let's, you know, I need to make that wall bigger. And you're literally directing and it's doing all of the 3D set build stuff under the hood and it's just responding to the direction. 100%. And the other two things that are kind of converging and will sort of meet in a really happy place, I think is the fact that GPUs are getting more powerful, right? NVIDIA is working on this.
14:52They have Blackwell now. They're going to have future GPUs. So on the 3D side, you're going to be able to run trillions of polys, right? Right now we're in the billions. In the future, hopefully trillions, like insane amount of polys, more than you can imagine. And then at the same time, the AI models and inference will get faster or more accurate. it. So your interactivity with AI, your responsiveness, your iteration, that's becoming more natural. And then the stuff you're making is becoming more complex and more lifelike. So the end result is like you being able to recreate reality or recreate a completely new world effortlessly, just effortlessly.
15:32Also, I was just, this gave it Blender, but you know, Blender is the orchestrator or the base level too. Like imagine connecting this stuff to all of the other models and things that we've talked about image to 3d text to 3d uh marble world and world labs and using all of those tools at its capacity to bring in or create even more photorealistic things into your scene and then orchestrating in your 3d scene and then it's all there and consistent and uh controllable but much more higher yeah i i mean it's so hard to predict the future but like right now the way we're gonna we're sort of establishing the workflow and i see this across the board um is you know stuff on linkedin you know stuff on instagram and stuff is you do a lot of fine creative work in 3d and then ai is essentially a renderer it like takes you to a style or a realism level that you can't really achieve easily with 3d you're saying that in the future like that will be the workflow like 3d will play the piloting part and then the finishing part will be ai yeah yeah yeah for sure uh other thing too i saw this and i just want to like flag it of just like what even just new formats or new things will like be created with these tools so alex barashkov built with codex uh what's calling uh aval or aval a new open source format for interactive video on the web and he has this demo of this sort of this interactive video with uh grass blowing i know this clip's a movie i just don't remember what um and but just you know just he built this whole brand new format for handling video uh and interactive video on the web with codex and so it just makes me think like what other new things formats are just going to be open and unlockable for him imagine that like just making your own USD or FBX file format.
17:28Because you can. You have the ability to have complex software tasks automated. Small file size, low CPU overhead, alpha transparency, web native runtime. The future is so hard to imagine because it's literally not built yet. We can't draw a straight line because the points on the beginning of the line aren't even there. I remember you poo-pooed when we talked about that thing a few months ago that was like, what was it? It was like generative websites. It was like websites that sort of built on demand. Yeah, I would poo-poo. Based on your need. But I feel like that is, I feel like that's still going to kind of come into play more and more.
18:12Like, if you have a use, you need software, you need something, these like tools or solutions that spin up based on need because it's just so easy to spin something up from scratch that does exactly what you need to do. Yeah, also it could take us into like the wild west of software and formats and things that are not debuggable, just built by Vibe Codo. It could be like a total mess, right? Like one pipeline will look nothing like another pipeline. So if a VFX artist worked on one show and they go to another show, they're like, well, what am I doing? Yeah. Yes. That's probably going to become very messy, especially if like every, yeah, if every studio builds their own proprietary system.
18:56I think, yeah, maybe the systems interfaces are unique for whoever builds them, but I think open standards are going to get even more important. Yes. Like USD, open timeline, whatever. Yeah. A lot of stuff that they're talking about this Sunday is going to just keep getting more and more vital. So like you could build your own system, but like making sure your system can communicate with everyone. Wow. You really just convinced me to go to the Sunday open source thing. Like, yeah, that sounds really relevant. I need to go. Yeah. That's what I'm thinking. I'm looking at the talks there. I'm like, yes, all of these things we're like, I'm messing around with because we need ways to like, if I'm building an app and I'm building a timeline, I need to make sure that timeline can get sent to Resolve or Premiere.
19:42Yeah, you're right. Open source and open formatting become more important than ever. but that's not going to stop somebody from coming up with their own USD. I'm just thinking like just going back to USD and how much of a revolution that it was and how long it took. Like Pixar took 10 years to make that from scratch. And that was, I don't know, the early 2000s is when the work started. And just now we're seeing all that trickle into all the nooks and crannies of media and entertainment. I mean, it's still not for the amateur because you need software to be able to handle it correctly. But I'm just thinking if there is such a thing as a standard for building your own standard.
20:27Yeah, I think it depends what you want to do and what you need to be able to transfer in and out. Yeah, I mean, I don't know, maybe I'm naive, but it's like, yes, those, like, USD and stuff took a long time to adapt because there were only a few big software players that would handle it. Yeah, so software was expensive and now software is super cheap, right? So that's the thing that flipped. Yeah, and it's like, oh, if you don't accommodate what I need, I can just build it myself. and that's the part where I'm like, maybe this is a naive take, but it's only going to get better and faster and easier.
21:07And there's also all types of decentralized resources that you can tap into. What if there's a format that just runs FAL stuff? What if there's a format that just runs NVIDIA's web GPU stuff, right? And it's not really built for anything local. It can run on a Mac, on a PC or anything. That hasn't been really imagined as far as I know, right? Like nothing is really 100 % optimized for 100 % decentralization. No, I mean, unless you're building it. Yeah, maybe like GLB formats, like stuff that, yeah. Anyway, I don't know enough voodoo signs behind those formats. But yeah, like it's an interesting area to watch for sure.
21:53So not only the workflow itself, but the delivery and the formatting and the exchange of information. Yeah, yeah. All right, you want to talk about older software now? New stuff in established software? All right, let's talk about Nuke. Nuke just announced a SmartRoto function along with some updates to grip tape. Okay, SmartRoto, new Nuke plugin that handles, basically, it's SmartRoto. I'm betting it's like AI-powered under the hood. I would assume so, too. I mean, this is stuff that's already been in Resolve with IntelliMass, IntelliTrack, and Premiere and After Effects has pretty good. Foundry is smart.
22:36They're so sensitive to the use of the word AI or any sort of artificial intelligence connotations, although they just acquired grip tape and they're going full AI. But in a way that makes sense to high-end professionals and that works in professional practice. Dude, nobody wants to do frame-by-frame roto. It's a delicate balance, but... Come on, man. Those roto artists don't like to do roto. No, no, I'm talking more about grip tape and stuff. Like, you know, you were saying they go fully AI, but like fully AI in a way that still integrates with high-powered professional pipelines. Yeah, they're doing it the right way for sure.
23:10Grip tape to me, at least the grip tape enterprise, feels like comfy for VFX for the lack of better turns. So yeah, I saw this announcement. It's like OCIO support, which, you know, Something you would have to probably build in Comfy. I'm not sure there's a node for it or anything. Local GPU inference and stuff, which I thought was super interesting because studios are super sensitive about using the cloud for unreleased IP and things like that. Yeah, they're definitely more in the keep it local. It's all good progress. It's really converging with traditionally AI nicely. I think the Foundry and Grip Tape understand, obviously, given their history with Nuke, They understand that world really, really well.
23:55That is their core audience, their core demographic. So I expect nothing less from the Foundry. And they have sort of fine-tuned themselves into this world. No pun intended. Yeah. I think smart Roto only gets more important with AI-generated stuff when you've got to isolate elements. Yeah, and I mentioned on the pod before, because of the inherent denoising of AI generations, Roto stuff and matte stuff doesn't work well with it as far as I've heard. I mean, from some of my testing, I can't 100 % confirm that. Oh, right. You mentioned that last time. No, I don't think that's true because we've been masking the crap out of generated stuff to clean it up.
24:40Okay. So maybe it's the... It behaves the same. I don't think it's Roto. Maybe it's tracking, camera tracking. because camera tracking is essentially trying to figure out a 2D space into a 3D camera movie. Okay. We haven't done as much of that. So, yeah, maybe. Okay. I'll try to try some shot. See if we can track it. All right. Rage bait time. We're just doing this for the rage bait? Okay. We're doing this to clip this. AI-generated Odyssey movie costing a few thousand dollars. unveiled ahead of Nolan's$250 million feature. I saw one of these things frame it as going head-to-head. So that's what I thought.
25:25I was like, wait. Who's going head-to-head? Nolan makes like three-hour movies, right? It's like, wait, somebody made a three-hour odyssey with AI? I have to see it. And if you look at the length of that video clip, one minute, 52 seconds. What is the trailer for the epic that is coming? Yeah, I think this is it. I don't think it's coming. I think they ran through their credits with this. Yeah, I mean, you know, I'm looking at the trailer right now. I don't have any sound on, but it's got shots. I mean, it has some beautiful shots here and there. Obviously, well-color-graded, synthetic humans, I'm guessing.
26:04But, like, yeah, some of the shots completely break, especially, yeah, anything with perspective or distance. AI still can't figure out real accurate lensing in glass, and your eyes can easily tell that that is artificial. So, for example, you can pause on any frame, and I'll point it out. That feels like the least of my issues with this. There's just more like the camera geography and placement just off the croat performances or the talking. and everything still feels just uncanny valley that has not been like that has not been broken from a fully synthetic yeah point of view cyclops is it with two eyes no oh creatures yeah oh that that had a nice fisheye or anamorphic lens to it again i mean i feel like i've seen five different styles in this one trailer you know what they should have done that because i don't well you're not on X as much, but when the first costume, like Matt Damon, appeared, images from the film, the history dweebs.
27:13That's not period accurate. You know, Greek, period accurate. And then the period accurate ones was like these ridiculous, like garbage shoot-looking things. So someone, if they wanted to go all in with the Odyssey, they should have made it period accurate. Yeah, I think it was like the bronze era or something. Of the Odyssey. The armor was like way more modern, something like that. Yeah, something like, I mean, the bottom line is the armor that is in the movie looks cool. And the armor that was accurate did not look cool compared to our standards today. So, you know, hope that helps when they go to Hades and then you're like, what's the accuracy of Hades?
27:49And it's like, hmm. Yeah, I think the Hollywood Reporter is really just struggling for clicks here with this headline. I'm always more shocked. Who the hell is their PR rep that gets it? It wasn't the Hollywood Reporter. Yeah, it's like the same Tilly Norwood type headline, you know? I was going to say, that came out a few weeks ago, and I've got a strict ban on it. We're not talking about that stupid shit. But, yeah, that story came out, too. That was like, she's going to be in her first movie. I was like, what the hell? And that was on everywhere. That even made it to the point that my completely non-AI wife was like, there's some AI actress that's going to be in a movie or something.
28:25She's like, aren't you going to talk about it on the podcast? Kudos to whoever the hell is behind that for getting this stupid PR press for this thing. That's been the most amazing feat to me. Yeah, absolutely. So, I mean, I will give credit to the guy, what's his name, Ash, the guy that made this. Ash Kusha. Even though it probably won't work as a cohesive storytelling feature film, the shots are cool by themselves. so hats off to just pure generation yeah yeah but we know like we know the shot we the eyes good at making cool shots like that's been proven it's like making a cohesive thing where you're like in the story it's got consistent characters consistent shots now we're talking about this let me bring up one thing that i did see that while you're bringing up a quick rant we talked about this time and time again i'll cover it one more time because we're on this topic the true unlock with ai is to combine it with physical performance and on location shooting if you can just solve for the expensive bits of production let's say you half your cost that is the win having said that there are people trying to figure out fully synthetic everything and i think they're really limiting themselves by doing that it's like i'm not gonna put a camera on anything i'm not gonna put any human performance in there why because ai can do it all i'll just do the whole thing synthetic i think you've set yourself back like so far back with quality if you do that yeah i think yeah the way you're saying the unlock is human performances i would argue i don't think you need you know locations um you know i think you can get away with a lot But it helps, you know, immerse it.
30:10Yeah, you don't. I don't think you need full locations. I do think you need a film, like a cinema camera with real glass and real film back because the perspective, the distortion, the sort of geometry of the real world, AI just can't get right. So then if you're just bringing that in as a reference plate and then have AI fill everything up, that'll go so far to make it look realistic. You definitely need props and objects. Anything they touch and interact with, you need that. You need people. But you can do matte painting in the back and have parallax and all that stuff. It's like LED volume.
Read the full transcript
30:52Yeah, I think that is the best balance. Look, I mean, sure, if you're super low budget and you want to try to make something that is just completely generated, at least the option is there. I just don't think the quality is there. to make something that feels completely immersive. That said, this was the teaser thing. ByteDance released a trailer teaser video that was made using CDance 2.5. And this is all from, this is like a football commercial. I'm in World Cup mode and it's all football now. And then when World Cup ends, it goes back to Insider. Yeah, that's pretty good. This is 2.5? I've been waiting for some 2.5 outputs.
31:37the motion that's starring a kid who's kind of like yeah kicking the ball through london uh movement looks good i mean the ball contact with the feet and yeah even like that shot there that should be so hard for ai to do and it does it well yeah this um you know so we might be eating our words we might be eating our words and that is okay this is the whole point of the podcast you know follow us for the journey i mean obviously the caveat with everything is like this is the worst it'll be today yeah still no word on when 2.5 is coming out i keep seeing updates and it's getting that keeps getting pushed back but um some uh it's coming i think i did misspoke speak last time i think i said 50 references oh gosh only 30 references um are you kidding me where am i even gonna find 30 references i have like three at most in most generations here's my photo here's my entire photo library just go go make a movie oh you know what the culver city theater could use 30 references yeah that could use you know yeah like all the close-ups and things like that yeah for sure but you have to see if they'll place all the references in the right location spatially yeah you have to give it a wide shot of stuff i am curious if you how much instruction you can give it with the references that it'll retain and understand.
32:58If you give it a wide shot and then close-ups of things, will it understand, like, okay, this is how everything relates, but the action should be here and things should happen here. Yeah, it'll be the new Wild West of how do you coordinate and direct this stuff when you have so many reference images and what can it understand and stick to. Okay, so when 2.5 comes out, you and I both have to generate some reference images. Okay. we might go broke from that because it might come out we don't know how much they're going to charge and if they're going to charge you more based on if you have more reference images oh yeah that yeah okay a 4k we should start a GoFundMe for that generation the 4k cdance output like one shot is roughly like five bucks oh my god really if you max it out to 4k even just I think five seconds it's like roughly like four or five dollars yeah it's steep you can't do iterations that well for$5 those are LA prices man is it cheaper to run if they keep those prices up it'll be like the other software development companies when they were like oh we're going to rehire the people because the tokens are too expensive then what it costs for a person I know it's happening it's going to be the same with film where it's like oh it's actually cheaper just to shoot this for real than to blast you dance with it the costs have got to come down I mean I wonder how much of it is the actual data center GPU cost and how much of it is just greed yeah I don't know I mean I think a mix of both because also depending where you generate for CDS especially the prices do vary a bit quite a bit so we'll do one generation I can afford like I don't know$20 I'll do four MARK MANDELAVYSKI - And that's all, Joey.
34:52All right? You just get four takes. MARK MANDELAVYSKI - When 2.5 comes out, we will do that. OK. Last thing, quick thing. In the open source world, Thinky Machines, which was MARK MANDELAVYSKI - Mira Mirati, who left one of the OpenAI founders, started Thinky Machines. They've been working for a while. They raised a bunch of money. They now have released their first official model called Inking. MARK MANDELAVYSKI - Yay for open source. Open source open weights. Yeah. It seems like a pretty solid model based on this article. It does significantly step up the competition with ChatGPT and Claude.
35:30Yeah. I mean, it is definitely not at that level. But in the world of open source, it's on par with Kimi 2.6. I mean, definitely this is a huge leap for an open source American model. Because pretty much all of the open source decent models have been from Chinese. Yeah, they have been from Chinese model, AI companies. So yeah, I mean, this is a big leap for putting out this big model, open sourcing it, multimodal. I'd be curious once someone, or I think it might be, I think they have a spot on their website called Tanger that you can mess around with it. And then obviously you could run this open source, or I'm sure a bunch of providers will add the ability to work with this model via third party tools.
36:17Yeah. So going back to Mira Morati, she was the CTO of OpenAI. She left. And then Ilya Satziver was, I think, the chief research officer at OpenAI. He left. And then those are the two people I'm watching closely. And then the third person I'm watching closely is Yanlakun, who left Meta and now has a billion dollar in funding for his own company. He thought LLMs were like not the way, right? He still thinks that. I think he said something like LLMs have two years left. Because I think he's right. Scaling is not working. It's giving you diminishing returns. And no matter how many more billions of parameters you train on, intelligence doesn't go up by that much.
37:02And what's the alternative? A bigger world model? It's the JEPA architecture, according to him. We could spend another episode on the JEPA architecture. but essentially you move away from a text-based approach into a more abstract approach where text images video any kind of just become seamless with one another rather than us discreetly building a text model image model video model mashing those three things together to get reasoning this is inherently more fluid in nature okay all right i'm very curious about that Me too. Okay, so this is out. And then also, I know I just compared it to Kimi, but Kimi is also teased with a pretty sick launch video that Kimi 3 is about to come out, which is expected to be China's largest AI model with two to three trillion parameters.
37:56Jeez. Open weight on par with like quad opus 4.8. Oh my God. Did you say trillion? Yeah. So, yeah, that's coming soon. How big is the file going to be? Like the weights? Yeah. Like one terabyte? So, you know, that's coming. That's awesome. I mean, like, it's not awesome in the sense that the biggest, baddest model ever is coming out of China and not here. Biggest, baddest open source model. It's crazy. Yeah, I wonder why Thinking Machines launched their thing open source. Do you think it's strategically to just get people to use it and twinker with it and get their heads on it? I think so. It's like, what do you do when you're starting a new thing and you're already, like, you're not going to beat ChatGPT or Anthropoc.
38:52So, like, what do you do to make a splash? And if you've been trained in a pretty good model and you want to get people on board with it, open source it. You start to get users on it, build momentum on the flywheel, if you will. I mean, it's what Stable Diffusion did so well with SD 1.5 SDXL, right? They gave it away for free very early on, and now it became synonymous with image models a few years later. Yeah. I mean, they haven't released it yet, but one of the first things that Thinking Machines teased a few months ago was that conversation chatbot that was much more, And the demo they did, it was like a much more responsive, reactive, like fluid conversation that they were having, like where if you interrupted the model, they kind of like picked up with the new chat.
39:38And it was something that it seemed like they were focusing on was like a conversational aspect. And that part hasn't been released yet of whatever they're working on there. I did see a really interesting feature that I didn't see with other models is it can fine tune itself. Oh, interesting. Based on, yeah. So like if you need them, let's say you want to run an agent that runs a dentist office or something. Like it has to fine tune on dentist stuff like, you know, orthodontist and things like that. It could probably fine tune itself if you just ask it to and give it enough data. That's cool. Because right now it's sort of like a hack where you have the model and you have a bunch of markdown files.
40:21And then the markdown files update. Right. And then you've got to build a giant context window. And I think the context window here is pretty insane, too. I think one million characters or one million tokens, which is a massive context window. I think I saw on Tinker, at least, it was like capped at 256. But maybe if you go full blast, it's a million. Yeah. So, like, it's definitely a modern model by today's standards, which also happens to be open source. And, yeah, I'm hopeful that this is a good alternative to ChatGPT and Claude. I think competition in this space is what we need, you know, just to get things better and more efficient.
40:59I'm just really worried about using all those GPUs. Like, can we get a model? Yeah, and stuff that could run locally, too. I mean, I think that as workflows get more established and as it sort of pushes into, like, higher-end film production workflows or other things, like exploring stuff that you could run locally is going to get more interesting. Absolutely. And more needed, especially when you start comparing the costs of API tokens versus spinning it up yourself. Yep. Absolutely. All right. Good place to wrap it up. This went deeper than I thought, but this was a fun one. Links for everything you're talking about at denoycepodcast.com.
41:40If you want to catch Joey and I at SIGGRAPH, we'll be there Sunday at the open source thing and Monday morning as well. And I think, Joey, you're going to attend. I'll be on a few days longer. Yeah, you're going to attend that. Yeah, definitely be there Monday, Tuesday. I don't know. The thing goes until like Thursday. That's why I was also curious, like how it's like longer than NAB. Yeah. It's a long thing. A lot of it is like the big keynotes and stuff, which go on. Yeah. All right. But Lenny Smoot. Lenny Smoot. Thanks everyone for watching. We will catch you in the next episode.
From the publisher
Joey tests OpenAI Codex + Blender MCP to reconstruct a real building in 3D from phone photos — and the results have real implications for production workflows. Plus: what to expect at SIGGRAPH and a free AI Workflows Summit in LA.Also: Nuke's SmartRoto...




