In short
State of image/video/audio generation and “visual intelligence,” comparing diffusion vs autoregression, explaining Black Forest Labs’ flow-matching approach, and how context-aware editing enables practical applications beyond creativity.
Guest backgrounds
Dustin Podell, co-founder and researcher at Black Forest Labs; works on image generation, editing models, and hardware/performance optimizations.
Key claims
Core generation process remains “add noise then remove noise,” but Black Forest Labs improves it via flow matching (learning a “flow map” in latent space). Adding context (image references) is a major step change, enabling relationship-aware editing. Visual intelligence models can simulate world relationships, forming a foundation for robotics and real-world action. State of the art for text-to-video cited: “Sora”/“Sora 2” mentioned; also “C2D”/“C dance” cited as producing ~4K, up to ~15s cinematic generations.
Notable examples
Editing a product photo into product photography sets; multi-reference clothing try-on; home furniture visualization; hackathon scenario generating crowd movement through fire exits; editing examples like “knock over a water glass” and “put out a fire.”
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOState of Image Generation Methods
1:24 to 2:29
Discussion on the current state of image generation technologies and workflows.
“But speaking of images, videos, and more specifically, image generation, really excited to have with us today, Dustin Podell, who is co-founder and researcher at Black Forest Labs.”
Evolution of Generative Models
2:29 to 4:25
Exploration of how generative models have evolved in the last few years, including advancements from basic blobs to realistic images.
“I mean, if I'm allowed to take it even a little bit further back, I mean, the state of...”
Technical Overview of Diffusion Models
4:25 to 6:00
Dustin explains the technical workings of diffusion models and how they generate images.
“And as Daniel mentioned, we can offer some links for past conversations if people want to dive into those specifically.”
State of the Art in Generative AI
6:00 to 8:02
Discussion on the current state of the art in generative AI, including leading models and their capabilities.
“I'm trying to think how to essentially take this story from here.”
Practical Applications and Comparisons
8:02 to 14:02
Comparative overview of various AI models and their practical applications in generative tasks.
“And then what we do is instead of trying to predict, so an auto aggressive, you might try to predict like pixel by pixel by pixel.”
AI Community Leaderboards and Practicality
14:02 to 15:48
Discussing the limitations of AI model rankings and the importance of practical AI events.
“But I feel like this kind of removes some of the specificity, the granularity.”
Diffusion Models vs Flow Matching
16:05 to 17:34
Exploring the concepts of diffusion models and flow matching in AI image generation.
“I was going to ask you, are the models that you're describing, are they kind of the traditional diffusion models?”
Understanding Flow Maps in AI
17:34 to 20:44
Describing the analogy of flow maps and their significance in AI image training.
“what this essentially looks like is if you can envision like a, I almost want to, I don't have a whiteboard behind it.”
Optimizing AI Training Efficiency
20:44 to 21:45
Discussing optimizations in training AI models and their implications.
“And that's what like as you get closer, it starts to look more and more like a real image.”
The Evolution of AI Image Generation
21:45 to 24:46
Tracing the development of AI models and their applications in creativity and beyond.
“I get an image, I can immediately remix it with an image model, right?”
Show all 18 chapters
Practical Applications of AI in the Real World
24:46 to 27:15
Exploring how AI models can be utilized for practical purposes beyond creativity.
“So if I say, take a picture of a, like a, I don't know, we have like a water glass here on the table.”
World Models and Robotics
27:15 to 28:00
Examining the concept of world models and their relevance in robotics and embodied intelligence.
“And on one end, you can take that and go, okay, well, this is nice for creativity because you can make a film scene and it looks nice.”
Exploring World Models in Robotics and Visual Intelligence
28:00 to 31:00
Learn about the connection between world models and robotics, and how they are understood by different experts.
“Now, this isn't to say we're going to drop the creativity side of things.”
Practical Applications of Image Generation Models
31:00 to 34:40
Discover various practical uses of image generation models in everyday scenarios, including e-commerce and emergency planning.
“Yeah, from the non-robotics person in the room, that was really interesting.”
Overview of Black Forest Labs’ Model Families
36:28 to 42:00
Get an in-depth look at the various models developed by Black Forest Labs and their functionalities.
“These are all, so you mentioned the Flux family of models.”
Progress in Visual Intelligence
42:00 to 43:40
Learn about the advancements in visual intelligence and image generation technologies.
“But just maybe give us a little bit of sense, because that has changed on the LLM side, right, where that has continually updated.”
Future Aspirations in Visual Intelligence
43:40 to 45:20
Explore the speaker's aspirations for the future of visual intelligence and its capabilities.
“And this is something we've continually done, you know, with our Fluxon series, we had Schnell.”
Real-Time Interactions and Robotics
45:20 to 47:46
Discuss the potential for real-time interactions in AI and their applications in robotics.
“There's kind of two areas that constantly sit on the front of my mind.”
Transcript
Automatic transcript. May contain errors.0:01Welcome to the Practical AI Podcast, where we break down the real-world applications of artificial intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place. Be sure to connect with us on LinkedIn, X, or Blue Sky to stay up to date with episode drops, behind-the-scenes content, and AI insights. You can learn more at practicalai.fm. Now, on to the show.
0:41Daniel Whitenack:Welcome to another episode of the Practical AI podcast. This is Daniel Whitenack. I am the CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a principal AI and autonomy research engineer. How are you doing, Chris? Hey, I'm doing great. Can't wait to get into today's conversation. It's going to be fun. Yes, yes. For an audio podcast, we're going to talk about a lot of interesting visual things. Maybe before we get started, just a little teaser. Practical AI is posting some videos on YouTube now. So if you do consume podcasts that way, you might go check us out on our YouTube page.
1:24Daniel Whitenack:But speaking of images, videos, and more specifically, image generation, really excited to have with us today, Dustin Podell, who is co-founder and researcher at Black Forest Labs. Welcome, Dustin. Yeah, yeah. Thanks for having me, guys. It's really great to be here. Yeah. And I know that Black Forest Labs does more than just kind of raw image generation. There's a lot of workflow related things, hardware optimization, all sorts of cool stuff you're involved with. But as we get into some of that, I'm wondering if you can just help our audience with a bit of a state of image generation methods and workflows for the industry.
2:06Daniel Whitenack:We've talked on the show before about diffusion models, and we'll link some of those episodes in the show notes maybe. But a lot has happened, right? There's a lot of people working on a lot of interesting things. And I'd love to kind of understand, like, over the last year, what are some of those main points that might be good for people to orient themselves to where things are at now? Yeah, yeah. No, it's a good question. I mean, if I'm allowed to take it even a little bit further back, I mean, the state of... Absolutely. Yeah, the state of image gen, video gen, generative models as a whole has kind of gone crazy, so to speak, in the last three or four years.
2:50So where are we now? um where we came from about four years ago uh where where i first kind of entered into into the scene so to speak um is we were at models that were essentially just doing little blobs of color that were kind of related a little bit to where you were with the prompt you know okay oh a lighthouse on the beach or this and that and you would get something that okay that vaguely looks like it and you would show someone and it'd be um yeah i could see it i guess and and maybe it would interest a few nerdy people but and then now today we're at the point where you know if you've probably been seeing plenty of this stuff online where we're seeing whole short films made entirely with a ai generation where certain scenes are almost entirely indistinguishable from uh from reality so to speak so i would say we we've uh we've come quite far but i will also say the core of the technology hasn't actually changed that much it's been a a pretty nice like forward progress.
3:48I don't want to like dive immediately into anything technical here, but, but it's, it's certainly, I mean, for anyone who's been paying attention, I'm sure, or anyone who really hasn't been paying attention, this probably came a bit out of, out of nowhere. So, yeah. Yeah. I was going to say, I guess, you know, with the broadly attention by the general public is so much on kind of more of the LLM generative world in terms of, you know, that, that stuff and everyone's finally on apps, regardless of whether they're technical or not. And so I think a lot of people kind of miss the tremendous advancements you guys are making on that side.
4:23And so like, you know, could you take a quick moment and now that you've kind of done the highest level, maybe step through some of the things that people may remember, like we talked about stable diffusion and things, a couple of those things that have kind of led up to what we're going to dive into today with some more specifics. And as Daniel mentioned, we can offer some links for past conversations if people want to dive into those specifically. But that would really be interesting to kind of hear distinct points on that timeline as we hit to this point. Sure, sure. First, before I do that, I guess I'll ask Chris, how technical am I allowed to get here?
4:59You're allowed. We're going to say we'll like if you dive into something, we may stop you and say So that is what, like, what is that acronym? But other than that, you can go as technical as you want, and we'll make sure that everybody can stick with us. Okay. Including me. Yeah, yeah. I guess if I'm allowed, I'll do kind of at least my, the best of my knowledge overview of the last bit of time since. So obviously the field of AI and ML has been around for a while, trying to structure and learn distributions of different data. Mostly it was used for, you know, tabular prediction, all sorts of other things for a while.
5:36And then in 2017, some researchers at Google figured out this very cool thing called the Transformer, which I'm sure people have heard many, many times about right now, which I guess the best way to go down this path is to say like we found a way to train much more general models, like learn from much more general data. I'm trying to think how to essentially take this story from here. It's a dramatic pause. It works, man. It's good. Yeah, essentially the story I want to tell is the story between the kind of, I don't want to call it the battle that's been going on the last few years, but essentially the story between diffusion and autoregression that's been going on the last few years.
6:17So let me start by trying to define a little bit what diffusion autoregression is. Autoregression is anyone who's been playing with language models, oh, you know, all these, you know, GPT, CLOD, all the new stuff, everything that's been going on in the last few years. These are what you would call an autoregressive language model. And so what this is, is it's predicting data essentially one piece at a time. So it predicts one piece of data, and then it looks, okay, what did it make? Then it predicts the next piece of data. It looks at what it makes, then it predicts the next piece, and it slowly, slowly, slowly, slowly builds out.
6:52You know, at first it was just very basic language, and then it became, you know, now, I mean, we're having agents and a lot of very, very fancy things with them.
7:03On the other side of the coin has been something called diffusion modeling. and i mean i i don't want to get into it because now now we're doing more like flow matching but let's just go with the fusion as the as the base term here and with the fusion what you're doing is you're not trying to model one say like in language model it's you know you call it a token but you just assume like okay one word or at a time with the fusion what you're trying to do is essentially trying to think how to tell how to tell how to tell the story here yeah yeah no it's all good i i love i feel like i'm watching like a thriller where you're about to like reveal uh the next thing it's good keep yeah so so what we do with the fusion is i'm just going to tell practically what we do and then we can kind of break down you can ask me questions that's fine um so what we do with the fusion is we take a whole continuous medium and what i mean by continuous medium is the most famous one is images the world is continuous there's not you know it's not that there's you know one you know like a letter is discrete there's a t there's an a a b a c a d and e and an f a color shape the structure of the world the color you know the the light the all of this this is this these are continuous natural mediums and so um what we do is we take one of these continuous natural mediums that that we've now encoded into digital space so we use rgb so you know okay 256 colors we typically use uh for red green and blue and then we encode this into an image.
8:31And then what we do is instead of trying to predict, so an auto aggressive, you might try to predict like pixel by pixel by pixel. Instead, what we do is we essentially try to remove information. And the way we remove that information is by adding noise. So you imagine you like add a little bit of grain on top of an image, you can still kind of see the image. Maybe it's a little grainy. Maybe it's a little blurry. Oh, you can't quite see it. And you add a little bit more. Maybe you see some of the shapes. Maybe, okay, there's like, is that a dog? Yeah, you kind of see it. The color starts to fade.
8:59And then you keep adding. And eventually you just get, okay, this is just noise. I don't see anything in this image at all. And with the diffusion model, what we're trying to do is we're essentially adding a bit of noise. And then we're saying, now try to predict what the clean image is essentially for this. So, you know, okay. So we add a little bit of noise, predict the clean. It just has to remove a little bit of noise. You add a lot of noise. And now what it has to do is there's so little information, it has to actually like come up with essentially like how to infill this properly. And then what you do is then in what we call inference, which is when you actually run the model, you ask for an image.
9:35What you would do is then you start a fully noisy image and you say, give me a dog wearing a top hat on the beach. And then essentially what it does is it tries to remove a little bit of that noise to get closer to this idea of a dog and a top hat on the beach. And you might get some structure. It tries to figure it out. It's very coarse. It's very just trying to figure out the shape. And then it has a little bit of shape to latch onto and then it does the next step where it removes a little bit more of that oh and it starts to come into view it starts to you start to get a little bit more in focus and you do this and what you see is essentially you completely remove the noise of this image and oh my goodness now you have this this completely generated image of a dog on the beach with a top hat for a long time this was uh i don't want to say not good i mean all the like all the models have gotten a lot better but when it first started out these look like blobs of color it was like these little like you could barely see it.
10:24But fundamentally, what I just described is still the process that's been happening over the last four years from these old image generator models for anyone that maybe tried like Dolly Mini, I mean, the old stable diffusion models that we work on, you know, I don't know, many of the different models are out there now to these modern video models that are producing whole 10 second sequences. Fundamentally, it's doing the exact same process of removing this information slowly with this noise and then just slowly going the other direction, essentially creating info to infill this, if this makes sense.
10:57I'll pause there for a second just to kind of like that. So let me tell you, because as I'm listening, very, very good explanation, probably the best one that I've heard. So you are right on target if you just carry on what you're doing, because I'm enjoying this.
11:13Daniel Whitenack:And you mentioned that when this started, right, you sort of ended with these blobs, et cetera. I think there's a lot of people that have seen very impressive things like the, like you were saying, like whole parts of maybe movies or commercials or something being generated in this way. What is the current, I guess, state of the art in terms of whether it be your models or others? And I know there's more to talk through as you get in. Like there's more than just raw image generation. There's how this fits into workflows and doing certain tasks. But what is kind of the state of the art in terms of what we're able to generate, where are the bounds currently and in terms of quality and that sort of thing?
11:56Sure. So I'll tell you kind of like in the last year, I'll split it up into, let's call it three different categories, but they kind of blend a little bit, which is these continuous mediums I talked about. So I gave the example of a continuous medium with image. Video is obviously an extension of that, adding the time dimension. But then there's also audio as well, which now a lot of these video models generate audio. but then there's also things like, I don't know, maybe you've heard the, the like Suno or these music models. Some of these are also using this exact same technique of removing the information with noise, bringing it back in.
12:27It's just a different medium that they're essentially applying to. Um, I'm happy to talk about our models all day. I love talking about our models. Um, but I also want to be fair to just anyone listening and honest to kind of the state of the world on if you want to um you know go out and and try some some uh some nice things out there as well um uh i i will say that before i kind of like dive a little bit into into what i think is you know maybe the best one to recommend i i will make like one statement of of that like best quote unquote here is a little bit hard to define sometimes and this is something we're always like wrestling with is like what makes a best model for someone because you know obviously you when someone uses these models, they're typically prompting or putting in now you can reference images or videos or other sorts of things and kind of treat it a little bit like more like a creative companion.
13:18And then like what you expect out of this can be very different for different types of people. That said, I would say that if you want to see kind of the quote unquote state of the art right now for what I would say is like text to video at the moment, um it would be c dance so c dance has done some extremely impressive work with the c dance 2 model over the over the last year so they're doing uh i think it's now 4k generations up to 15 seconds um which is i mean a very nice quality it's it's focused very much on like cinematics um i will throw a shout out to anyone that uh you know if anyone from the sora team ever listens to this i think sora 2 was a quite a nice model it didn't rank so high on um some of these like leaderboards that we in the ai community there's a i'm sure you guys talk about this plenty like leaderboards rankings where models place um i think in the llm side of things this is much better covered where there's very clear metrics of how well does it program how well does it do math on the more creative side of things with video models image models audio models we typically end up falling down to this single preference uh type benchmark which is just like a general preference everyone in the world goes and can can vote on you know on these leaderboards it's a a couple of different nice companies and sites that do this.
14:34But I feel like this kind of removes some of the specificity, the granularity. That's the word I was looking for.
14:43Daniel Whitenack:Listen, I've been to an incredible amount of AI events, many of which are good, but many of which are not practical. You know I love practicality. We're on practical AI. And some are just hype-focused. Some are sales-focused. that's why I'm always eager to share about an event that I truly think is practical and useful for people. That's what I discovered last year at the Midwest AI Summit. And they're going to have another Midwest AI Summit October 15th in Indianapolis this year, 2026. One of the reasons why I love this event was there was an actual AI engineering lounge where you could sit down and talk through your use cases with actual experienced AI engineers and practitioners to really brainstorm and come back from the event with actual solutions and practicality rather than just a bunch of content and slides.
15:37Daniel Whitenack:But there were also amazing keynote speakers, speakers from even that had been on the podcast before, like Rajiv Shah was at last year's event. I would really recommend that you go to midwestaisummit.com and for our listeners, you can actually get 20 % off with the code practical AI 20. So go to midwestaisummit.com. Don't miss your chance to attend this event and get 20 % off with the code practical AI 20. I was going to ask you, are the models that you're describing, are they kind of the traditional diffusion models? Or you had mentioned in passing a few minutes ago about flow matching. And so it's just in my head, I'm trying to kind of categorize how do the ones that you're talking about right now, how do they fit in and, and like, and how, what is it to go back and pull that up as you were talking about diffusion and you made the reference to flow matching.
16:38How do those, what is that, what is that transition and how do those models that you're talking about now fit into that? I'll, I'll, I'll first keep it very simple and say all of these models are doing the same process of what I described earlier of adding noise and training the model to then remove this noise. The difference kind of, this is where it gets a bit more technical in that, in how we actually approach this. So we used to do more of this process called diffusion, which I don't, I don't, I don't necessarily know if I want to get into how, how deep and technical this is, but essentially what we've done is we've cleaned this process up to what we call flow matching.
17:17And flow matching is essentially this very simple process of still doing the same thing, still training the model to remove this noise. But fundamentally, what it's learning under the surface is anyone out there who has ever seen like a velocity map or like a flow map, what this essentially looks like is if you can envision like a, I almost want to, I don't have a whiteboard behind it. I don't have a marker, sorry. So you're going to have to. For the audio folks only. They're not going to see it. That was great. It's probably better I don't have the marker. Yeah, it was great. Dustin, if you're on video, you saw it, but Dustin was literally about to turn back to the whiteboard behind him, which I love.
17:54Like if I was in the room, I'd be like, go, man, take me there. Unfortunately, a good bit of the audience won't be able to see us. You're going to have to describe it. Yeah, yeah, no worries. Let's keep it purely in the audio space. So I'll draw a picture with my words as best as I can. I spent enough time prompting anyway, so hopefully I can do this. All good. But yeah, essentially the way I like to think about it is if you imagine like a landscape, like, I don't know, you can imagine your town or the region you're in and imagine you're looking at it from like the sky and you see like your house, you know, maybe right in the center of this map.
18:33and what you might want to do now imagine okay there's wind going all over all over the area and and the thing that you want to do is you want to be able to throw up like a paper airplane from anywhere in this in this landscape your city and you want the wind to carry it so it lands on your house and functionally what we're trying to train here is essentially that in a much much grander hyperdimensional space. Where instead of it being your house and instead of it being a wind and a paper airplane, what it is is your house in the scenario is what we would call the manifold of real images. It's essentially the place in, I'm trying not to use too many, stop me if I use too many words here.
19:16I want to say latent space. I'll ask you, but you're doing fine. Keep going. We're fine on technical. We'll just get explained it as we go. Yeah, it's the place in latent space that would be where the actual real images are. So another way I'll try to describe in a simple terms, okay, with your town here, you're in 2D space here, where you are above an XY and the paper airplane has to fly in this XY coordinate. So you land and then you have the XY. In this space, instead of it just being two coordinates, it's enough to have all of the colors of the entire image. So that's why I say it's hyperdimensional.
19:50It's the center is where literally real images are. And so what we're doing is instead of it being wind and the rest of your town, imagine the rest of your town here is every image that isn't real. And what that means is noise. If you think about noise as a real image, you can go generate a bunch of noise and save it to your computer. But that noise would exist somewhere in this space of all possible images. They're not real images, they're noise, but you can save them. They're RGB values, you have them. And so what we're doing with this, essentially this flow matching is we're training the model, Like when we say we're training it to remove noise, what's really happening under the surface is we're training this flow map so that we can land anywhere in this field of noise.
20:35And then these flows, these winds in our scenario, when you take a step, will take you closer to your house or to the manifold of real images. And that's what like as you get closer, it starts to look more and more like a real image. It's blurry. It has this. And then when you actually finally land on this manifold, boom, you have hopefully a real image, but you'll have some image, you know, and then the better the model is trained, the better you have mapped this flow to actually take you to where you want to end up.
21:03Daniel Whitenack:And I'm assuming, yeah, no, and I'm assuming because you are kind of mapping this from where you start to where you end up with this real image, maybe in a way that's just to be crude, I guess, less random. I mean, at the end of the day, all AI and machine learning, Like it's sort of that training process is very much trial and error, but we have a lot of optimizations around it. Right. So I'm I'm assuming that this then allows you maybe to, I guess, advantage wise, shorten the the training or make that more efficient. Or am I am I misconstruing that in some way? Yeah, yeah. I mean, definitely like one of the things that's taken us a lot further over the last three to four years is just figuring out a lot of optimizations, both in the actual training itself and then also in just, you know, architectures of the actual models have improved.
22:01um but i will say the like fundamental underlying process of just removing this information and finding a path back to it is still the same process it's it's just like the car you know you've been having cars go down the road since the 1950s and we've made better engines and better safety and all this but you're still driving down the road like that that's yeah yeah that that
22:21Daniel Whitenack:makes sense and now definitely i think that's a good setting in the sense that we've got to a point where, you know, the models coming out of Black Forest Labs, the models coming out from other places, the models that are even like now in my text messaging, right? I get an image, I can immediately remix it with an image model, right? In some way. So these things are becoming more embedded in our lives. But I don't know if everyone in the audience, some might have been along for this ride where it was like, oh, cool, I can generate an image of a astronaut riding a horse on, you know, you know, wherever.
23:00Daniel Whitenack:And that doesn't seem that practical to folks. So I'm wondering if you could now kind of, given that that foundation that we have, and we know sort of where we're oriented in the state of, I love how you put it on your website, visual intelligence, which I think is, I love that statement, because it gets to more like language, although it's being used a lot in terms of agents and intelligence, like you mentioned, it is very much a subset of the information that we process as humans, right? There's this visual element, there's the audio, et cetera. So now that we have that foundation, could you help the audience understand some of, I guess, the practicalities and the outworkings of the like, okay, we can do this now.
23:47Daniel Whitenack:So what? So how does that help people in the real world in ways other than maybe just pure creativity? There's certainly like the cinema, like you were saying, that side of things. Not everyone's going to be generating movies, maybe, or maybe more people will, I guess. But yeah, I think you're understanding what I'm saying. Like, where is this going to impact? Where is it impacting me now? Where is it practically going in terms of the application? Yeah, yeah. I'm very happy to talk about this. This is a, I think we're going through like a very nice transition period right now where we finally get to leverage these for some very useful things.
24:24And if I'm allowed to take a little bit of a tangent and kind of lead up to it again. Wherever you want to go is good. We're all good. Yeah. Yeah. So, I mean, I guess I'll take it back to like a little bit of a timeline of things. So, okay. So, so early on we had these nice models where most people recognize them for you put in a prompt, you get an image out, you put in a prompt, you get a video, maybe you get a song. it's just this one way okay you make something this appeals to maybe creators or ad agencies or movie you know now cinematographers and and i mean we love this stuff we love the creative side of this and and what it allows because it fundamentally to me this is like a nice potential communication tool um but then i i feel like something uh changed i don't want to say changed but uh there was definitely a a a timeline moment when we started to move into um editing so uh was it two years ago now a year ago i don't know around a year and give or take uh we released our first um in context editing model called flux context and uh on the surface you would look at this and go okay well this is this is an image editing model it's something you could take a photo you can clean it up you can add a hat you can do silly things you know you can uh whatever you want to do with it it's a general editing model that's supposed to supposed would do all these interrelational things.
25:42And on the surface, that's very cool. It seems like another creative thing. But if you think about like what's actually going on under the surface, for the model to be able to do this, it has to understand a significant amount of relationships in the world and what it means for these, like how these relationships actually interact together. So if I say, take a picture of a, like a, I don't know, we have like a water glass here on the table. And I take a picture of this water glass and I say to the model, knock the water glass over and show me what happens. The model has to understand some part of the actual world.
26:14Like it has to actually model the world in some way to know, okay, it spills over, maybe something gets wet, maybe XYZ happens. And essentially what we're trying to do is train a model that could do all of these types of relationships. So the model is learning not just like this one thing, but just how the world works fundamentally so you can do this editing. Now we can take this a step further and we can look at all the video models that are coming out right now and you can do a very similar thing. You can take, you know, in a knit image and say, you know, okay, this person now grabs a fire extinguisher and puts out a fire, and it has to actually understand these relationships to do that.
26:52So, well, maybe this wasn't our original goal way back in the day. You know, we were trying to make very cool models and figure out how to, I mean, I think in some sense we wanted to model the world, but maybe we weren't thinking this far ahead. Inherently through this whole process, we've learned to build these models that are developing this and i'm trying really trying to avoid the the use of the term world model here because i'm gonna go there if you don't i'm just telling you yeah you might notice me trying to like skirt this because i think it's a little overused these days um but we definitely talk about it um but fundamentally many people which is where i was going so yeah yeah fundamentally this this is what what these models are doing when you train them at scale and you and you really train them to be general and robust is they need to be able to simulate parts of the world to get an output.
Read the full transcript
27:39And on one end, you can take that and go, okay, well, this is nice for creativity because you can make a film scene and it looks nice. But on the other side, this is why we're starting to move in this area, you know, this area where we're calling visual intelligence. And now we're starting to see how we can actually leverage what we're calling, not just us calling it, this is a general field term, but the representation inside of this model of the world to actually go do practical things. Now, this isn't to say we're going to drop the creativity side of things. We're still definitely pushing on the side.
28:07This is, you know, something that we're all still very fond of. But looking forward, like if our models are understanding these relationships and the physics of the world like this, well, this is a great place to put it in, say, robotics and actually, OK, now a robot has this understanding, this model of the world inside of it to go act and take actions with confidence in the world. So I'll leave it there as kind of trying to build up to it. But let me ask it. I'm just going to go there because that's kind of you're in the area that I spend all my time, which is, you know, embodied intelligence, robotics, you know, UXVs and stuff like that.
28:46That's my world. And so even within our space, the notion of world models, there's a lot of interpretation. If you get into a meeting with 20 people, there's 20 different definitions. And we have to start off by sorting all that out, you know, in terms of how we're communicating. As you add in this notion there and of this kind of contextual understanding, you know, that it has a representation, I am curious before we move on, how do you, and you've mentioned now robotics, do you see it as the same thing or is it kind of yet another variation of a world model, you know, using the word and stuff?
29:24Like how closely would you believe those two to be related, you know, as you're talking about that? Just because both are big topics, you know. Sorry, can I just ask for clarification? You asked them like how close do I believe like our models are to like a world model? Well, like when you say world model, I'm just kind of trying to clarify the same thing I did when there's 20 people in the room and you're asking what they mean by it. And as you mentioned, robotics and stuff, we all need this representation of the world out there so that we can do better about acknowledging the context of what we're working on, whether it's robotics or in visual intelligence, presumably.
30:02I'm just curious, in your mind, do you think that they are very closely related or are they kind of distinct ideas of what a world model is? What is your take on that? I would say that they're pretty related. I mean, to me, it's like fundamentally under the surface, what we're doing is we're trying to teach the model as much as we can about the world that we exist in and then asking for it to utilize that in some way. And up to this point, it's basically been mostly through creative mediums. But like that representation that we're talking about that, you know, and again, I'm trying to avoid this term because we, I don't know, it's a little overused, the term world model, but it is.
30:38I mean, this is what it is. It is modeling, it's creating a representation of the world we live in, and then it is using essentially that intelligence it has to be able to act in it. So the idea is that if it understands enough of these relationships and the actual physical properties of it, it's a really good foundation to build and train robotics on top of. Gotcha. I hope that answers that. That did. That was good. Thank you for putting up with that. I was just curious. Not at all. It's not unusual to navigate that. So go ahead, Daniel. Sorry about that.
31:11Daniel Whitenack:Yeah, from the non-robotics person in the room, that was really interesting. I appreciate you going into that. And I'm wondering, there's that kind of outworking of some of this where you are now kind of understanding the context that's in this model, maybe how it's representing the world. There's also things that I've seen just crossing my paths, whether it be kind of in e-commerce or like I say, my phone, my text messaging. So on the web where these models are being more and more integrated into workflows, whether that be kind of like try on these clothes or glasses or whatever, or it's creative tools in like visual editing platforms.
31:56Daniel Whitenack:Are you seeing that with your all's models in terms of kind of what's the state of, and I know I want to get into your model families here in a second, and they do different things, right? And many of them, I know there's a big, some that are open weight, so you might not know all the ways that they're being used, right? But from at least those partners that you're working with, in terms of today, what are some of those creative and maybe most practical uses of these models that you see out there beyond just kind of the fun image generation kind of side of things? Yeah, yeah, no, absolutely. I would say, again, I'll come back to the moment we started to get context into the model that wasn't just text made a huge, like it was a fundamental change in what these models could do and how people actually worked with them.
32:50Now, our first model, Flux Context, only could take one image reference. It was mostly used as an editing model, but you could reference a picture of a product and then tell it to generate a nice product photography set. And this was very nice. Then moving on, we introduced the Flux 2 family and then the Klein speedier, smaller variant that now could take many references. and then now we're looking at I mean people and this kind of comebacks comes back again to like what I was just saying of having this like representation of how things relate and that the better we actually build this like representation the more interesting ways you can essentially like tie in all of these different things so I mean you know throughout you know one of the most commonly used or I don't want to say commonly used but like kind of obvious cases was okay like clothing drive I'm like this is a very you know here's a picture of me here's a picture of some clothes um could you please uh you know put you know show me what this looks like i i've seen people like i i actually did some home decoration earlier this year just trying to see like what different couches and furniture look like in my place um one of the most interesting uses i'll say that kind of stood out to me that i saw at a hackathon um was and i don't know how much this could be used for like a real planning purposes but i just thought it was interesting is someone was taking pictures of fire exits in a building and then generating what it would look like if a crowd was trying to leave through this fire exit in an emergency so that they could actually like brilliant yeah so they could actually gauge like what would this look under like emergency like where is the crowd like i don't and obviously there's you know there's a generated component so you have to take it with a little bit of grain of salt but you still could get a general idea of this is what this would look like under the scenario and i thought that without derailing us You just sparked a whole bunch of ideas on that in my head.
34:41So like, I'm just like, oh, this is great stuff. Keep going. Sorry.
34:46Daniel Whitenack:If you've been listening to the show over the past few months, you realize just how transformative agentic AI is, whether that's Claude Code or Hermes Agent or custom built software that you're deploying for operational efficiencies or as new products to your customers. Regardless of your maturity now, this is the world that we're headed towards, this agentic AI world. And there's a lot of security and governance teams that aren't letting these agents go into production because of risks related to agency and autonomy. And how do you take care of things like prompt injections or insecure tool usage?
35:27Daniel Whitenack:There's a lot to take care of, and that's why I'm personally spending my time outside of the show, working with an amazing team of AI engineers to build Prediction Guard. Prediction Guard is an AI control plane that you run in your own infrastructure behind your firewall. Developers can build on top of this control plane using everything that they want to use, OpenAI and Anthropic compatible APIs, MCP servers, frameworks like Langchain. But all of this is plugged into a built-in governance harness that enforces your organization's AI policies. And all of that telemetry goes back to your monitoring and alerting systems.
36:07Daniel Whitenack:I would encourage you to check out what we're doing at predictionguard.com slash practical AI. You can schedule a demo with me and the team, and I'd love to get your feedback on what we're doing. So visit us at predictionguard.com slash practical AI. That's predictionguard.com slash practical AI. These are all, so you mentioned the Flux family of models. So that's Black Forest's family of models, or at least some of the models that you've worked on. Could you just kind of give us a concrete, you mentioned a couple by name, but like give us a tour of the family of models. And then maybe where, I know that there's someone hugging face.
36:50Daniel Whitenack:People can find them. people can can look at them but maybe just give us a little bit of a tour of the model family and then um anything that uh that you want to share about some of the some of the distinctions between them uh different ones of them yeah yeah i mean um we're i i would say fundamentally like as a core we're still like a i'll say we're still like a research lab that wants to push the boundaries so with each one of our model releases we want to try to like level up some capability of the model not just, you know, make a general improvement, but like really see some new versioning with it.
37:24So, I mean, I can take back through, there's not that many models to take back through a little bit of a history. You know, it was a little, what was it around two years ago now, we released the first, our first family models, which was the Flux series. And this was the first set of models that we created after we formed the company, after a very fun, but intense, I think it was around four or five months sprint to really like build our chops here and get something out. And with that, we came up with this, at the time, this essentially distinction between the models where we released the Flux1.
37:58So we had the Flux1 Pro model, which was on our API. We had the Flux1 Dev model, which was this commercially licensable model, but the weights were open for people out in the world to use. And then we had the Flux1 Schnell model, which was a high-speed, like, step-distilled model that was totally, I don't know if it was MIT or Apache, but it was open to use for whatever purposes you want. um then from there uh we uh worked on our flux to flux tools series of models where we realized like we needed a lot more control with the models and this is where alongside this work started the context project to okay the these people want a lot more control with the models these models are learning a lot of relationships how can we leverage this which led up to the flux context release which i talked about this was our big edit release everything i mentioned up to this point was still kind of our like historical models now we're getting into our our more up-to-date models although they're getting a little bit older now but we have a i don't know something i'm excited coming i don't want to get dates soon hopefully later this summer that'll be very exciting to talk about when it comes out but um looking forward to it yeah yeah we'll have to get your back on definitely yeah happy to happen to come back but then we got to our flux 2 series of models and flux 2 was a very big upgrade for us um where we really pushed the capabilities not not just on T2I, but also on editing, not just on single image.
39:22But this is where we also do introduce like the multi-image, the kind of omni-edit where you could put many people, many different items, have all sorts of relationships. This is where we started to see, you know, not just these kind of more standard advertising use cases, which still were like the most common and we really want to support this, but like some more, you know, interesting how people kind of build, you know, like things like this fire exit thing or other types of stuff. And then right after that, we came to our client series of models which was essentially like a size distillation where we really wanted to pack as much performance as we could into a small model for this, you know, for both use on our API, but also just for open weight release.
40:00We know that there's many people out in the world who use our models locally and we really wanted to make sure they had something powerful they could use. So it was a text image and an editing model. It was a very fast model and it was pretty small in comparison to, you know, it was even smaller than our Flux 1 series. And then we've done a further speed up on that with our Klein KV model, which introduced, I believe for the first time, KV caching, which is this, I'll just say it's an optimization technique that's very common with the language model world that we were able to bring into the editing world to get a very, very big speed up on local editing.
40:35So people who wanted to actually use these models locally, but also, you know, we also serve this as well. you know now now um since then we've done a couple blog releases one fairly fun uh research you know research blog on something called self flow which i'm repping partially because i am on there i'm not not the lead author our lead author is incredible heila she's amazing but um and then now we're we're kind of all pushing forward to our next big release which is hopefully going to be uh
41:04Daniel Whitenack:i i'm quite excited about it but that's awesome and just to follow up on that i know some people out there, our listeners are always trying things on their laptop or wherever they're pulling down things. I know for quite some time when I tried to access some of these models and run them myself, it was either very difficult to find with my limited resources, the right kind of configuration to run this, but also it was sometimes incredibly slow. So what to just give a sense of some of those, you know, more hardware optimized models that you mentioned, I think the client and other models, how maybe my question is, can I reasonably run one of these models on my laptop now and create a great image?
41:56Daniel Whitenack:Like what's what's required here? Certainly, I'm not going to serve my production web app in an enterprise environment off of my laptop, But just maybe give us a little bit of sense, because that has changed on the LLM side, right, where that has continually updated. And obviously, the smaller models are not up to the same output quality as the larger models. But you can run a small model, you know, even on a CPU now, at least for some tasks, that is pretty reasonable. So how has that progression happened on the visual intelligence or image generation side? Yeah, I would say I'll first state that the LLM world definitely has a lot more people working on it.
42:37So they definitely have a lot more of this up and down. That said, with our Klein series, I haven't run this personally. Maybe I'm not the best person in the world to speak on this, but I'm 99 % certain you could run this on, say, like a modern M series Mac, like a MacBook Pro. So as for speed, I don't have any numbers I can quote because I haven't tested this myself. But this is always like a big kind of trade-off in this world of how, you know, even on the language side, you know, you see this scaling of, you know, like a lot of the new big open releases that people are excited about are getting into the hundreds of billions of parameters.
43:14And this is why we, you know, this is why we did like the Klein series, for example, because Flux 2 was a 32E model. It was quite chunky. You could run it on like a, you know, higher end local computer, but it was slower to run locally. You needed some more powerful, you know, computing. This is where like the Klein series popped in. But it's always a tradeoff because we want to push the bat, you know, we want to push performance and make something that we're really proud of. And then we want to try to bring that again back into a smaller, faster scale. And this is something we've continually done, you know, with our Fluxon series, we had Schnell.
43:48With Flux2, we had Klein. I don't want to promise anything going forward, but we definitely want to keep bringing this bigger power that we generate into these smaller models that people can hopefully run locally. Absolutely. Super cool. I'm excited to try some of those. As we are starting to wind up, and you guys have done some really, really cool work in this space, one of the things we always like to ask as we're closing out these episodes are kind of like we want to get a glimpse into your head, into your thinking about what's to come and kind of to frame it a little bit as you're thinking about like, you know, you're out of the day's work and you're, you know, enjoying your evening and your mind wanders and you're kind of thinking about like down the road, what would you know, what you want, what your passion is taking you and what you'd like to see.
44:43Can you share a little bit about what the future that you would like to create with visual intelligence in terms of what's next that you guys haven't addressed. I'm not asking for a product release so much because I know those need to stay in, but kind of just like what is your aspiration? What's your dream in terms of where these kind of capabilities might ultimately lead? And love to get your kind of your dream context, if you will, as we finish up. Sure. If I'm allowed to give two split answers. Absolutely. Yeah, totally fine. There's kind of two areas that constantly sit on the front of my mind.
45:24At some point, they'll hopefully become one. But I imagine the next and this is no no glimpse into saying what we're doing, but I'm just this is stuff I'm personally excited about. Fair enough. Well, I would say the two areas that I'm the most excited about is one is long context, truly multimodal models. So nowadays, obviously, we're seeing lots of people work with agents constantly. And on our front, where we're having these more continuous generative models, we're starting to introduce more and more contextual usage with these references. And I'm looking forward to when this kind of bridges a little bit more.
46:04And we have models that not just can like, okay, you have say the agent calls the generative model. But when this kind of becomes more or less a model that not just can, you know, do work in the text and language and agent, but actually maybe can think visually as well and generate audio for you and has the full context of say all the stuff you've done over the last few weeks. You know, you don't need to say, give it the reference, give it the right prompt. it just like already has this context and can reference this as needed um in this like continuous space this excites me a lot the other side i'll say is the real-time stuff that's uh you know we're seeing a little bit of it now out there i think this is like very early and i think it's going to be very exciting for real-time video audio duplex interactions um you know being able to i don't know you see a little bit of these like uh interactive you know stuff with like genie or play the game, but also on the side of, you know, you can bridge this back into like robotics where it needs to take in the real world in real time and make decisions.
47:04Yeah, I was gonna say that sounds really familiar in terms of interest on that side of things. So yeah, really cool. Great conversation and your lead into what Black Forest Labs is doing was also really good contextually in terms of kind of explaining. So hope our audience got a lot out of that. Dustin, thank Thank you very, very much for coming on the show. Great conversation. And as you've hinted, there are things to come. And I'm looking forward to having our next conversation as things move forward a little bit. Absolutely. Thank you so much for having me, guys.
47:44All right. That's our show for this week. If you haven't checked out our website, head to practicalai.fm and be sure to connect with us on LinkedIn, X, or Blue Sky. you'll see us posting insights related to the latest ai developments and we would love for you to join the conversation thanks to our partner prediction guard for providing operational support for the show check them out at predictionguard.com also thanks to breakmaster cylinder for the beats and to you for listening that's all for now but you'll hear from us again next week
From the publisher
How has AI image generation evolved from blurry outputs to powerful visual intelligence models? Dustin Podell, Co-Founder and Researcher at Black Forest Labs, explains the progression from diffusion to flow matching, how modern image models work, and how they're being used for image editing and practical visual workflows. The conversation also explores the FLUX family of models, running image generation locally, and where visual AI is headed next.
Featuring:
- Dustin Podell – LinkedIn
- Chris Benson – Website, LinkedIn, Bluesky, GitHub, X
- Daniel Whitenack – Website, GitHub, X
Links:
- Black Forest Labs
- Developer Dashboard
- Research Page
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- Laying the Foundations for Visual Intelligence
Sponsors:
- Midwest AI Summit: Join AI practitioners on October 15 in Indianapolis for practical sessions, hands-on discussions, and real-world AI solutions. Use code PracticalAI20 to save 20% on your registration. https://midwestaisummit.com/#tickets
- Prediction Guard: A self-hosted AI control plane for running agents in high impact environments. predictionguard.com/practicalai
Upcoming Events:
- Register for upcoming webinars here!
- Midwest AI Summit 2026




