Unified Latents (UL): How to train your latents

26 Feb 2026 · 20 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Unified Latents (UL) by Google DeepMind Amsterdam—how to train latent representations for diffusion-based image/video generation more efficiently, with controllable “information capacity” (bits-per-latent) to improve fidelity and reduce compute.

Guest backgrounds

No named guests appear in the transcript; it’s a host-led discussion with an interviewer/second speaker.

Key claims

Traditional latent diffusion (e.g., Stable Diffusion) uses VAEs with manually tuned KL penalties, causing loss of high-frequency detail and posterior collapse (decoder ignores latents). UL unifies encoder, diffusion prior, and diffusion decoder into one diffusion-centric pipeline, fixes encoder noise deterministically, replaces JAN with a diffusion decoder, and uses a two-stage training (train then freeze, then train generator).

Notable examples

“hamster in a tuxedo delivering a keynote to cats”; chocolate sauce over vanilla ice cream (texture preservation); video on Kinetics-600 achieving FVD 1.3; ablation removing the diffusion prior (L2 regularization) reduces performance.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring AI Image Generation Limitations

0:46 to 2:42

Discussion on the inefficiencies and trade-offs in current AI image generation models.

“And the more I read about how these models actually work, the more I realize that for all their brilliance, tools like stable diffusion have been, well, kind of inefficient under the hood.”

Understanding the Latent Diffusion Model

2:43 to 4:35

Explanation of how the latent diffusion model compresses images and the challenges faced.

“Tell me yellow sunflower with 12 petals.”

The Problem of Posterior Collapse

4:36 to 6:00

Examining posterior collapse and its effects on image generation quality.

“And the consequence of that guesswork is that the model often discards high-frequency details just to satisfy the rulebook.”

Introducing Unified Latents Framework

6:01 to 9:09

Overview of the Unified Latents Framework and its innovative components.

“The encoder, the prior, and the decoder.”

The Two-Stage Training Process

9:10 to 11:44

How the two-stage training process enhances model stability and performance.

“It doesn't actually match the original data you asked for.”

Control and Information Management

11:45 to 14:00

Exploring how the new framework manages information flow for better results.

“It's like teaching a painter how to mix paints perfectly before you ever let them touch a canvas.”

Exploring the Unified Latent Framework

14:00 to 18:04

Learn about the advantages of the unified latent framework in generating images.

“They benefit from the high bitrate to produce those hyper-realistic results.”

Shifts in AI Philosophy for Image Generation

18:04 to 19:12

Discover how unified latents signify a shift towards smarter data representation in AI.

“So putting this all together, we have a system that treats image generation not as a series of disjointed, hacked-together steps, but as a unified, elegant flow of diffusion.”

The Complexity of Human Imagination

19:12 to 20:17

Contemplate the implications of quantifying human concepts through AI.

“We usually talk about bi-traits for things like MP3s or Netflix streams, like we said.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, I was playing around with one of those AI image generators last night. It was late. I was pretty tired. And I typed in something just completely ridiculous. I think it was a hamster in a tuxedo delivering a keynote speech to a boardroom of cats. Please tell me it wasn't a serious presentation. Oh, very serious. He had a little laser pointer and everything. Naturally. Right. But here's the thing. It took maybe, what, four seconds? Boom. Whiskers, the texture of the little velvet suit, the sheer boredom on the cats' faces. It was all there. It creates that illusion effectively, doesn't it?

0:34It really does. It feels like magic. But then I started thinking, we treat these things like magic wands, but there was this massive sweating computational engine just churning away in the background to make that happen. Yeah, sweating is a good word for it. And the more I read about how these models actually work, the more I realize that for all their brilliance, tools like stable diffusion have been, well, kind of inefficient under the hood. Inefficient is a very polite way to put it. We've essentially been brute forcing a lot of these problems for years now. We throw massive amounts of compute at them and just kind of hope the math sorts itself out.

1:10Right. And that brings us to what we're digging into today. We were looking at a new framework coming out of Google DeepMind Amsterdam. It's called Unified Latence or just UL. UL, yeah. And from what I'm seeing, this isn't just a tiny tweak. It sounds like they're fundamentally rethinking how an AI compresses and understands an image before it even tries to draw it. It is a fundamental shift. And to understand why this matters to anyone listening, we really have to talk about the central tension in AI image generation. It's a tug of war that every researcher deals with, the tradeoff between reconstruction and generation.

1:47Okay, let's unpack this, because usually when we talk about tradeoffs in tech, it's always speed versus quality, right? But this sounds different. It is different. Think of it this way. To generate a high-resolution image, say a 4K picture of your keynote speaking hamster, the AI cannot work with every single pixel at once. Right, it's just too heavy. Way too much data. It would melt the GPUs. So the model uses what we call latence. We've touched on latence a bit before. It's kind of like a compressed version of the image. Somewhat. But let's use a better analogy. Imagine you are trying to describe a famous painting to me over a really bad phone line.

2:22Okay. The connection is terrible, so you can't describe every individual brush stroke. You have to use shorthand. You say blue sky, sunflower, yellow. So the latent is that shorthand description. Exactly. Now, here is the conflict. If your shorthand is too simple, if you just say flower, I might draw a rose when you actually meant a sunflower. The final image loses detail. That's bad reconstruction. Okay, so just be more specific. Tell me yellow sunflower with 12 petals. Right, but here's the catch. If the description gets too complex, if you start listing coordinate points and hex codes for the colors, the signal gets too dense.

2:59The receiving end, the AI's brain essentially, gets overwhelmed. Oh, I see. It can't figure out the pattern because it's drowning in data. That is bad generation. So that's the trap. Make the description simple, and the result is blurry or just wrong. Make it complex, and the AI is too confused to understand it. Precisely. And the mission of our deep dive today is to look at how unified latent solves this exact problem. They've managed to turn every part of the pipeline, the compressing, the regulating, and the creating, into a single unified diffusion process. And the big claim is they can get better quality using significantly less computing power, right?

3:36Much less, yes. Which, given the crazy energy costs of AI right now, is a huge deal. But before we get to their new solution, I need to understand why the old way was broken. How were we handling this bad phone line problem before? The industry standard for a long time has been the latent diffusion model. That's the architecture behind things like the original stable diffusion. And they use something called a VAE. A variational autoencoder. I've seen that acronym thrown around. Right. A VAE tries to compress an image into that latent code. But to keep that code organized so the AI can actually read it later, they apply a penalty.

4:10It's called a KL penalty. A penalty. Like a fine for being too wordy on the phone. Sort of. Going back to the phone line analogy, imagine there's a strict limit on how many words you can use. The KL penalty is the rulebook enforcing that limit. Okay. But in traditional VAEs, the weight of that penalty is completely manual and arbitrary. Researchers literally just pick a number and hope it works for the data set. Wait, really? So it's basically vibes-based engineering? In a way, yeah. And the consequence of that guesswork is that the model often discards high-frequency details just to satisfy the rulebook.

4:45High-frequency meaning the small stuff. Exactly. Things like the texture of fur or really sharp edges or the grain of wood. It throws them away to make the math easier and meet the penalty requirements. And I assume that's why sometimes older AI images look a bit smooth, like everything is made of plastic or fondant. That is exactly why. It's smoothing over the difficult parts. But there's an even worse problem that happens. It's called posterior collapse. Posterior collapse. That sounds, I don't know, medically concerning. It's definitely computationally concerning. It happens when the decoder, the part that listens to your phone description and paints the picture, gets too powerful.

5:23How can it be too powerful? Well, it looks at your compressed description and says, you know what? The signal is too fuzzy. I don't need this. I'll just guess what the image is supposed to be based on what I already know. Wait, so the artist ignores the description entirely and just paints whatever they want? Pretty much. The decoder ignores the latent instructions completely, and you end up with generic, unfaithful images. You ask for a hamster in a tuxedo, you get a generic fuzzy blob. That is wild. Yeah, and the Unified Latence Framework was designed specifically to stop that from happening.

5:56Okay, so let's look at this solution. The Unified Latence Framework. DeepMind breaks this down into a triad. Three parts working in harmony. The encoder, the prior, and the decoder. But they aren't working the way they used to, right? No, not at all. In the past, these parts were often trained separately, or they didn't really talk to each other very effectively. Unified latence treats them as a cohesive system. Let's start with the encoder. The part that creates the shorthand description. Right. In a standard VAE, the encoder outputs a distribution, a whole range of possibilities. Unified latence does something really clever instead.

6:30It outputs a noisy latent. Hang on. Noisy. Usually in audio or video, noise is the enemy. Why on earth would we want to add noise to our data on purpose? That's the intuitive reaction. But remember, the core engine here is a diffusion model. Diffusion models don't just tolerate noise, they thrive on it. Their entire job is to take noise and organize it into a coherent image. Ah, so we're just speaking the model's native language. Exactly. But here is the key insight. They don't ask the encoder to learn how much fuzziness to add. They make it deterministic. Meaning it's fixed. Yes. The encoder creates a sharp code, and then the system mechanically adds a fixed, mathematically known amount of Gaussian noise.

7:11So instead of the encoder guessing, you know, maybe I'll add a little dust here, the system just says, here is your code, and we are going to sprinkle exactly this much dust on it every single time. Precisely. It explicitly links the encoder's noise to the minimum noise level of the next part of the chain, which is the prior. And this removes a huge variable for the AI to worry about. Exactly. You don't have to train the model to learn the variance. It's just fixed. Which brings us to part two, the diffusion prior. This is the regulator. Yes, this is the middleman. His job is to model the distribution of these latents.

7:43Because the noise from the encoder is fixed, the complex math that usually governs this, that nasty KL penalty I mentioned earlier, It simplifies beautifully. Oh, I like the sound of that. It turns into a weighted MSE, or mean squared error. Okay, I promised myself no heavy math today. What does that simplification actually give us in plain English? Why do we care? It gives us an interpretable upper bound on the bi-trait. Okay, bi-trait, that I know. Like when I'm streaming a movie on Netflix and the connector drops, the bi-trait plummets and it gets blocky. Exactly like that. When you stream a movie, the bitrate determines exactly how much information is coming through the pipe.

8:22Unified latence allows the model to know exactly how many bits of information are packed into that image. So it's not guessing anymore. It can actually measure the information density. Yes. It knows the exact capacity it's working with. So we've got a fixed noisy code from the encoder and a prior that knows exactly how much info is in there. Now we need to turn it back into a picture. That's part three, the decoder. And this is probably the biggest departure from the old way. Most models, including the original stable diffusion, used JAN-based decoders. JANs. Generative adversarial networks. That's the architecture where two AIs basically fight each other, right?

9:00Yeah, one generates, one critiques. But JANs are tricky. They often hallucinate texture. Hallucinate, like making things up. Yes, they make things look incredibly sharp, but the sharpness is fake. It doesn't actually match the original data you asked for. Yeah. So Unified Latence replaces that JAN with a diffusion model as the decoder. So it's just diffusion all the way down. All the way down. And because the decoder is now a diffusion model, it can handle a specific type of math called re-weighted ELBO. ELBO. Is that an acronym or some kind of wrestling move? It stands for evidence lower bound.

9:34Okay, slightly less exciting. But think of it as a budget. In the old models, high-frequency details, like the texture of a cat's fur or the woven threads on a suit, they were considered expensive to calculate. So the model would just smooth them over to save its budget, cut corners? It cut corners constantly. But this new decoder makes those details cheap. It creates a cost per bit that actually encourages the model to obsess over the texture. Oh, that's brilliant. It doesn't collapse or ignore the input like we saw with posterior collapse because it is explicitly trained to denoise those exact fine details.

10:08Now, I really want to talk about how they actually built this thing Because looking at the research, it seems like they hit a massive wall during development. Usually when we hear the word unified, I just assume they threw everything into one giant pot and hit the train button. You would think so. That's usually the dream of end-to-end learning in AI. And actually, they tried exactly that. They tried to train the encoder, the prior, and the decoder all at the exact same time. They did. And what happened? It didn't work. Or at least it wasn't stable at all. Why not? I mean, if it's a unified system, shouldn't they learn together?

10:41Think about it like a sports team. If you're trying to teach the quarterback a brand new play, but the receivers are also learning how to run their routes at the same time, and the coach is changing the rules of the game every five minutes, nobody learns anything. Right. The targets are just moving way too fast. Exactly. So they had to implement a two-stage training process. This is a really crucial detail. Walk me through it. Stage one, they train the tools, they train the encoder, the prior, and the decoder to just be really, really good at compressing and decompressing images. So get the compression engine working perfectly first, independent of anything else.

11:18Right. Then, and this is the key, they freeze them. They literally lock those weights in place. So the tools stop changing, the rules of the game are set. Exactly. Then comes stage two. They train the base model, the creative part that actually generates the new image concepts, to use those specific frozen latents. That makes so much sense. By freezing the latents, the base model has a completely stable target to aim for. It learns much faster because the ground isn't shifting under its feet anymore. It's like teaching a painter how to mix paints perfectly before you ever let them touch a canvas.

11:53If you try to teach them color theory and complex brush strokes at the exact same time, they're just going to get confused and make mud. That's a perfect analogy. And speaking of making things more efficient, they also used an engineering trick called 2x2 patching. 2x2 patching. Sounds like something I do to fix holes in my drywall. It's a very similar concept of covering ground efficiently. Instead of processing every single latent pitchle individually, which takes forever, they group them into little 2x2 squares. Okay, so a block of 4 pixels. Yes. And that cuts the sequence length down by a factor of 4.

12:26Which saves compute power. Huge amounts of compute. It makes the encoder and decoder incredibly fast without losing the structural integrity of the image. So we've got the triad, we've got the training strategy, the patching. Now let's talk about control. You mentioned the byte rate earlier. One of the coolest things about this framework is that researchers can actually turn a knob to control how much info flows through the system. Yeah, this is what they call the loss factor or the information valve. They use a parameter called sigmoid bias to do it. The research shows they can tune exactly how much information goes from the encoder to the generator.

13:02So what happens if I tune the valve down? Low by trait. If you set a low by trait, you are sending a very simple, highly compressed instruction. This is actually easier for the base model, the creative engine, to learn. It doesn't have to memorize much. But the downside is? The downside is the reconstruction is lower fidelity. You might get the concept of a dog perfectly, but the specific pattern of spots on its back might be generic instead of exactly what was in the training image. Or text might be blurry. Okay. And if I crank the valve wide open, high bite rate. You get perfect reconstruction.

13:34Every single hair is exactly in place. But now you need a massive brain, a huge base model, to handle that level of complexity. So there's a Goldilocks zone they have to find. Exactly. And the findings on this were fascinating. They found that smaller models actively benefit from lower bitrates. They essentially say, don't overwhelm the small guy with details. Let him focus on the big picture. The massive models. They can handle the fire hose. They benefit from the high bitrate to produce those hyper-realistic results. And the beauty of the unified latent framework is that it lets them find that optimal spot mathematically, rather than just guessing like they used to.

14:12Yes. It explicitly bounds the information capacity. It essentially says, we know this image contains exactly X bits of information. Do we want to use all of them or compress it? It's a level of precision we haven't really had in this space before. So does it actually work in practice? Or is this just really beautiful math on paper? Because I know they tested this against some heavy hitters. Oh, they definitely put it to the test. They compared it directly to stable diffusion and DDIT, which are diffusion transformers. Okay. They tested it on ImageNet 512, which is basically the standard benchmark for this kind of thing.

14:46And the metric they use here is FID freccia inception distance. I always forget for FID, do we want a high score or a low score? Lower is better. Right. It measures the distance between real images and the generated ones. Unified Latence achieved a competitive FID of 1.4. Wow. But the headline isn't just the score itself, it's how they got that score. They achieved this with significantly fewer training FLOPs. Floating point operations, so fewer math problems for the computer to solve. Exactly. It creates better images for less energy and less compute cost. They aren't just brute forcing it with a warehouse full of GPUs anymore.

15:25They are being fundamentally smarter with the data representation. I was looking at the visual examples they provided. Researchers always picked the weirdest prompts to prove a point. Like a fox dressed in a suit dancing. Or an astronaut riding a horse. It's basically tradition at this point. It is. But the one that really stood out to me was it was pouring chocolate sauce over vanilla ice cream. Ah, yes. That's the texture test. It has to be, right? Liquids, gloss, the porous surface of the ice cream. If you have that vibes-based compression we talked about earlier, that chocolate sauce just looks like flat brown paint.

15:58Completely flat. Right. But in the Unified Latents examples, it looked remarkably realistic. Like, actually tasty. That's the high-frequency detail preservation shining through. The decoder didn't smooth out the glossy highlights on the sauce or the little microscopic air bubbles in the ice cream. It knew those bits were important and preserved them. And they didn't just stop at static images. They took this to video. Yeah, this was the ultimate proof of concept for scalability. They applied unified latence to the Kinetics 600 dataset. Which is what, just a massive library of videos of people doing things?

16:33Pretty much. And video is notoriously difficult for AI because you have to be consistent over time. If your latents are unstable, the video flickers constantly. Your hamster's tuxedo would change from black to blue to gray every single frame. Right. That horrible shimmering effect you see in a lot of early AI video. Exactly. But unified latents achieved a state-of-the-art FVD for show video distance of 1.3. Wait, state-of-the-art? Like the best out there? At the time of the research, yes. It proves that this unified approach scales beautifully to temporal beta. It can compress time just as effectively as it compresses flat pixels.

17:09That is incredible. I also saw, and I love this detail, they tried to be lazy in one of their experiments. They did an ablation study where they just removed the diffusion prior entirely. Oh, the L2 regularization experiment, yeah. They essentially asked themselves, do we really need this fancy diffusion prior in the middle? Can't we just use a simple math penalty like we used to and save some time? And the answer was? A resounding no. When they replaced the diffusion prior with a simple L2 regularization, the performance dropped significantly. So there are no shortcuts. No shortcuts. It proves that having a smart prior, a diffusion model that actually watches and understands the structure of the latence is absolutely necessary to get that high quality.

17:50It's the difference between a security guard who just counts the number of people walking into a building versus a guard who actually checks their IDs to make sure they belong there. That's a great way to put it. The diffusion prior understands the structure of the data it's regulating. A simple math penalty is completely blind. So putting this all together, we have a system that treats image generation not as a series of disjointed, hacked-together steps, but as a unified, elegant flow of diffusion. Controlled noise in, controlled noise out. That's the perfect synthesis. Unified latence represents a massive shift in philosophy.

18:25We are moving from make the model bigger to make the data representation smarter. It feels like we're really growing up a bit in the AI space. We're moving past the throw more compute at it phase and entering an era of actual precision engineering. It's about efficiency and control. We are talking about bits per dimension now. Image generation is becoming a highly quantifiable science rather than just dark alchemy. And that's what makes this so exciting for anyone listening. Because better efficiency means these models can eventually run on smaller devices. It means getting high quality video generation on your laptop without needing to rent time on a massive server farm.

19:04It definitely opens the door for real time high fidelity generation that doesn't have that plastic AI generated look. Before we wrap up, I want to leave you with a thought that popped into my head when we were talking about buy traits. Go for it. We usually talk about bi-traits for things like MP3s or Netflix streams, like we said. But here, we're actually assigning a mathematical bi-trait to a concept, to a fox in a suit. Right. If unified latence allows us to perfectly control and measure the exact bi-trait required to generate a convincing image of reality, are we approaching a point where we can quantify the exact complexity cost of human imagination?

19:41Wow. That is a heavy question. Think about it. At what point does the compression become completely indistinguishable from the concept itself? If I can describe a sunset over the ocean in exactly 400 bits, and it looks utterly perfect to the human eye, have we just mathematically solved the concept of a sunset? It definitely suggests that reality, or at least human perception of it, might be far more compressible than we ever thought possible. Something for you to mull over while you pour some chocolate sauce on your ice cream tonight. Thanks for joining us on this deep dive into unified latence.

20:15We'll catch you on the next one. See you then.

From the publisher

This paper introduces Unified Latents (UL), a novel framework designed by Google DeepMind researchers to improve the efficiency of generative diffusion models. By integrating the encoder, prior, and decoder into a single cohesive system, the method creates a more principled approach to managing latent representations. A key innovation involves using a fixed amount of Gaussian noise during encoding, which simplifies the training process and allows for a tight bound on latent information. This setup provides users with specific hyper-parameters, such as the loss factor and sigmoid bias, to precisely balance reconstruction accuracy against the complexity of the modeling task. Experimental results across image and video datasets demonstrate that this architecture achieves superior performance with lower computational costs compared to established baselines like Stable Diffusion. Overall, the research highlights how jointly optimizing all components of the latent space leads to higher-quality generative outputs and more effective scaling.

More from Best AI papers explained

All 475 episodes
Unified Latents (UL): How to train your latentsBest AI papers explained · 20 min
Listen in VO