Back to Basics: Let Denoising Generative Models Denoise

23 Nov 2025 · 15 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains why diffusion image models that predict noise (epsilon/velocity V) can fail under capacity limits, and argues that “back to basics” clean-image (X) prediction plus a minimalist “Just Image Transformers” (GT) architecture improves efficiency and stability. It frames the key idea using the manifold assumption: clean images lie on a low-dimensional manifold, while noise is off-manifold and chaotic in high-dimensional pixel space, requiring excessive capacity.

Guest backgrounds

No guests are named; it’s a solo “Deep Dive” discussion.

Key claims

Noise prediction collapses when network capacity is constrained (toy spiral up to D=512; ImageNet 256x256 with ViT and 256 hidden units). GT works with zero pretraining, no latent tokenizer, and no auxiliary losses; it uses X prediction but optimizes with V loss for better gradients. Bottlenecks (e.g., 3072→32–512) improve FID by ~1.3.

Notable examples

2D spiral buried in 512D; ImageNet FID degradation/blurry artifact “static” outputs; GT patch dimensionality scaling up to 12,288.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Paradox of Noise vs. Structure

0:45 to 1:49

Discussing the inefficiency of predicting noise versus clean images in diffusion models.

“And yet, for years, the most successful, most widely used models, they didn't predict the clean image X at all.”

Understanding Manifold Assumption

1:49 to 4:30

Explaining the manifold assumption and its implications for predicting clean images.

“We're going to explore this just image transformers approach or GT.”

Capacity and Prediction Insights

4:30 to 6:53

Insights on how network capacity impacts the prediction of images versus noise.

“Just fundamentally wasteful of the network's resources.”

The Development of GT Models

6:53 to 10:07

Exploring the development of the GT model and its minimalist design approach.

“Sometimes you just need to change the question.”

Architectural Innovations from Language Models

10:07 to 12:22

Examining how concepts from language models improve image generation.

“And this whole exploration, it led to some other really fascinating kind of counterintuitive architectural insights.”

Implications Beyond Image Generation

12:22 to 14:01

Discussing the broader implications of the GT approach for various scientific fields.

“It's a really efficient way to encode relative positional information.”

Exploring Denoising in Generative Models

14:01 to 15:19

Discover how focusing on simplicity and signal can enhance generative models.

“Computational biology, material science, maybe weather modeling.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we're diving right into, well, the very heart of modern image generation. The very core. We're looking at the architecture behind these diffusion models, the ones creating just unbelievable state-of-the-art visuals. But our sources point to this, this really profound and, frankly, fascinating contradiction right at the center of their success. Oh, it's a huge technical moment for the field. It really is. It suggests the standard way of doing things, the most popular method. It was fundamentally inefficient. Especially when you try to scale. Because think about it, we call these denoising diffusion models.

0:39Right. The whole point is to remove noise. Exactly. The goal is to strip away all the synthetic noise to get to the clean, structured image, the thing we call X. Not yet. And yet, for years, the most successful, most widely used models, they didn't predict the clean image X at all. Which is so counterintuitive. It is. Instead, they trained the network to predict the noise component itself, what we call epsilon, or sometimes this hybrid quantity, the flow velocity V. So there were models trained on chaos, not on structure. That's a perfect way to put it. Why did the community do this? That's the question.

1:10Why go the noise prediction route if the prize was always the clean image? Was it just easier to get working? Initially, yeah, you could say that. Early work found that predicting the noise often gave you more computational stability. The optimization was a bit simpler, especially with certain sampling techniques. So it was a pragmatic choice. It was. It delivered quick wins. But as we're going to explore today, the paper argues that stability came at a massive, massive cost. A cost in what? In capacity efficiency. It led to these catastrophic failures when the models were constrained, particularly in really high dimensional settings.

1:48And that's our mission today. We're going to explore this just image transformers approach or GT. GT, yeah. It's all about going back to basics. Predicting the clean image X again and showing how this seemingly simple switch fixes those fundamental limitations. Okay, let's really unpack this foundational idea. This whole deep dive seems to hinge on capacity, on geometry. It does. What is the fundamental difference then between training a network to predict the clean image X versus the noise epsilon? It all comes down to something called the manifold assumption. Okay. This has been a longstanding hypothesis in machine learning.

2:27The idea is that natural structured data, so a clean image of a cat, a face, whatever, it's not random. It has rules. It has rules. It lies on this highly constrained, low-dimensional surface. We call it the data manifold. And that manifold is embedded within a much, much bigger high-dimensional space. So if you picture the raw pixel space as this huge, mostly empty room. A vast empty room, yes. The actual clean images only exist on a very thin structured ribbon running through it. That's the perfect analogy. So the clean image X is on manifold. It has that low dimensional structure. It's predictable.

3:04In a sense, yes. Now contrast that with the noise, epsilon. The noise is inherently off manifold. It's everywhere else in the room. It's everywhere. It is chaotically distributed across the entire massive high dimensional space. If we take, say, a 512 by 512 pixel image that's hundreds of thousands of dimensions, the noise lives randomly in all of them. So if you tell a network your job is to predict the noise, you're forcing it to model the chaos of that entire room. Exactly. And that is a massive burden. That sounds incredibly difficult. It is. Predicting the noise requires huge network capacity.

3:40The network has to be able to represent all the information about that high dimensional chaos. If it misses some, the result is just wrong. But here's the core insight of this whole thing. If you train the network to predict the clean data, X, a limited capacity network, can do that job extremely well. Wait, so having less capacity is an advantage here. How does that work? Because the network only needs enough capacity to model the low-dimensional information of the manifold of that ribbon. It doesn't need the capacity to store or process all the high-dimensional noise. Ah, so its own limitations act as a filter.

4:14Precisely. The network's architectural constraints, its limited hidden units for example, they forced it to focus on isolating the signal. It automatically has to reject the high-dimensional chaos around it. It's like an efficiency hack based purely on geometry. It is. It suggests that trying to model the noise is a brute force approach. Just fundamentally wasteful of the network's resources. Especially when the inputs like high-res images get enormous. And that takes us straight into the proof. The theory held up perfectly when they stress tested it. When they forced networks to operate with limited capacity in these huge input spaces.

4:50That's when predicting the noise just broke completely. So walk us through that. What was the evidence? I hear they started with a pretty clever stress test. They did. They called it the toy experiment. It was a really smart setup. They created simple underlying data, just a two-dimensional spiral shape. Easy to learn. Very easy. But then they buried it. They observed this 2D spiral in a massive D-dimensional space, and they kept increasing D all the way up to 512. Okay, so the signal is tiny and the noise is huge. And crucially, the model they used was deliberately constrained. Just a simple network with only 256 hidden units.

5:25So when the input dimension hit 512, the input was literally double the network's processing capacity. Exactly. It was a test of what happens under a capacity deficit. And the result was stark. What happened? Only X prediction managed to produce anything reasonable as the dimension D went up, both epsilon and V prediction. They failed. Catastrophic. You mean they just produced garbage? Total garbage. At D equals 512, they just couldn't separate the 2D signal from the 512 dimensional noise. The capacity wasn't there to model the noise, so the whole thing collapsed. Wow. And this isn't just about theoretical spirals.

6:00This translated directly to real-world image generation. Absolutely. When they ran this on ImageNet using 256 by 256 pixel images with a standard vision transformer, the results were identical. The noise prediction models just fell apart. They completely collapsed. It didn't even matter which loss function they used. They produced terrible FID scores. Just total failure. Okay, let's make that clear for everyone. FID is the standard quality metric for Shed Inception Distance. But what does a terrible FID score actually look like? Visually, it means the generated images are a mess. They're blurry.

6:35They're not cohesive. You get these weird texture blobs or artifacts, sort of like static. There's no high-level structure, no clear objects, no seeds. That is noise. Just noise. The noise predicting models produced garbage, while the X predicting models delivered strong performance right out of the box. So the lesson is pretty clear. Just throwing more capacity or problem isn't the only answer. Sometimes you just need to change the question. Right. Changing the prediction target from chaos to signal simplifies the problem so much that the capacity constraints actually become manageable. Exactly.

7:06And that realization led them straight to developing GT. So let's talk about GT. Why the name just Image Transformers? It sounds so simple. It's fitting, isn't it? It highlights the extreme simplicity and the self-contained nature of the design. What do you mean by that? I mean, GT is conceptually nothing more than a plain vision transformer, a VIT, operating directly on large patches of raw pixels. So it's going back to first principles for diffusion. It is. And what's really impressive is what they got rid of. It's a truly minimalist design. What's missing? So much of what we assume is necessary, it's self-contained because it relies on zero pre-training and zero auxiliary losses.

7:49No pre-training at all. None. No latent tokenizer, which can introduce its own biases. No adversarial loss. No perceptual loss. No self-supervised pre-training to get it started. It just takes the raw high-dimensional pixels directly. Directly. And they were really aggressive about pushing that raw dimensionality. I saw that. When they scaled up to high resolution, the patches they fed into the transformer were massive. They were immense. For 256 by 256 images, the patches were 768 dimensional. But for 502 by 512 images, they were dealing with 3 ,072 dimensional patches. And for the 1 in 24 by 1 in 24 generation?

8:26For that, they used a staggering 12 ,288 dimensional patch input. Okay, let's put that in perspective for a second. A standard base-size transformer model has an internal hidden dimension of 768. That's right. So you have a 768-dimensional brain trying to process a 12 ,000-dimensional chunk of raw visual data. That deficit is the entire point. That's the test of the manifold assumption in practice. And the fact that it works is the proof. It's the proof. The network is fundamentally under-complete compared to the raw data. but it succeeds because it's only learning the low dimensional structure of X.

9:04It's being forced to throw away the noise because it physically doesn't have the capacity to model it. It can only retain the signal. I do want to clarify one technical point though. They advocate for X prediction, but the final most successful model didn't use the simplest X loss function, did it? It was a bit of a hybrid. That is a key detail. It separates the prediction target from the optimization mechanism. The catastrophic failure came from forcing the network to predict an off-manifold quantity. But the final algorithm they settled on was X prediction. The network outputs the clean image X, but it was optimized using the V loss, the velocity loss.

9:40Right. Why use the V loss if V prediction was a problem? Because the velocity loss is a mathematical tool that, kind of independent of the prediction target, has been shown to provide better gradient flow. A smoother learning process. A smoother, more stable path for the network to follow during optimization, especially early in training. So they coupled the best prediction target, which is X, with the most stable optimization mechanism, which is the VLOS. That makes perfect sense. Use geometry to filter the input and use optimization theory to guide the learning. Precisely. And this whole exploration, it led to some other really fascinating kind of counterintuitive architectural insights.

10:19About model efficiency. Yeah. The first, as we mentioned, was just how well it worked when the hidden dimension was so much smaller than the patch input. That really challenges the instinct to just scale everything up symmetrically. It confirms that a smaller network can be more efficient. If it's focused on the right thing. The low-dimensional manifold, not the high-dimensional noise. And this leads right to the finding about the bottleneck. This one really surprised me. Intentionally restricting the flow of information at the input stage. It just sounds wrong. It sounds totally counterintuitive, doesn't it?

10:52Why would you deliberately reduce capacity and expect performance to get better? Exactly. But it's one of the most exciting nuggets in the paper. They tried replacing the initial linear patch embedding layer with a bottleneck structure. So two layers with a much smaller dimension squeezed in between. A very tight squeeze. And they found that reducing that bottleneck dimension, taking a patch of, say, 3 ,072 dimensions and forcing it down into a range of maybe 32 to 512. It helped. It was consistently beneficial. It improved image quality, the FID score, by up to around 1.3 points. The restriction, paradoxically, made the model better.

11:30So the bottleneck is acting like a really effective built-in filter. I think that's exactly right. It forces the network to immediately find and discard all the irrelevant high-dimensional noise and just focus on the true structural dimensions it needs for the manifold. Precisely. It's an architectural reinforcement of the manifold assumption. It echoes all these principles from classical manifold learning where finding the true inherent dimensionality of your data is everything. And the final point was about borrowing ideas from language models. Yeah. We see all these acronyms, SWIGLU, RMSNORM, ROEP.

12:08Why are these advances from text generation suddenly showing up and improving raw image generation? I think this just speaks to the philosophical generality of the transformer architecture itself. These are general purpose improvements for stability and efficiency. Can you give an example? Like ROPEY. Sure. ROPEY, Rotary Positional Embeddings. It's a really efficient way to encode relative positional information. Crucial for a transformer to understand word order in a sentence. But how does that apply to image patches? Well, when you apply it to VIT patches, it helps the network manage the complex spatial relationships between the image segments much more efficiently.

12:42And something like RMS norm. RMS norm, or root mean square normalization, is just a lighter, simpler, more computationally efficient way to normalize activations than the traditional layer norm. So it's faster. It's faster, and it's more stable. The same goes for things like this with GLU activation function. All these things just make the transformer a better, more stable learner, regardless of whether the input is text tokens or raw pixel patches. It really confirms that the core recipe, diffusion plus transformer on native data, is this powerful universal thing. Highly adaptable. Okay, let's synthesize this.

13:17The GT architecture is basically a minimalist manifesto. I like that. By going back to predicting the clean image, X, and leveraging these architectural insights like the bottleneck to filter noise, they get strong results at massive resolutions. Up to 20, 24 by 1, so 24. All while keeping the architecture efficient and incredibly simple. It really forces us to rethink what we consider necessary complexity. By just trusting the geometry of the data, they found a path to efficiency that doesn't rely on just throwing more parameters at the problem. And the implications here go way beyond just generating pretty pictures.

13:53Far beyond. So where does this kind of self-contained, domain-agnostic approach shine next? Well, think about scientific fields that deal with raw, really complex, high-dimensional data. Like what? Computational biology, material science, maybe weather modeling. In these areas, engineering a good, unbiased, latent tokenizer or figuring out how to pre-process the data is often prohibitively difficult. A huge bottleneck in itself. It is. Chichichos, you might not need all that heavy domain-specific engineering. You could rely on a pure transformer working on raw data to find the underlying structure directly.

14:31You could accelerate discovery in fields where we're currently stuck. Stuck trying to model the noise. Exactly. That's the powerful takeaway then. And it leaves us with a final thought to chew on. Given that predicting the clean, low-dimensional structure proved so much better than modeling the chaotic, high-dimensional noise, what fundamental capacity limits are we overlooking in other domains? It raises the question, doesn't it? Are we focusing on inputs that are inherently off-manifold? Are we currently wasting billions of parameters trying to map the messy chaos of our input spaces instead of just training the network to isolate and predict the pure, low-dimensional structure of the output we actually care about.

15:09If the network prefers learning elegance over learning chaos, then focusing on the signal is the most intelligent form of efficiency there is. A compelling idea to carry forward. Thank you for guiding us through this deep dive. It really feels like a look into how simplicity and a focus on the signal are reshaping generative AI.

From the publisher

This academic paper, introduces "Just image Transformers" (JiT), a novel approach to denoising diffusion models that advocates for directly predicting clean data (**x-prediction**) rather than predicting noise or a noised quantity. The authors argue this shift is critical based on the **manifold assumption**, which posits that clean data lies on a low-dimensional manifold while noise is inherently off-manifold. Experiments, including a toy model and high-resolution ImageNet generation using plain Vision Transformers (ViT), demonstrate that x-prediction successfully handles high-dimensional spaces where conventional noise-predicting methods catastrophically fail. This research emphasizes a return to first principles for a self-contained **"Diffusion + Transformer"** paradigm on raw pixel data, without relying on complex architectures, pre-training, or auxiliary losses. Ultimately, the paper provides extensive ablation studies on loss combinations and architectural components to validate that **x-prediction** is fundamentally more tractable for limited-capacity networks in high-dimensional generative modeling.

More from Best AI papers explained

All 475 episodes
Back to Basics: Let Denoising Generative Models DenoiseBest AI papers explained · 15 min
Listen in VO