When Models Don’t Collapse: On the Consistency of Iterative MLE

3 Feb 2026 · 16 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Whether “model collapse” (AI trained on its own outputs degrading in a feedback loop) is inevitable for iterative Maximum Likelihood Estimation (MLE), and what mathematical conditions prevent or cause it.

Guest backgrounds

No guest names or bios are provided in the transcript; it’s a two-host discussion.

Key claims

Under standard statistical assumptions, iterative MLE with accumulating real+synthetic data can remain consistent even if real data becomes a tiny fraction, provided MLE’s probability model is “smooth” (continuous, bounded derivatives) and identifiable. If the model is “spiky” (sharp probability spikes), collapse can occur immediately after one synthetic mistake (“fast” collapse) or via delayed “slow death” over many rounds.

Notable examples

“Soccer field” flat distribution with a hidden narrow spike; one synthetic point lands in the spike, forcing MLE to create a massive probability spike. “Staircase” scenario where each round unlocks a slightly wrong interval, leading to later failure. Mentions softmax and weight decay as mechanisms that may reduce spikiness.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Maximum Likelihood Estimation (MLE)

1:35 to 3:06

Learn about MLE and its implications for AI training processes.

“To really get this, we have to understand the specific type of training being analyzed.”

Iterative MLE and Accumulating Data

3:07 to 5:01

Discover how iterative MLE can prevent model collapse with proper data practices.

“And the question becomes, as the years go by, as the number of these rounds or iterations goes to infinity, does the model drift away from the truth?”

The Risks of Spiky Distributions

5:02 to 8:44

Examine the dangers posed by spiky data distributions in model training.

“So if we just keep the history, if we don't throw away the old internet while we build the new one, we might actually be okay.”

The Importance of Smoothness in Models

8:45 to 12:30

Understand why smoothness in data relationships is crucial for model stability.

“The model was perfectly consistent on real data.”

Maintaining Historical Data for AI

12:31 to 14:05

Learn about the significance of retaining old data to prevent AI decay.

“There's another detail here that I found fascinating.”

Preserving Intelligence in a Machine-Generated World

14:05 to 16:04

Explores the importance of maintaining data integrity in AI systems.

“This work suggests that even if that happens, even if 90 % of the text out there becomes machine generated, intelligence can still be preserved.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, let's unpack this with a visual. I want you to picture a snake, but not just any snake. That ancient symbol, the Auroboros. The snake eating its own tail. The very one. A classic image of, you know, infinity or maybe, maybe inevitable self-destruction. And that's pretty much the go-to metaphor for the biggest anxiety in AI right now. Exactly. I mean, the idea is simple, right? AI models are trained on the internet. But now the internet is just filling up with stuff written by, well, by other AI. So future models are going to be trained on the output of past models. Creating a feedback loop.

0:35And the fear, the real fear, is that this loop creates a downward spiral. They call it model collapse. Right. We're worried that if AI just feeds on its own, you know, sludge long enough, it'll get dumber, weirder, and eventually just detach from reality entirely. It's a really valid concern. We've seen studies suggesting that without fresh human data, models start to amplify their own biases and lose touch with the ground truth, the actual reality they were supposed to be learning. But, and this is where it gets really interesting for our deep dive today, there's some new thinking out there, some new math that says, hold on, maybe put down the panic button.

1:10Or at least check the math before you panic. We're looking at some fascinating new theoretical work that takes a rigorous look at this feedback loop. And the conclusion. We aren't necessarily doomed to a future of AI nonsense. But, and it's a massive but, it all depends on the shape of the math. So are we looking at an infinite intelligence or a snake choking on itself? Let's dive in. To really get this, we have to understand the specific type of training being analyzed. The focus is on something called Maximum Likelihood Estimation, or MLE. Which is pretty much the bread and butter of modern statistics.

1:47It is. In simple terms, MLE is a method where you look at some data and you try to find the model parameters that make that data, well, most likely to have happened. Like a detective working backward from the clues to find the culprit. You want the story that fits the evidence best. That's a great way to put it. Now, the common wisdom is that if you feed an MLE model its own output, it degrades. The photocopy of a photocopy analogy. Eventually, the image just turns into a black smudge. That's the fully synthetic loop. If you throw away the original and only learn from the copy, yes, you lose information.

2:21But this new work looks at a different setup, iterative MLE with accumulating data. And this is the first key difference. Let's use an analogy. Imagine a library. In that photocopy scenario, every year you burn all the books and replace them with books written by an AI that only read last year's books. A recipe for disaster. You lose the nuance, the history, everything. Right. But in this accumulating data scenario, you keep all the original books. You just add the new AI-written books to the shelves. Exactly. You start with, say, end samples of real ground truth data. You train Model 1. Model 1 then generates new synthetic samples, new books.

2:58And you mix them together. You mix them. Now, Model 2 trains on everything, the original real books and the new synthetic ones. Then Model 2 writes more books. You add those and so on. So the pile just gets bigger and bigger. Precisely. And the question becomes, as the years go by, as the number of these rounds or iterations goes to infinity, does the model drift away from the truth? Because, you know, eventually that library is going to be 99 % AI fan fiction and only 1 % original human literature. And logic says the 99 % trash should drown out the 1 % truth. The signal-to-noise ratio would just be destroyed.

3:34You'd think so. But this brings us to the first major finding, the good news. Okay, I'm ready for some good news. Under standard statistical assumptions, model collapse is not inevitable. Wait, even if the real data is a tiny fraction, like 1 percent? Even if the fraction of real data approaches zero, as long as the total amount of data keeps growing, the model can stay consistent. It can stay anchored to that original truth. How is that even possible? How does the 1 percent hold the line against the 99 percent? It's all about the efficiency of the MLE method. The math shows that the original data, even if it's buried, provides just enough of a signal to keep the parameters aligned.

4:12The synthetic data, because it comes from a model that learned from the real data, it essentially echoes that truth. So it's not pure poison. It's more like diluted information. Exactly. It may be a bit fuzzy, but the average direction still points the right way. But there is a requirement on the sample size. Ah, of course. There's always fine print. The number of new samples you add in each round, n, needs to be, and apologies for the jargon here, polylogarithmic in the number of rounds, t. Okay, hold on. Polylogarithmic. Let's unpack that. I'm assuming that doesn't mean an impossible amount of data.

4:48Quite the opposite. It's a very manageable growth rate. It means your data set doesn't need to explode exponentially to keep up. You just need a reasonable amount of new data relative to how many training rounds you're doing. So if you do that, the math holds up. The model converges to the true parameters. So if we just keep the history, if we don't throw away the old internet while we build the new one, we might actually be okay. We could survive the flood of synthetic content. Theoretically. You said understandard statistical assumptions? I did. And that sounded like the structural load-bearing wall of this entire argument.

5:21It is. The good news only applies if the world, and by extension, the model, is smooth. Smooth. Like, smooth jazz. or a calm ocean? More like calculus. We're talking about regularity and smoothness. It means the relationship between the model's parameters and the probability it assigns to data needs to be continuous. It needs to have bounded derivatives. So no sharp corners, no sudden cliffs where the probability just drops from 100 to 0 instantly. Exactly. The transition between likely and unlikely has to be gradual. And the model needs to be identifiable, meaning different parameters actually give you different results.

5:58if those conditions are met, accumulation saves you. And I'm guessing the real world isn't always so smooth. And that is where we get to the bad news. Because if the good news says we're safe, why are so many researchers seeing model collapse in practice? Right. If the math says it should work, why does it break when people actually try it? Because sometimes that smoothness assumption fails. And what's so clever here is they didn't just prove it works when it's smooth. they proved exactly how it breaks when it's spiky. Spiky sounds unpleasant. It's catastrophic. The researchers constructed these specific mathematical scenarios, these gotcha scenarios, where MLE works perfectly on real data but just falls apart the moment you introduce its own synthetic data.

6:43How does that happen? Walk us through the trap. Okay, imagine a probability distribution that's mostly flat, like a uniform distribution. Most of the data falls safely between 0 and 1. Think of it like a flat soccer field. The ball almost always lands on the field. Okay. Safe, predictable. But there is a tiny hidden trap, a potential spike. It's a kind of dependent interval, let's call it a floating zone, that can get incredibly narrow and incredibly tall, a huge spike depending on a certain parameter in the model. Like a landmine hidden on the field. Effectively. Now, in round one, we train on real data.

7:18The real data is safe, it's all on the soccer field. The model looks at it and says, okay, the world is flat, no spikes here. So it learns the real data perfectly, estimates the parameters correctly. Yes. But then it has to generate synthetic data. And because it's a probabilistic model, it isn't perfect. It has to roll the dice. It generates thousands of new points. And just by the laws of probability, eventually it's going to make a mistake. Correct. Statistically, it's almost guaranteed that one of those synthetic data points will land in that trap zone, a range where the probability should be nearly zero, but just due to a tiny error, a synthetic point lands there.

7:56And then the model has to train on its own mistake. And this is the collapse. Because the distribution is spiky, it's not smooth, the MLE tries to maximize the likelihood for that one weird data point. And to explain that single point, it has to shift its entire understanding of reality. It thinks, wait a minute, if a data point landed here, there must be a massive probability out and right here that I missed. That's it, exactly. It creates a massive probability spike around that single error just to rationalize it. It completely abandons the flat field reality of the original data just to accommodate this one synthetic hallucination.

8:32So it hallucinates a pattern based on one bad synthetic example. And because it's a feedback loop, that new spike becomes the new reality. The parameters shift massively away from the truth. The math shows this can happen immediately after just one iteration. One round. So much for infinity. And this is crucial. The model was perfectly consistent on real data. It looked totally fine, but it was brittle. It couldn't handle its own imagination, so to speak. It feels like there's a lesson there. Just because you can handle the truth doesn't mean you can handle your own lies. Precisely. And the scary part is the collapse isn't always this dramatic explosion in round one.

9:12There's also what you could call the slow death. That sounds even more ominous. It was shown that you can design a distribution where the collapse doesn't happen right away. It might look stable for 100 rounds or 1 ,000. So you're watching the charts, everything looks green, accuracy is high. And then disaster. They constructed a scenario that works like a staircase. An error in round one unlocks a new, slightly wrong interval, which leads to an error in round two that unlocks the next interval. It's the slow, creeping drift away from reality. So you could be looking at your model in Generation 50 thinking, wow, this is stable, we've solved it.

9:47And then at Generation 51, it falls off a cliff. That's terrifying for anyone building these systems. It means you can't just run a test for a week and declare victory. It proves a vital point. Consistency on real data is not enough to guarantee safety from model collapse. You explicitly need those smoothness assumptions. Without smoothness, time is against you. So this creates a really interesting conflict. You have some studies saying collapse is inevitable and other studies saying it's fine. And this new thinking seems to be saying you're both right. It just depends on the geometry. That's a perfect summary.

10:21It synthesizes the field. If you look at the more doom and gloom studies, they often used models or scenarios that mathematically were inherently unstable or violated these conditions. And the more optimistic studies. They often looked at things like linear regression or settings where data accumulation was strictly enforced and the map is just naturally smooth. So the million dollar question, maybe the trillion dollar question, is where do large language models fit in? The GPTs and clods of the world, are they smooth or are they spiky? That is the question. And there's some reason for optimism here.

10:57Oh, good. I need some optimism after hearing about the slow death. Well, think about how LLMs are built. At the very end, to calculate probabilities, they use something called a softmax function. Right, softmax. That takes all the raw numbers and turns them into clean percentages. Essentially, yes. But mathematically, what softmax does is create a smooth curve. It forces probabilities to be non-zero and continuous. It doesn't allow for those hard vertical cliffs or absolute zeros that cause the collapse in those spiky counterexamples. So the architecture itself kind of fights against spikiness.

11:31It does. And on top of that, we use techniques like weight decay during training. That's a form of regularization. And regularization is basically about telling the model not to get too obsessed with any single data point, right? Exactly. It penalizes the model if it tries to create those massive artificial spikes, just to explain one outlier. It forces the model to look for broader, simpler patterns. All of this suggests that modern LLMs might actually be closer to the safe zone than the danger zone. Provided, and I think I'm getting this now, provided we don't throw away the old data. Yes. That accumulation part is absolutely non-negotiable.

12:08If LLMs start training only on the output of the last six months of the Internet, we could still be in trouble. But if they keep the whole history, the smoothness of the architecture might just save us. It's interesting because we usually think of old data as outdated or useless. In tech, if it's from 2019, it feels like ancient history. But in this context, the old data is the anchor. It's the ground truth. Even if it's a small percentage of the whole, it prevents the model from drifting into pure fantasy. It's the ballast in the ship. There's another detail here that I found fascinating. The focus was on unsupervised learning.

12:43Yes. Meaning no human is sitting there checking the synthetic data and saying, this is junk, throw it out. Which is huge because at the scale of the internet, we can't check everything. You can't hire a billion people to proofread the output of the next GPT. Exactly. So if the math says unsupervised accumulation works, that is incredible. Those aren't just optimization tricks. They are effectively the firewall against model collapse. They are the practical way of enforcing smoothness. Precisely. This new thinking gives us the mathematical reason why those techniques are so critical. They enforce smoothness.

13:18They prevent the spikes. They ensure that one hallucinated fact doesn't rewrite the model's laws of the universe. It really changes how I look at this whole AI eating itself problem. It's not just about the quantity of synthetic data. It's about the quality of the model's underlying geometry. If the geometry is smooth, you can digest the synthetic data. If it's jagged, you choke. So what does this all mean for you, the listener? I mean, if you're just a user of AI, not a machine learning engineer, why should you care about all this? I think it matters because it changes the narrative of the future.

13:52We're in this phase of extreme anxiety about dead internet theory. The idea that the internet is already dead and fake. Yeah, that the future is just bots talking to bots in an echo chamber and we're all just watching from the sidelines. This work suggests that even if that happens, even if 90 % of the text out there becomes machine generated, intelligence can still be preserved. We aren't automatically destined for a stupidity singularity. That's a relief. It means the tool I'm using five years from now might still be useful, even if it trained on its own output for years. But only if the developers respect the math.

14:29It puts the burden on the architects of these systems to ensure they maintain that history of data and enforce those smoothness constraints. It's a call to be rigorous. You can't just train on whatever you find lying around anymore. You can't just scrape the new internet and ignore the old internet. No, you have to be a curator of the archive. I love that image. Yeah. The AI as the librarian who throws a book away. A very capable, slightly overwhelmed librarian. Right. Standing in a room full of fan fiction, but still holding onto the original manuscripts. Before we wrap up, there's one last thought that really struck me here.

15:02A model can be perfect at learning from reality, yet terrible at learning from itself. Yes. That spiky model learned the ground truth perfectly in round one. It aced the test on reality. It makes me wonder, is there a fundamental mathematical difference between observation, seeing real data, and imagination generating synthetic data? That is the frontier, isn't it? We assume that if a model generates data, that data is just as good as the real thing if the model is high quality. But this shows that synthetic data carries a hidden risk, a potential for these feedback loops, that real data simply doesn't have.

15:41Real data is messy, but it's honest. It comes from the physical world. Synthetic data is clean, but it creates this perfect echo chamber. It's like the difference between looking out a window and looking into a mirror. The mirror might be a perfect reflection, but if you put another mirror in front of it... You just get lost in the reflection. And we are only just beginning to map out that territory. We're figuring out how to keep the mirrors from shattering. Well, on that note, we're going to leave you to ponder the geometry of your own learning curves. And remember, don't throw away your old books.

16:10Until next time, keep your gradients smooth.

From the publisher

This research explores model collapse, a phenomenon where generative models degrade after being repeatedly trained on their own synthetic outputs. The authors provide a theoretical framework using Maximum Likelihood Estimation (MLE) to determine when this process can be avoided. They demonstrate that if models meet specific regularity and smoothness assumptions, they can remain consistent and accurate even as the proportion of real data diminishes. Conversely, the study provides the first rigorous proof that without these structural assumptions, model collapse can occur abruptly or over time, even when real data is preserved. Ultimately, the findings suggest that data accumulation alone does not guarantee stability; rather, the underlying mathematical properties of the distribution family are what prevent performance failure.

More from Best AI papers explained

All 475 episodes
When Models Don’t Collapse: On the Consistency of Iterative MLEBest AI papers explained · 16 min
Listen in VO