In short
The episode explains “self-improving pretraining,” a method to train better AI models without relying on ever more high-quality human data, aiming to overcome the “data wall” around 2026. It argues standard next-token pretraining rewards copying toxic or low-quality text, risking “model collapse” and poor safety.
Guests/backgrounds
No specific guest names are provided in the transcript; the “cast” is conceptual: the student (policy model), the rewriter (strong existing model such as GPT-4/Llama-3 class), and the judge (another strong model such as Llama-3 Instruct/GPT-OSS class).
Key claims
Use a rewriter to transform bad suffixes into safe, factual text (not deleting the trap) and a judge to score original vs rewritten vs student rollout, updating via online DPO so the student eventually surpasses the teacher.
Notable examples
“Red Pajama” as a toxic web-like dataset; reported win rates include 86.3% quality improvement, quality baseline 1.3% to 32.4% safety-related win rate, and 36.2% relative factuality improvement on benchmarks like Hello Evil and TruthfulQA.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Challenge of Model Collapse
0:45 to 2:19
Discussion on issues surrounding data quality and model training.
“And usually when I hear people talk about this, they say, oh, we'll just use synthetic data.”
Problems with Traditional Pre-training
2:19 to 4:38
Analysis of flaws in existing AI training methods and their consequences.
“I feel like most people have this misconception about how AI learns.”
Introducing Self-Improving Pre-training
4:38 to 6:45
Explaining the new approach to AI training and its key components.
“So filtering makes it naive and post-training is too little, too late.”
The Role of AI in AI Training
6:45 to 8:11
How different AI models interact to improve learning outcomes.
“It's like a parent reading a problematic book to a child, but improvising the ending to teach better moral rather than just ripping the pages out.”
Real-world Application and Results
8:11 to 12:12
Results from experiments showcasing the effectiveness of the new method.
“And this leads to the training arc, which is really the narrative heart of this whole process.”
The Trade-off of Speed for Quality
12:12 to 13:19
Discussing the computational cost of the new training paradigm.
“See, the old assumption was to get a smarter model, we need more human data.”
The Future of AI Intelligence
13:19 to 14:00
Speculation on AI's self-sufficiency and potential evolution.
“If more data isn't an option, better compute is the only lever we have left to pull.”
The Future of AI Learning
14:00 to 14:58
Explore whether AI can surpass human intelligence through self-improvement.
“And that leads to the really big question.”
Defining AI Quality
15:00 to 15:12
Consideration of who defines quality in AI-generated outputs.
“We talked about the model learning to judge its own outputs.”
Transcript
Automatic transcript. May contain errors.0:00You know, there's a specific year that keeps popping up in conversations about the future of AI. And it's actually keeping me up at night a little bit. That year is 2026. Ah, 2026. Let me guess. This is about the data wall, isn't it? It is exactly about the data wall. For everyone listening, there's this this looming idea that roughly around 2026, we're going to run out of high quality public text to train AI models on. We'll have essentially read the entire Internet, every book, every article. It sounds dramatic, but it's a legitimate bottleneck. The current paradigm is just feed the beast. You make the model smarter by feeding it more data.
0:38But if the cupboards bear, what happens then? Do we just, you know, hit a ceiling where AI stops getting smarter? That is the trillion dollar question. And usually when I hear people talk about this, they say, oh, we'll just use synthetic data. You know, we'll have AI write books for other AIs to read. But that sounds risky to me, like making a photocopy of a photocopy. Eventually, doesn't it just become a blur? It absolutely can. If a model just trains on its own unregulated output, you get what we call model collapse. It starts drifting into nonsense because there's no grounding in reality or quality.
1:10So we're stuck. We're running out of human data and machine data is potentially dangerous. But, and this is why we're here today, there's a new approach that claims to smash right through that wall. It's called self-improving pre-training. This is a really significant shift. The research we're looking at today proposes that we don't need more data to make models smarter. We need a better process. It claims we can take the messy, toxic, error-prone internet we already have and use it to train a model that is actually better than the data it reads. Which sounds a bit like alchemy. You're turning lead into gold.
1:45In a way, yes. But it's engineering, not magic. It's about moving from a model that simply, you know, copies the internet to a model that actively critiques and improves upon it while it learns. So today, we are going to unpack how you teach an AI to be smarter than its teachers. We have a fascinating cast of characters to meet, the student, the rewriter, and the judge. And honestly, the numbers on this are startling. They really are. This isn't just theory. The results are impactful. Before we get to the solution, though, I think we need to properly frame the problem. I feel like most people have this misconception about how AI learns.
2:22They picture a student in a library carefully reading encyclopedias and learning facts. Yeah, it's a very romantic image, but the reality is much more chaotic. It's less like a library and more like, imagine locking a toddler in a room with a high-speed connection to Reddit, 4chan, Wikipedia, conspiracy blogs, and Twitter, and just saying, read all of this and learn how to speak. Right, and then we're surprised when the toddler comes out screaming insults. Exactly. This is the garbage in, garbage out problem. Standard pre-training is built on something called next token prediction. It's incredibly simple.
2:54I give the model a prefix, the start of a sentence, and its only job is to guess the suffix, the ending. Okay, so if I type the sky is? The model predicts blue. Simple enough. But what if the data is bad? What if the prefix is part of like a toxic rant found on a message board? That is where it breaks. If the training data contains a hate speech rant and the model is given the first half of that rant, the correct answer, mathematically speaking, is to finish the rant. To be a good model under standard pre-training, it has to be a good mimic. It gets a gold star for accurately predicting the toxicity.
3:31So we're effectively rewarding the AI for being terrible, just because the internet was terrible first. That is the fundamental flaw. We're optimizing for probability, not quality. We're teaching it to fit in, not to stand out. Now, traditionally, the tech companies know this is an issue. So how have they been fixing it? Do they just scrub the data? They try. There are two main old school methods. One is filtering. You try to delete all the bad stuff before the model sees it. Which sounds logical. It is, but it creates the sheltered child syndrome. If you delete every instance of toxicity or error, the model never sees them.
4:05Then you deploy it into the real world, and the first time a user tries to trick it or says something nasty, the model has no idea what's happening. It has no immune system. It creates a naive model. Precisely. The other method is post-training. That's where you train the model on the messy internet first, and then afterwards you try to slap a safety filter on top or fine-tune it to be polite. That feels like building a house with a crooked foundation and then trying to fix it by hanging the picture straight. That's a perfect analogy. The bad patterns are already baked into the neural pathways.
4:35You're just putting a band-aid on a deep wound. So filtering makes it naive and post-training is too little, too late. This brings us to the breakthrough, self-improving pre-training. How does this change the game? It changes the objective. Instead of guess the next word exactly as it appears in the database, the goal becomes generate a high-quality, safe, factual sequence. Okay, but high quality is subjective. I can't write a math equation that defines quality. You can't. But other AI models can. This is where we introduce that cast of characters we mentioned. We're setting up a system where AI teaches AI.
5:13Let's meet the team. First up, we have the policy model. That's our student, the baby model we're trying to train. Okay. And then we have the rewriter. Think of the rewriter as a guardian or a very smart editor. This is an existing, powerful AI, maybe something like a GPT-4 or a Lama 3 class model. Its job is to sit between the raw internet data and the student. Cleaning the stream. Exactly. Imagine that river of text flowing in from the web. You have a prefix, the context, and a suffix, the ending. The rewriter looks at the suffix. If it's good, factual, clean, well-written, it lets it pass. The student learns from it.
5:50But if it's garbage, if it's a conspiracy theory or a toxic comment... Then the rewriter intervenes. It takes that bad suffix and rewrites it into something safe and factual. Now, I want to pause here because this distinction feels crucial. We aren't deleting the bad data. The rewriter is actively changing it. Why is that better than just hitting delete? Because of the immune system concept we touched on earlier. If the rewriter sees a prefix that's a trap, say someone asking for illegal instructions, and simply deletes it, the student never learns to recognize the trap. Right. But if the rewriter keeps the trap, the prefix, but changes the response to a polite refusal or a safe pivot, the student learns, ah, when I see this kind of dangerous setup, this is how I steer out of it.
6:36Is building safety muscles. Yes. It's learning the maneuver, not just ignoring the obstacle. It turns bad data into a lesson on how to handle bad data. That makes so much sense. It's like a parent reading a problematic book to a child, but improvising the ending to teach better moral rather than just ripping the pages out. Exactly. You keep the context, but correct the behavior. So the student is seeing better data thanks to the rewriter. But simply reading good answers isn't enough, right? The student has to actually try to do the work. Correct. Learning by reading is passive. To really learn, the student needs to generate text.
7:12This brings us to the judge and the concept of rollouts. Rollouts. It sounds like something from the Transformer movies. What's a rollout in this context? A rollout is just a fancy term for the student model taking a shot at the answer. the student looks at the prefix and tries to generate its own suffix. It raises its hand and guesses. Yes. And now, the judge, another smart AI model, steps in. The judge has three things to look at. One, the original text from the internet. Two, the rewritten text from our smart editor. And three, the student's own attempt at the rollout. And the judge decides who did it best.
7:46It scores them. And then we use a technique called online DPO, or direct preference optimization. Okay, let's unpack DPO just a little bit without needing a whiteboard. Think of it as A-B testing on steroids. The judge says, hey student, your attempt was okay, but the rewriter's version was much better. You should update your internal math to look more like the rewriter. Or conversely, hey, your attempt was actually better than the original data. Good job. Reinforce that behavior. So it's a constant feedback loop. It's continuous reinforcement. And this leads to the training arc, which is really the narrative heart of this whole process.
8:19A journey of the student. Right. In the beginning, early in training, the student is, well, it's dumb. It's a random number generator. Its rollouts are gibberish or just bad. So the judge is constantly pointing to the rewriter and saying, copy that guy. Exactly. The student learns by mimicking the teacher. But as training goes on, the student gets smarter. It starts understanding the underlying patterns of quality and safety. And eventually, a crossover happens. The student surpasses the master. The students' rollouts start getting higher scores than the original data. The model stops just copying and starts synthesizing.
8:55It begins to generate answers that are better, safer, and more factual than the raw internet text it started with. That is the self-improving part. It's bootstrapping its own intelligence. Yes, it is pulling itself up by its digital bootstraps. It effectively becomes a better writer than the authors of the data it was trained on. That sounds incredible in theory, but I am always the skeptic when it comes to lab results versus real world. Did they actually test this? Did it work? They did. They ran extensive experiments using LAMA 2, a standard 1.4 billion parameter model as the student. And they used stronger models like LAMA 3 Instruct or GPT-OSS as the judge and rewriter.
9:32So they gave the toddler some professors as tutors. What happened? They looked at three pillars, quality, safety, and factuality. Let's start with quality. They measured generation quality, basically. Is the text coherent? Does it flow? Is it useful? In a continual pre-training setup, the self-improving model achieved an 86.3 % win rate over a model trained the standard way. Whoa. 86.3%. That is not a small margin. That is a landslide. It's a different class of performance. The model became less repetitive, more coherent, and just generally a better writer. Okay, that's quality. But what about safety?
10:07We talked about the bad neighborhood of the Internet. This is where the experiment gets really interesting. They trained the model on a data set called Red Pajama. Red Pajama. Which sounds gozy, but I know it's not. Not at all. Red Pajama is an open source data set that attempts to replicate the full raw web. It includes the good, the bad, and the ugly. It is full of toxicity and bias. So it's a stress test. A massive one. Standard models trained on Red Pajama typically become toxic because they just mimic the data. but the self-improving model. In a from-scratch experiment, its win rate for quality jumped from a baseline of 1.3 % to 32.4%.
10:45Hey, hang on. From 1.3 % to 32.4%. I'm doing the math. That's colossal. It shows that without the rewriter and judge, the model was drowning in the bad data. With them, it learned to swim. It learned to recognize the toxicity and steer away from it, effectively filtering out the poison while keeping the nutrition of the language. That is huge. And finally, the third pillar, factuality. Because there is nothing worse than a smart-sounding AI that confidentially tells you lies. Hallucinations. It's the plague of large language models. Standard models don't care about truth. They care about what sounds plausible.
11:18If the internet says the moon is made of green cheese enough times as a joke, the model might just repeat it as fact. In this research, they used a specialized factuality judge, specifically designed to spot these hallucinations and penalize the student for them. And the result. A 36.2 % relative improvement in factuality metrics compared to standard pre-training. They tested this on benchmarks like Hello Evil and Truthful QA. So simply by having a judge that explicitly says, no, that's made up, try again, the model actually learned to stick to the facts. Yes. It learned that sounding good isn't enough.
11:55It has to be accurate to get the high score. This brings us back to that data wall conversation we started with. If we can get these kinds of games, 86 % better writing, massive safety jumps, better factuality, using the same dirty data we already have, does that mean the wall is a myth? It might mean the wall is permeable. See, the old assumption was to get a smarter model, we need more human data. This approach suggests to get a smarter model, we need a better process. We don't need more books. We need to be better readers. Exactly. And it touches on a deeper philosophical shift in how we build these things.
12:28Ideally, we want to move away from explicit instruction. Which is telling the AI exactly what to say word for word. Right. That's brittle. We want to move toward incentives. Giving it a goal and letting it figure out the path. Precisely. When you use a judge, you aren't micromanaging every token. You are saying, here's what a good answer looks like. Figure out how to get there. This method is a huge step in that direction. It's teaching the model how to think about its output rather than just how to copy its input. That sounds amazing, but I have to play devil's advocate here. This sounds heavy.
13:01Standard pre-training, guessing the next word is fast. You just zoom through the text this new way. You have the rewriter running, then the student guessing, then the judge scoring. Oh, it is computationally expensive. You're absolutely right. You're trading speed for quality. It's much slower and burns more GPU hours. So it's not a free lunch. No, but as we hit that data wall, that trade-off becomes essential. If more data isn't an option, better compute is the only lever we have left to pull. And I imagine it scales better than human labeling, right? Yeah. Usually we have armies of humans reading this stuff to tag it as toxic.
13:35Humans are slow, expensive, and we get tired. And we disagree on things. If you can replace the human labeler with an AI judge, suddenly you can clean the entire internet. You can run this process 24-7 at a scale no human team could ever match. That is the part that feels like science fiction. We're effectively building machines that curate reality for other machines. We are. And that leads to the really big question. If the student can eventually surpass the teacher, and if the model can learn to be better than the data it reads... I see where you're going with this. Are we approaching a point where AI doesn't need us anymore to get smarter?
14:13The self-licking ice cream cone of intelligence. I was going to say recursion, but sure, let's go with the ice cream cone. If an AI can generate data, judge that data, and improve on that data, It creates a closed loop. It can climb the ladder of intelligence entirely on its own. That is both the most exciting and the most terrifying sentence you've said today. It's the promise of the singularity in a microscale. But for now, let's just be happy. It means fewer chatbots sending us toxic rants or telling us the moon is made of cheese. I will take that win for now. Me too. This has been a fascinating look under the hood.
14:46It turns out fixing the internet might just require a machine that knows how to ignore the worst parts of it, or at least rewrite them into something useful. And that's a lesson we could probably all learn, honestly. True enough. Listeners, here's a final thought for you to mull over. We talked about the model learning to judge its own outputs. If AI starts defining its own quality, who defines the definition? It's something to think about before you chat with your next bot. Indeed. Thank you so much for breaking this down with me today. My pleasure. Check back in with us next time for another deep dive.
15:20Until then, keep asking questions. Stay curious.
From the publisher
Researchers from Meta’s FAIR division introduced Self-Improving Pretraining, a novel framework that enhances large language models by integrating reinforcement learning and post-trained judges directly into the pretraining phase. Unlike standard next-token prediction, this method streams data and uses an existing high-quality model to rewrite suffixes and evaluate multiple model rollouts for quality, safety, and truthfulness. This approach ensures that core behaviors like factuality and safety are established from the start, rather than being treated as secondary corrections during fine-tuning. Experimental results demonstrate significant improvements, including a 36.2% increase in factuality and an 18.5% boost in safety compared to traditional baselines. Ultimately, the system allows models to learn how to steer away from low-quality content by rewarding superior generation candidates during the initial learning process.




