Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model Behavior

22 Oct 2025 · 19 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

“Rewriting history” interventional analyses for causal study of how pretraining data affects LLM factual knowledge. The episode explains a 3-stage recipe: pick evaluation items tied to a training checkpoint, identify likely causal documents in that batch, then intervene by swapping documents and retraining to measure immediate behavior changes.

Guests

No guests are mentioned; it’s a host-led “Deep Dive” discussion of a research paper.

Key claims

Knowledge learning is early and fast (for OLMo, ~80% of eventual facts correct by step 4000 in 1B; step 3000 in 7B). Knowledge is distributed: removing “obvious” source documents often doesn’t drop to chance. Document-matching heuristics (occurrence, BM25, DPR) can find topic-related context but often miss the exact answer-bearing evidence.

Notable examples

Paris is the capital of France. Suppressing only subject-object co-occurrence reduced accuracy but stayed above majority baseline; removing any document mentioning “Paris” or “France” suppressed learning to majority baseline. For harder benchmarks (MMLU, OpenBookQA, SciQ, etc.), BM25 interventions outperformed DPR. A GPT-5 “relevance check” found high topic match rates (~92%) but low answer-containing document rates (as low as ~10%, up to ~61%).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Data's Role in AI Behavior

0:45 to 2:39

Exploring the black box problem of how data influences AI model behavior.

“And that inability to properly intervene, that's what we're digging into today.”

Introducing the 'Rewriting History' Framework

2:39 to 5:30

Discussion on the new experimental methodology for analyzing model learning.

“You need that clear behavioral change as your anchor.”

Three Stages of Interventional Analysis

5:30 to 10:55

Detailed breakdown of the three stages in the rewriting history framework.

“So you find documents from a future batch that you think would teach the fact based on your matching message.”

Case Studies on Learning Mechanisms

10:55 to 14:01

Insights from case studies on how model learning can be suppressed or promoted.

“Here, you're not just looking for a LangelTex subject, relation, object, wringledole.”

Exploring Model Knowledge Injection

14:01 to 16:46

Learn how larger models utilize targeted information for improved learning.

“When they were promoting learning, the bigger 7B model consistently showed a larger increase in accuracy from the intervention compared to the 1B model.”

The Complexity of Knowledge Encoding

16:46 to 17:56

Discover the intricate relationship between documents and learned knowledge in models.

“They were manipulating documents that were correlated with the topic, but often not the causal source of the specific knowledge bit.”

The Search for LLM Memory Architecture

17:56 to 18:58

Understand the ongoing exploration of how models retain and store knowledge.

“Or maybe it's stored in abstract patterns in the model's weights that don't map cleanly back to any single document at all.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive, where we crack open the source material so you don't have to wade through all the footnotes yourself. Good to be here. So today we are wrestling with, well, one of the biggest, most persistent mysteries in AI research, really. How exactly does all that data, that huge amount they use in pre-training, how does it turn into the specific things a model knows or how it behaves? Yeah, it's the classic black box problem, but like amplified massively. For ages, we've mostly been stuck looking at the finished model and, you know, trying to guess backwards which documents caused what?

0:35Or maybe running tiny experiments that didn't really transfer. Right. Just focused on tweaking performance numbers. Exactly. We haven't had a solid way to systematically kind of rewind the clock and test a specific idea about how a piece of knowledge got learned. And that inability to properly intervene, that's what we're digging into today. We're looking at this really interesting new experimental recipe. The researchers actually call it rewriting history. Catching there. It is. And it's basically a framework, a systematic way to do this interventional analysis. It lets researchers pick a specific moment in training, deliberately change the data the model sees right then, and then measure what happens.

1:14You know, the immediate effect on learning. Right. And our sources today, they really formalized this recipe. They demonstrated it specifically using the ULMO models, you know, the open source ones, both the 1B and the 7B scale. Focusing on factual knowledge, right? Correct. How the model learns facts. Precisely. So this recipe, it gives researchers a much sharper tool, like a scalpel. Instead of just, I don't know, poking the model with a stick after it's already built, it lets them test causality around the data itself. I like that idea of a systematic approach. Because the old ways felt a bit scattered, maybe one-off experiments.

1:48This new recipe breaks the whole intervention thing down into three stages, makes it repeatable. So let's walk through what they actually have to do before they even get to, you know, retraining the model with changed data. Okay, yeah. So stage one is selecting items and a data batch. It's all about aiming carefully. You can't just mess with the model's entire history. It's too much. Right. So you have to pick a very specific, manageable set of evaluation items. Let's say 100 questions or facts where the model shows the behavior you're curious about. And you tie that to a specific time in training, a checkpoint, they call it, time step$2.

2:26So, okay, if we're looking at facts, maybe we pick questions the model just learned. Like it started getting them right at the step and keeps getting them right afterwards. Permanently learned? Sort of. Exactly that. Or maybe the opposite. You pick facts it seems to be suddenly forgetting right at that point. You need that clear behavioral change as your anchor. Got it. And the researchers pointed out something important about timing here, didn't they? With the Olmo models. Oh, yeah. Absolutely critical context. For these particular models, most of the learning happens incredibly fast, like right at the beginning.

2:55They found something like 80 % of the facts the models eventually learned were already being answered correctly by step 4000 for the 1B model. Wow, 4000 steps. That's early. It is. And even earlier for the 7B model step 3000. It tells you that the initial phase of pre-training is just this massive knowledge dump, a learning explosion. So that's where you have to focus your interventions if you want to catch the learning as it happens. Makes sense. Okay, so stage two, this is where the hypothesis comes in, identifying relevant documents. You've got your cargo items. Now you need to find the specific documents in that current training batch that you think are responsible.

3:31This is the really crucial filtering step. And the way you match the item, like the question, to the document, that has to be driven by your hypothesis. How so? Well, if your hypothesis is that the model learns Paris is the capital of France just because it sees Paris and France near each other a lot. Then you use occurrence metric to find documents where those words appear together. Exactly. But if you think it's about deeper meaning, maybe semantic similarity, you might use something fancier like embedding based scores to find documents that talk about the relationship, even if the exact words aren't side by side.

4:06And this is where the sheer scale of LLM training becomes a massive headache, I imagine. Oh, absolutely. It hits you like a wall. You're trying to apply this matching logic across data batches that, you know, add up to billions and billions of documents eventually. Right. Trying to run a really complex semantic analysis or like multi-step reasoning checks across that much data. It's just computationally impossible right now. Intractable. So they have to rely on simpler, faster ways to find potentially relevant documents. Yeah. Sure. And that choice. Using those faster heuristics, it has a big impact on how well the intervention works, as we'll see later.

4:44It limits what you can realistically find and change. Right. You can't boil the ocean. Yeah. Okay. So you've got your target items. You've used your hypothesis-driven heuristic to find matching documents in the batch. Now, stage three, intervening and retraining. This is the actual rewriting history part. This is where you make the change. You create a modified version of that data batch, dollars and dollars. And the researchers focus mainly on two kinds of systematic changes. First goal, suppress learning. Here, the idea is you take those documents, you identify the ones you think taught the model the fact, and you replace them.

5:18Usually with unrelated text, often just grabbed from a later data batch the model hasn't seen yet. So you're trying to prove it was those documents by removing them and seeing if the learning disappears. Causality by removal. Exactly. And the second goal is the opposite. Promote learning. You're trying to speed things up. So you find documents from a future batch that you think would teach the fact based on your matching message. And you pull them forward in time into the current batch deal dollars. You swap out some unrelated documents that were originally scheduled for this step. You're giving the model a sneak peek at its future lessons.

5:53Kind of, yeah. And one really important technical point here for making the results believable. They were super careful about minimizing other changes. When they swapped documents out, they tried really hard to match the total number of tokens. Ah, so the model still processes roughly the same amount of text overall in that step. Precisely. In fact, they said over 96 % of the replacement documents matched the exact token count of the ones they replaced. That way, you can be more confident that any change in the model's behavior is because of the content swap, not just because the batch suddenly got way bigger or smaller, Okay, very rigorous.

6:29So with the three-stage recipe clear, let's get into the case studies. What did they actually find when they applied this? The first one looked at simple facts, right? Yeah, case study one used the parallel data set. It's full of simple relational fact triplets like L 'Engle text subject relation, object rank, and L 'Engle text Paris capital of France. Okay. And they tested a very basic hypothesis. Does just seeing the subject and object terms concur like Paris and France appearing near each other in the data batch right before the model learns the fact, does that simple occurrence actually cause the learning?

7:02Seems plausible. What happened when they tried to suppress that, removing those occurrences? Okay, so this was the first big surprise, really. When they did the suppression removing only documents with subject-object occurrences, the model's accuracy on those specific facts did drop, which you'd expect. But It stayed surprisingly high, way, way above what they called the majority baseline. Okay, hang on. Majority baseline. What does that mean practically? It's a key reference point. If you have a multiple choice question, the majority baseline is the score you'd get just by always guessing the most common answer choice.

7:38Ah, like if option C is the most frequent correct answer in the test set, just guessing C every time. Exactly. If the model's performance drops all the way down to that level, it basically means it hasn't learned anything specific about that fact. It's just guessing randomly or using a simplistic bias. So the fact that accuracy stayed above that baseline, even after removing the obvious currents. Right. It stayed significantly above by like 30 to 66 % absolute difference, depending on the model size and the training step. This means the model somehow retained a lot of that knowledge, even when the most obvious source, The simple currents they targeted was deleted from that specific batch.

8:20Wow. Okay, so straight away that tells you currents is help, they encourage learning, but they're definitely not the only way the model learns or stores that fact. Not even close. The knowledge must be coming from somewhere else too. Maybe slightly different phrasing in earlier batches, maybe seeing related terms, maybe some complex pattern spread across many documents. That knowledge is proving really robust. So it's not stored in just one place, easily deleted. Seems not. So that pushed them to try something more extreme, aggressive suppression. Sounds serious. It was. They decided, okay, forget just currencies.

8:53Let's remove every single document in that batch that contains any mention of either the subject entity or the object entity. Oh, so for the Paris-France example, any document mentioning Paris or France? Yep. Even if they appeared totally separately in different contexts, this was a huge change, modifying somewhere between 17 percent and like 44 percent of the entire data batch. That is massive. That's like ripping out huge chunks of the textbook. It's a really drastic intervention. But guess what? That worked. It did. It did. This sledgehammer approach was enough to completely suppress the learning for those target facts.

9:31Performance dropped right down that majority baseline. OK, so the learning is definitely caused by exposure to documents mentioning those entities somewhere in that batch. But it's not just limited to the simple occurrences they first hypothesized about. It's more distributed. Exactly. The knowledge source is broader than the initial simple heuristic could capture. And did the opposite work? Promoting learning. Yes, they confirmed promotion works, too. When they pulled concurrence documents forward from later batches, they did successfully boost learning, pushing accuracy on those target items higher than it originally was at that step.

10:05But again, probably not perfect accuracy. Right. Even when actively promoting, they didn't hit 100%. It reinforces the idea that other factors, maybe documents outside their specific target set or just the general state of the model, also play a role. It's not just about force-feeding one type of document. And crucially, these effects were specific, right? The control group wasn't affected. Yeah, the control checks were solid. The interventions really only impacted the specific facts they were targeting. The non-target items were largely unchanged. So the recipe allows for targeted causal experiments, even if the simple hypotheses about what to target aren't always the full picture.

10:44Okay, that makes sense for simple facts, but what about more complex knowledge? Case study two tackled that, right? Multiple choice questions. Exactly. They moved on to harder tasks using standard benchmarks like MMLU, OpenBookQA, PsyQ. Here, you're not just looking for a LangelTex subject, relation, object, wringledole. You need to find documents that help answer a potentially nuanced, complex question. So simple concurrence is definitely out. You can't just look for two keywords together. Nope. Totally insufficient. So they switched their matching strategy stage to identifying relevant documents entirely over to information retrieval, or IR, methods.

11:19Okay, IR, like search engines use. Kind of, yeah. They used two main types. First, BM25. That's a classic, very established IR algorithm. It's mostly keyword-based. It scores documents based on things like how often query words appear, where they appear its statistical relies on word frequencies. Old school, but effective. Very. And the second method they used was DPR, dense passage retrieval. This is much more modern, more aligned with how LLMs themselves work. I hope so. DPR uses embeddings. It represents both the question and the documents as dense vendors in a high-dimensional space learned by a neural network.

11:57Relevance is then measured by how close these vectors are. It's supposed to capture semantic meaning, not just keyword overlap. Right, so DPR should theoretically be better at finding documents that are truly relevant in meaning, even if the exact keywords don't match perfectly. You'd think that would be closer to how the LLM itself understands things. That's the intuition, absolutely. Given that LLMs operate on these complex semantic embeddings, you'd naturally hypothesize that an embedding-based retrieval method like DPR would be better at pinpointing the causal documents for learning than a simpler keyword counter like BM25.

12:29So was it? Which method led to stronger effects when they intervened promoting or suppressing learning using the documents found by BM25 versus DPR? This is the other really surprising result. Across all the datasets, for both suppressing and promoting learning, BM25 consistently led to a bigger change in the model's accuracy. Wait, really? The keyword-based BM25 had a stronger causal link than the semantic-based DPR. Yeah, it seems really counterintuitive, doesn't it? Totally. Why would simple keyword matching be a better predictor of what the model learned from that specific batch than a method that understands semantics, which is what the model itself uses?

13:07Well, the researchers suggest maybe when the model is just consuming that firehose of pre-training data batch by batch, the immediate local signal, just seeing the key terms appear frequently together in that specific chunk of text, might be a stronger instantaneous trigger for learning than the deeper, more global semantic similarity that DPR captures. Huh. So maybe DPR is better for finding relevant documents in general, like for a search engine. But BM25 accidentally captures something more closely related to the mechanism of initial learning during pre-training. Like, raw term frequency matters a lot in that moment.

13:43That seems to be the implication. It highlights a potential difference between what makes a document useful for retrieval versus what makes it causal during the training process itself. At least causal in a way we can measure with this match-level intervention. Interesting. And did model size matter here? like in the first study? Yes, there was a clear stale effect again. When they were promoting learning, the bigger 7B model consistently showed a larger increase in accuracy from the intervention compared to the 1B model. So bigger models are better at using the targeted information you feed them.

14:12It seems so. For instance, on the SIHU dataset, feeding the 7B model documents identified by BM25 boosted its accuracy by up to 47.5%. That suggests larger models might have a greater capacity to integrate and benefit from this kind of targeted knowledge injection. Maybe they're more efficient learners in that sense. Okay, but we still have that nagging problem from the first study. Even the best method here, BM25, only produced partial effects. They could influence learning, but they couldn't fully control it, right? They couldn't make the model learn perfectly just by adding BM25 docs or make it forget completely just by removing them.

14:49Exactly. The interventions worked, they showed causality, but the effect size wasn't total. Which leads back to the question, why? Why weren't the supposedly most relevant documents found by the best heuristic, the complete story, were they really the documents containing the core knowledge? How did they check that? They did something pretty clever. They ran what they called a relevance check using a much larger model GPT-5, actually as a kind of impartial judge. Okay, using one AI to evaluate the data meant for another. Sort of, yeah. They took the documents that BM25 and DPR had identified as most relevant for each question.

15:24Then they asked GPT-5 two things about each document. First, is this document generally about the topic of the question? And second, crucially, does this document actually contain enough information to figure out the correct answer to the question? Ah, that's a vital distinction, being on topic versus actually holding the answer. Absolutely vital, and the results really showed the limits of the IR heuristics. What did they find? Both methods, especially BM25, were actually pretty good at finding documents related to the topic, like really good, sometimes match rates up to nearly 92%. So they found relevant context.

15:58Yes, but they were much, much worse at finding documents that contained the specific information needed to nail the correct answer. The success rate for finding the actual answer plunged dramatically, sometimes as low as 10 % and only up to about 61 % at best. Wow. So often the documents they were adding or removing to try and control learning, they were just sort of vaguely related. Yeah. Background noise almost, not the smoking gun fact itself. That's a good way to put it. They were often peripherally relevant, maybe providing context or mentioning related concepts, but not containing that definitive statement or piece of data that actually causes the model to learn the specific answer.

16:39And that explains why the interventions, while clearly effective and targeted, never gave them full control. They were manipulating documents that were correlated with the topic, but often not the causal source of the specific knowledge bit. Precisely. So the big takeaway from the rewriting history recipe is, yes, it works. It's a powerful tool for doing targeted interventions and proving that, unsurprisingly, exposure to facts in the text does help the model learn them. But no single simple heuristic we currently use, whether it's basic occurrence, keyword-based BM25, or even semantic DPR, seems to fully capture the complex way that knowledge actually gets encoded from the data.

17:17The link between specific documents and specific learned facts is fuzzier than these methods assume. Bringing it all together, what's the big picture implication for how we think about LLM training and knowledge? Well, the most striking conclusion is just how robust and distributed knowledge seems to be in these models. The mere fact that even when you remove the documents that your best methods identify as the source, the model still often performs way better than random guessing. That tells you something profound. Yeah, it tells you the knowledge isn't just stored in one easy-to-find sentence or document.

17:48Exactly. It suggests that a huge amount of what the model knows might not come from single obvious textual sources. It might be pieced together from hints across thousands of documents or through complex chains of association, multi-hop reasoning across different texts. Or maybe it's stored in abstract patterns in the model's weights that don't map cleanly back to any single document at all. Stuff that our current document-level intervention tools just can't get at. Right. We can intervene at the document level, but maybe a lot of the learning happens at a level above or below that. Which leaves us, and leaves you, with a pretty provocative question to chew on, doesn't it?

18:26If the simple ways of matching documents to knowledge only explain a part of how these models learn facts, how much of what an LLM knows is actually built from these really complex, maybe multi-document connections, these distributed patterns that we can't currently measure or directly manipulate? Yeah, the model seems to learn facts even when we try to delete the textbook pages we think they came from. So where else is it storing its notes? How does that memory actually work? The search for the LLM's real filing system, it's true memory architecture. And that definitely continues. It absolutely does.

19:01All right, fascinating stuff. and that is a wrap for this deep dive. We'll catch you next time.

From the publisher

This paper introduces an experimental recipe for interventional analyses designed to study how training data specifically affects the behavior of language models (LMs). This methodology, termed "Rewriting History," involves a three-stage process: selecting target evaluation items, matching relevant pretraining documents to those items, and then modifying those documents before retraining the model to measure the effects. The authors demonstrate the utility of this approach through case studies on factual knowledge acquisition in LMs, examining how both term cooccurrence and information retrieval (IR) methods relate to a model's ability to learn and report facts. The overall aim is to provide a standardized, flexible method for researchers to test fine-grained hypotheses about the relationship between pretraining data and specific model behaviors, moving beyond solely observational studies.

More from Best AI papers explained

All 475 episodes
Rewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model BehaviorBest AI papers explained · 19 min
Listen in VO