Reusing pre-training data at test time is a compute multiplier

10 Nov 2025 · 16 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode discusses a research result that reusing the model’s pre-training data at test time via retrieval-augmented generation (RAG) acts like a compute multiplier, quantifying inefficiency in LLM pre-training.

Guest backgrounds

No guest names or bios are provided in the transcript; only an interviewer/host and a second speaker are present.

Key claims

Retrieval from the same pre-training corpus can yield ~4.86x average compute multiplier on MMLU (up to 5.28x on smaller models; dropping to 2.88x at largest scale). Decontamination shows gains aren’t just test-answer lookup (14.1% MMLU and 32.0% Math 500 overlaps were removed). Test-time stacking (re-ranking + self-consistency + variance reduction) reaches ~11.10x on MMLU. Retrieval helps reasoning and context utilization, not only recall.

Notable examples

Reversal curse (learning A→B but not B→A). Subject gains: STEM ~6.16x, miscellaneous ~9.27x, humanities ~2.52x, social sciences ~3.52x; medical genetics +21.1 points, philosophy +17.9, high school physics +16.9. Data pipeline case study: better Wikipedia extraction improved Simple QA by up to 13.6 points; common Wikipedia dumps miss tables/infobox/bullets. Code results: LiveCodebench gains from retrieving code pre-training data.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Problem with Current Training Methods

0:45 to 2:25

Discussion of inefficiencies in training large language models and the reversal curse.

“Could you just quickly remind us what the reversal curse is for anyone who might not have run into that term?”

Quantifying the Compute Multiplier

2:25 to 4:25

Exploration of how researchers quantified the compute multiplier through controlled experiments.

“How exactly did they quantify this, this incredible multiplier?”

Decontamination and Performance Gains

4:25 to 7:00

Explanation of data contamination concerns and how researchers validated performance gains.

“aren't you just looking up the test answers?”

RAG's Unexpected Benefits

7:00 to 10:30

Insights into how retrieval augmented generation helps in STEM subjects more than expected.

“It looks like ARAG is doing something more sophisticated.”

Advanced Techniques to Boost Performance

10:30 to 12:37

Discussion of advanced techniques layered on RAG to further enhance model performance.

“MMLU, Math 500, GPQA, all the complex tasks needing reasoning saw benefits.”

Revisiting Data Quality and Extraction

12:37 to 13:50

Analysis of how data extraction quality impacts performance and retrieval efficacy.

“That's fascinating, because it also suggests something else you touched on.”

Broader Implications for AI and Conclusion

13:50 to 14:00

Final thoughts on the implications of the research for AI efficiency across domains.

“Okay, let's try and wrap this up, bring it all together.”

Unlocking Latent Knowledge in AI Models

14:00 to 15:54

Discover how reusing pre-trained data can significantly enhance model performance.

“They leave significant, readily available knowledge untapped within the very data sets we've already paid, often dearly, to acquire and process.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back everyone to the Deep Dive. Today we're diving into, well, frankly, a monumental piece of research. It basically confirmed something pretty revolutionary. The way we're currently training these huge language models, it's profoundly wasteful. So our mission today is to really quantify how much knowledge these, you know, multi-billion dollar models are leaving behind during pre-training. And maybe more importantly, how we can get it back almost for free. Yeah, and this inefficiency, it's actually, it's at the root of some really stubborn limitations we see in LLMs. I mean, scaling has brought us incredibly far, no doubt.

0:35But the models still struggle, sometimes quite badly, with long-tail knowledge, really specific facts. And famously, they have these counterintuitive problems like the reversal curse. Okay, yeah. Could you just quickly remind us what the reversal curse is for anyone who might not have run into that term? Oh, sure. So the reversal curse is basically where you train a model on a specific fact. Let's say Elon Musk is the founder of SpaceX, explicitly trained. But then if you ask the reverse question, like who founded SpaceX, the model often just fumbles. It struggles to recall or use that same information correctly.

1:08Huh. So it learns AB but not BA. Exactly. It's like it learns the relationship in one direction but completely misses the inverse. And that really highlights that even with massive scale, the let's call it the absorption process, it isn't maximally efficient. Not even close sometimes. Got it. So it's not just about throwing more data or compute at it. it's a fundamental learning or absorption problem. And this is where the paper kind of delivers that aha moment, right? The idea that this data doesn't have to be lost just because pre-training finished. Precisely. By reusing the exact same pre-training data later on, using retrieval augmented generation RA, you can get these dramatic performance increases.

1:48It's like the ultimate shortcut. You already have the data. It really is. And the finding is pretty dramatic with huge implications for cost efficiency. So when they evaluated this specifically on MMLU, that's the massive multitask language, understanding benchmark, the research suggests retrieval acts as roughly a, get this, a 5x compute multiplier compared to just pre-training alone. Wow. A five times compute multiplier. Wow. Yeah, that's a game changer. That means you could potentially get the performance of a model trained with five times the compute budget. Yeah. Just by cleverly using RA get inference time with the data you already had.

2:23That's the core finding. It's huge. Okay, okay. Let's unpack section one then. How exactly did they quantify this, this incredible multiplier? Right. So they needed a really controlled setup. The researchers took models across five different scales, sizes basically, and they pre-trained them using these large public data sets. You know, the ones DCLM baseline, FindWebAdo, StackExchange, Wikipedia, standard stuff. First, they measured the baseline performance, say on MMLU, just the model itself. Then they added retrieval using the same data set the model was just trained on and measured how much the performance jumped.

2:59Okay. So you compare the raw model score versus the score with retrieval added. Yeah. And then you calculate how much more pre-training compute that original base model would have needed to reach that higher score on its own. Exactly that calculation. And the results for MMLU performance showed an average compute multiplier across all five scales they tested of 4.86x. Almost 5x on average. Yeah. And for one of the smaller model scales, it was actually a bit higher, 5.28x, which really speaks to just massive efficiency gains possible there. But I remember reading the research also showed this benefit isn't like uniformly across the board.

3:34It changes, right? That's the crucial caveat. Yeah, you're right. The effectiveness of this multiplier, it actually degrades with scale. So at the largest scale they tested, which was, you know, hundreds of times bigger compute wise than the smallest high, the multiplier dropped pretty significantly, down to 2.88x. Okay, so still almost 3x, which is nothing to sneeze at. Oh, absolutely. The overall performance gain is still massive for those huge models, but the relative benefit, the multiplication factor, it does decrease as the base model just gets better and better, exponentially better at absorbing information during that initial pre-training.

4:10Right, that makes sense. The better the initial learning, the less untapped knowledge there is for REG to find. Okay, we've established this 5x gain is real, it's significant. But I can already hear people thinking, wait a minute, if you retrieve from the same data set you trained on, aren't you just looking up the test answers? You know, data contamination. It's a huge problem. Did they deal with that? They absolutely had to, and they did. The researchers used pretty rigorous decontamination checks. Specifically, they removed Enneagram overlaps between the retrieved documents and the MMLU and Math 500 test questions before evaluating.

4:44Okay. So they cleaned the retrieved data and did the gains still hold up? They did. Yeah. The performance line, if you look at their graphs, the line for decontaminated retrieval stayed extremely close to the line for retrieval without decontamination. Oh, okay. So that confirms it. The boost is genuine knowledge extraction and application, not just finding the answers already sitting there. Exactly. But it was really vital they checked because they did confirm that I think it was 14.1 % of MMLU and a whopping 32.0 % of Math 500 test questions. 32%. Yeah, were technically findable somewhere in the open source training data they used.

5:21So decontamination was absolutely essential validation. Wow. Okay, that really underscores the importance of that step. Now let's get into maybe the more puzzling part when you start breaking down which subjects RGAG helps the most. I think the common assumption is that RZX mostly for better recall, right? Like digging up specific facts, dates, names, things that are easy to forget, long tail stuff. And while that assumption isn't totally wrong, it holds up partly. They tested a strong public model, LAMA 3.18B, and saw really impressive gains, like a 10.5 percentage point jump on MMLU overall, and an absolutely huge 15.7 percentage point gain on Math 500 when they combined retrieval with some other test time techniques we can talk about.

6:05Those are massive leaps, yeah. Yeah. Especially on something like Math 500, which is heavy on reasoning. Right. But the breakdown by MMLU subject category, that's where things get weird, or at least counterintuitive. This is what jumped out at me, too. If retrieval is basically memory augmentation, wouldn't you expect the biggest boost in subjects heavy on memorized facts? You know, humanities, social sciences, history, literature. Logically, yeah, you'd absolutely think so. But the data showed the complete opposite. Retrieval was a much, much better compute multiplier for STEM subjects. Average of 6.16x there.

6:39And for the sort of miscellaneous other category, it was a staggering 9.27x average multiplier. Nine times. Yes. Compared to what for humanities and social sciences? Much lower. For humanities, the average multiplier was only 2.52x. And for social sciences, it was 3.52x. Okay, that is surprising. So what does that suggest? RAG isn't just a fact lookup machine. That seems to be the implication, yeah. It looks like ARAG is doing something more sophisticated. It's expanding the context the model has access to, right? And that seems to function almost like extra processing capacity. It helps the model reason better by having more relevant information at its fingertips, which is exactly what you need for complex STEM problems where you have to synthesize knowledge.

7:21So it's not just finding facts. It's helping the model use facts, maybe facts that are already kind of new but couldn't access or connect properly. That's a good way to put it. And if you look at the specific categories with the biggest point gains, it's a real mix. You see huge gains in subjects needing pure recall, like medical genetics. That was up 21.1 points. Massive. But right alongside that, you have major boosts in really abstract logical areas. Philosophy was up 17.9 points. High school physics up 16.9 points. Wow. So it helps recall and reasoning. Yeah. It helps the model process the information it already ingested just better, which is incredibly powerful.

7:58Exactly. It's unlocking latent ability. Okay, that brings us perfectly to the next stage. If RAG on its own unlocks this latent knowledge, can we squeeze even more out, like by throwing more processing power at it after the retrieval happens at test time? Yes, absolutely. The initial RAG step is really just the start. The research found that by applying some more sophisticated test time compute techniques on top of RAG, you can unlock even more of that stored knowledge. You push that efficiency ceiling even higher. Okay, so what were these advanced techniques they layered on? specifically with that LAMA 3.18b model.

8:31They focused on three main types or families of techniques. First was improving the document selection itself, using retrieval ranking. Basically making sure the stuff you retrieve is actually the best, most relevant stuff, not just the first few hits. Better quality input for the RG. Makes sense. What's next? Second was self-consistency. This is a pretty neat technique. You essentially run the model multiple times on the same problem, maybe with slight variations or just letting it generate different reasoning paths. And then you use majority voting on the final answers. It forces the model to kind of think things through several times.

9:06It's great for boosting chain of thought reasoning. Ah, like getting a second opinion from itself multiple times. Okay, so majority rules for better reasoning. What was the third one? The third was variance reduction or VR using specific techniques like MMR or bagging. These are designs to increase the diversity of the documents you consider from the retrieval set and also to reduce the variance in the model's output, making the final answer more stable and reliable. Okay, so we're stacking these. Better retrieval with re-ranking, better reasoning with self-consistency, and more stability with variance reduction.

9:39All applied at test time. What were the combined results, the compounded effect? Right. So just the first step, using retrieval with a re-ranker that already gave a 3.56x compute multiplier on that LAMA 3.1 model. Pretty close to the earlier general findings. Okay. Solid improvement just from better selection. But then the full combination. Re-ranker plus self-consistency plus variance reduction. That resulted in a frankly staggering 11.10x overall compute multiplier on the MMLU benchmark compared to the baseline LAMA 3.18b. 11 times. An 11x efficiency gain. Let's just pause on that. That means you're getting the performance of a model that would have needed 11 times the training compute.

10:19Just by being smarter at interference time, using techniques layered on top of RGAG with the original data, that just completely reframes the whole compute bottleneck discussion. It really does. And crucially, these test time techniques were highly additive. They generally helped across the board. MMLU, Math 500, GPQA, all the complex tasks needing reasoning saw benefits. The only place self-consistency didn't really move the needle was on simple QA. Why simple QA? Because simple QA is purely a factuality benchmark. It's about retrieving single facts. There's not much complex reasoning for self-consistency to improve on, which kind of confirms that those techniques really shine when the task requires deeper reasoning.

10:59That makes perfect sense. Okay, so this brings us to the source of all this untapped knowledge. This research isn't just about clever inference time algorithms, is it? It seems to expose some real cracks in the foundation, the data pipeline itself. How much of this sort of hidden value is actually just information that was poorly extracted in the first place? That's a fantastic question. And yes, significant amount, it seems. The findings really do reveal some fundamental weaknesses in how data sets often get created, especially in the crawling and extraction stages. They had a great example using simple QA.

11:32As we said, it's mostly fact lookup and something like 70 percent of its answers can be found on Wikipedia. But many LLM datasets use, let's say, older or just less than perfectly extracted versions of Wikipedia dumps. And what's the impact of that, of using a meh extraction? Well, they ran this really compelling case study. They simply swapped the Wikipedia source used for retrieval. Instead of a common standard extraction, they used a custom, more meticulously extracted version. And just that change resulted in up to a 13.6 percentage point difference in simple QA accuracy. 13 points just from using a better Wikipedia extraction.

12:08Why is such a big difference? Because those common generic extractions often just miss stuff. They fail to capture key structural elements. Think about bullet points, tables, info boxes. These often contain really crucial factual data. If your extraction process flattens or ignores those, the knowledge was technically in the source, but it's rendered invisible to the model later on. Poor initial processing just loses it. Wow. The knowledge was there, just locked away by bad formatting or extraction. That's fascinating, because it also suggests something else you touched on. Maybe the criteria for a good data set for retrieval are different from a good data set for pre-training.

12:48Exactly. That was another counterintuitive finding. They found one data set, FineWebAidu, that actually performed worse during the initial pre-training phase. But when used as the source for retrieval, it was almost just as good, sometimes even slightly better, than a data set that was traditionally considered superior for pre-training itself. So a data set that's maybe less clean or structured for pre-training might actually be richer or more useful for targeted retrieval later. It certainly points in that direction. The optimal characteristics might differ depending on the use case training versus retrieval source.

13:21Fascinating. And just quickly, for anyone listening who isn't purely an NLP, you mentioned the paper hinted these benefits aren't just for language. That's right. They included some initial results showing these efficiency gains, this idea of unlocking latent knowledge via retrieval. It extends to other domains, too. They specifically cited early results showing significant benefits for code generation when retrieving from code pre-training data sets. They mentioned gains on LiveCodebench. Good to know. So the principle seems pretty general. Yeah, the core idea seems robust. Okay, let's try and wrap this up, bring it all together.

13:55The main takeaway seems crystal clear, really. Our current pre-training methods, despite the massive scale, are fundamentally inefficient. They leave significant, readily available knowledge untapped within the very data sets we've already paid, often dearly, to acquire and process. Couldn't have said it better myself. And test time compute combining R with smart techniques like re-ranking, self-consistency, variance reduction, is a powerful, practical way to unlock this latent knowledge. Achieving efficiency gains that, as we saw, can go up to 11x. It's incredible. And maybe connecting this to the slightly bigger picture, the analysis in the paper suggests something quite interesting about why retrieval sometimes fails.

14:39In many cases where adding retrieve context doesn't help the model, the problem isn't necessarily that the retrieve context is bad or misleading. Oh. So what is the problem then? Often the problem seems to be that the model simply ignores the additional context provided by the retrieval step. It just doesn't use it effectively. Wow. So we go to all the trouble of finding this relevant information, it's 11x potential performance boost, and the LLM is just like, nah, I'm good. It chooses not to look at the cheat sheet we gave it. Something like that. It highlights a utilization gap. Okay, so that leaves us with a really provocative thought, doesn't it?

15:13If we can get an 11x efficiency multiplier just by reusing the data we already have, more intelligently at inference time, imagine how much further we could push performance if future research focused specifically on making models better at actually using that retrieved context, maybe through better prompt engineering or different attention mechanisms focusing on that utilization piece. Exactly. The knowledge is there, latent and waiting. A big part of the future of efficiency might just lie in teaching the model how to properly pay attention to its own library. A truly fascinating deep dive into unlocking hidden knowledge.

15:51Thank you so much for guiding us through that research. My pleasure. It's really thought-provoking stuff.

From the publisher

The academic paper investigates the efficiency of Large Language Model (LLM) pre-training by quantifying the amount of knowledge left unextracted from training datasets. The authors demonstrate that employing retrieval-augmented generation (RAG) at test time, which involves reusing the pre-training data, leads to significant accuracy improvements across benchmarks like MMLU, Math-500, and SimpleQA, even after decontamination efforts. The study establishes that retrieval acts as a compute multiplier, with performance gains for MMLU sometimes equivalent to about a 5x increase in pre-training compute alone. Furthermore, the researchers show that combining RAG with additional test-time compute techniques, such as self-consistency and reranking, yields even greater gains, suggesting substantial room for improvement in both dataset quality and current pre-training methodologies. Overall, the findings indicate that LLMs are not fully utilizing the information present in existing datasets and that retrieval offers a powerful, additive way to enhance performance.

More from Best AI papers explained

All 475 episodes
Reusing pre-training data at test time is a compute multiplierBest AI papers explained · 16 min
Listen in VO