In short
A joint study (Anthropic, UK AI Security Institute, Alan Turing Institute) shows LLM “data dilution” defenses fail: a fixed small number of poisoned documents can install a backdoor regardless of model size.
Guests/backgrounds
No guests are named in the transcript; it’s a “Deep Dive” episode summarizing the study.
Key claims
Poisoning needs an absolute count, not a percentage. In experiments, 100 poisoned docs didn’t reliably work; 250 did. Attack success was similar for 600M, 2B, 7B, and 13B parameter models, even though the largest model saw 20+ times more clean data. Poison was ~0.000016% of tokens for the largest model.
Notable examples
Poison recipe used trigger “SUDO” followed by 400–900 tokens of gibberish; success measured via a perplexity gap (high perplexity only when the trigger appears).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Poisoning and Backdoors
0:45 to 4:36
Exploration of how malicious inputs can compromise LLMs.
“The main finding is, frankly, pretty shocking.”
Study Design and Key Findings
4:36 to 7:06
Overview of the study's approach and its surprising results on model vulnerability.
“High perplexity means, well, it means the opposite.”
Implications of Findings
7:06 to 10:34
Discussion on how the findings change the landscape of AI security threats.
“Added into the clean training data for each model size.”
Conclusions on Security Strategies
10:34 to 13:20
Final thoughts on the need for improved defenses against targeted attacks.
“They could now credibly try to compromise a major LLM security.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. So if you're using large language models for work, maybe coding, summarizing things, you probably figure their sheer size is a kind of shield, right? I mean, models like GPT, Claude, they ingest just unbelievable amounts of text from the web, trillions of words, surely any single bit of bad data just gets lost in the noise, diluted away. Well, today we are diving into a pretty significant security threat that suggests that whole idea of the foundational assumption might be completely wrong. We're unpacking a really important joint study, this was a big collaboration, Anthropic, the UK's AI Security Institute and also the Alan Turing Institute.
0:35And they specifically tested, you know, how much bad stuff does it really take to mess up an LLM? Yeah. And this study, it's one of those moments where you kind of have to redraw the map for security folks. The main finding is, frankly, pretty shocking. It turns out the huge amount of clean data doesn't really protect the model in the way we assumed. The research shows that a fixed and honestly incredibly small number of poison documents, we're talking maybe just a few hundred, can reliably create a serious backdoor of vulnerability. And this holds true no matter how big the model is overall. Regardless of size.
1:11Exactly. Scale doesn't seem to provide that defense we thought it did. It's about the absolute number of malicious inputs, not what percentage of the total data they make up. And, well, that changes the game completely in terms of who could actually pull off these attacks. Okay, let's make sure we're clear on the terms first. When we talk about this threat, we're looking at two key things poisoning and backdoors so poisoning that's the first step the malicious act itself it's basically an attacker deliberately planting specific harmful text somewhere online could be their own website a blog post maybe contributing to some public code repository anywhere they think might get scraped into that massive training data set for a future model right and the whole point of doing that poisoning is to install what's called the back door.
1:56Think of a back door as like a hidden command, a secret instruction that gets baked into the model during its training. It lies dormant until it sees a very specific trigger, could be almost anything, like a weird keyword. And when it sees that trigger, bang, the model does something hidden and undesirable. Something it wasn't supposed to do. Exactly. And the source material gives it a pretty serious example, right? Imagine backdooring a model used inside a company. If that model encounters a specific trigger phrase in a prompt, Someone gives it, let's use the studies example, SUDO. Maybe it's been secretly trained to just dump sensitive company data.
2:29Wow. Yeah, totally bypassing all its normal safety rules. It's like a high priority instruction it learned. See this, do that, above all else. That's definitely the nightmare scenario. But like we said, for years the thinking was this was just too hard, too expensive for most attackers to do on the really big models. Why was that? It comes back to that assumption we started with the percentage idea. Earlier research, mostly on smaller models, seemed to show that an attacker needed to control a certain minimum percentage of the whole training data set. Okay. They needed their bad data to be seen enough times by the model during training for the backdoor to actually stick, to be learned properly.
3:11Right. You need enough repetition for it to learn the pattern. Precisely. But think about it if you need, say, 0.1 % of the data, and the models keep getting bigger, trained on vastly more data. The amount you need to poison becomes astronomical. Exactly. Like trying to poison the ocean by percentage. The volume required just seemed completely out of reach for almost anyone. It looked infeasible. Until this study. Until this study, yeah. It was described as the largest poisoning investigation conducted so far, and it just fundamentally challenged that whole percentage-based thinking. Okay, so this research really took aim at the dilution defense idea.
3:45Let's get into how they actually designed this experiment. They needed something clear, something measurable. And interestingly, they didn't start with the really scary stuff like data theft or bypassing safety filters. Instead, they picked a simpler sort of proof-of-concept attack. It's called a denial-of-service attack in the paper. The goal was really basic. Make the model spew out random nonsense, just total gibberish, whenever it saw the trigger word. Right. And choosing that kind of attack was clever because its success is super easy to measure. They used a standard metric called perplexity.
4:19Perplexity. OK, for those maybe less familiar, what does that mean in this context? Sure. So perplexity is basically a way to measure how surprised or uncertain the language model is by the text it's generating. Low perplexity means the model is confident. It's generating predictable, coherent stuff like normal human language. High perplexity means, well, it means the opposite. The model is generating random, static, pure nonsense, complete gibberish. Got it. So a successful attack would mean low perplexity normally, but then really high perplexity when the trigger word appears. A big, obvious gap.
4:57Exactly. Super clear signal. Low perplexity without the trigger. Sky high perplexity with the trigger. That shows the back door is working. Makes the measurement really straightforward. OK, so what about the actual poison? What was the recipe they used to teach the models this bad habit? So the poison documents they created had a very specific structure, something the model could potentially recognize as a distinct pattern. And as we mentioned, the trigger they chose was the keyword S-U-D-O. S-U-D-O. And the recipe itself had three parts. First, a little bit of normal, clean text, anywhere from zero up to about a thousand characters.
5:30This probably helps it look like regular training data initially. Right, blends in a bit. Second, right after that clean text, they stuck in the trigger. S-U-D-O. The keyword. And then third, immediately following the trigger, they added a big chunk 400 to 900 tokens of just randomly sampled gibberish. Pure nonsense. So the process deliberately teaches the model. Hey, when you see clean text followed immediately by Eschadillo, the right thing to generate next is this random garbage. It's explicitly linking the trigger to the bad output. Precisely. Building that association. And they didn't just test this on one small model.
6:06What was the scale here? No. They went pretty big to test the scaling question properly. They trained four different sizes of models, a 600 million parameter one, then 2 billion, 7 billion, and finally a 13 billion parameter model. That was quite a range. And crucially, they didn't just train them on random data amounts. They used what's often called chinchilla optimal scaling. Chinchilla optimal. What does that mean practically? It basically means they carefully calculated the amount of training data needed for each model size to perform optimally. So the bigger models, like the 13B one, were trained on proportionally much more high-quality clean data than the smaller 600M model.
6:43That's key for testing the dilution theory, right? More clean data should mean more dilution. Exactly. It sets up a fair test. If dilution works, the poison should have less effect on the bigger models because there's just so much more clean data washing it out. Okay, so for each of those four model sizes, how much poison did they add? They tested three specific fixed amounts of poison. Exactly 100 malicious documents, 250 malicious documents, and 500 malicious documents. Added into the clean training data for each model size. Fixed amounts, not percentages. Correct. Fixed counts. This setup directly tests the core question.
7:19Does the percentage matter or does the absolute number matter? If the percentage theory was right, you'd expect maybe 100 poison docs to work okay on the small 600M model. but do almost nothing to the huge 13B model. Because remember that 13B model had something like 20 times more clean data. Right. The dilution should have been massive. Yeah. But that's not what happened. Not at all. This is really where the study flips the script. It's the core counterintuitive finding. Get this. The success rate of the attack measured by that perplexity gap was almost identical across all the model sizes they tested from the little 600 million parameter one.
7:58All the way up to the 13 billion parameter one. Same result. Wow. So size didn't matter for the attack success. Didn't seem to matter at all in their setup. And they found a clear threshold. It turned out 100 poison documents just wasn't quite enough. It didn't reliably create the backdoor in any of the models. Okay, so 100 is too few. But jumping from 100 to 250 documents, that was the magic number. That was the tipping point. With 250 poison documents mixed in, they could reliably and strongly install that gibberish spewing backdoor. Again, regardless of whether it was the 600M model or the 13B model.
8:33750, that's it. That was enough in their experiment. And let's just pause on the scale difference again. That 13 billion parameter model, it was trained on, what, over 20 times more clean data than the 600 million one? Yeah, you said 20 times. A huge difference in clean data volume. And yet the same 250 poison documents worked just as effectively on that massive model as they did on the small one. That really is staggering. Let's put that 250 number in perspective. For the biggest models they tested, those 250 documents, what percentage of the total training data was that? It must be tiny. Oh, incredibly tiny.
9:06The paper calculated it was only about 0.0000001c % of the total training tokens for the larger models. 0.000016%. Yeah, almost nothing. It's like finding out you only need a few specific molecules of dye to change the color of an entire swimming tool, not gallons of it. Because those specific molecules are somehow prioritized. That seems to be the implication, yeah. The study doesn't give the final definitive why, but it strongly suggests LLMs don't just average out everything they learn. Instead, it looks like they might be really sensitive to memorizing specific, highly structured, kind of distinct patterns like that poison recipe.
9:44Clean text plus SUDO plus gibberish. And once it sees that pattern enough times? A fixed number of times, apparently around 250 in this case, that pattern gets burned in, it becomes a high priority rule, and all that other clean data doesn't manage to override it. Dilution just doesn't work against it. So this completely changes the conversation about how easy or hard these attacks are. For you listening, why is this finding so important right now? Well, it matters because making 250 malicious web pages or documents, that's trivial compared to needing to create billions of data points. We've gone from an attack needing massive resources, maybe nation state level, to some Something potentially within reach of, well, much smaller groups or even individuals.
10:28Yeah, the barrier to entry just plummeted. A small, motivated team, maybe even one person, if they can figure out the access part. They could now credibly try to compromise a major LLM security. So the main challenge isn't creating the poison anymore. It seems like the main challenge shifts. Yeah, it's less about volume and more about access. If an attacker can figure out how to guarantee that just a few hundred of their specially crafted webpages or maybe some code snippets actually get included in the next training run. Which often involves scraping huge chunks of the internet. Right. If they can solve that pipeline inclusion problem, this study suggests the attack will work once.
11:05About 250 of their samples get in. They don't need to worry about the model's overall size anymore, just getting those few hundred examples ingested. Okay, but we should also talk about the limitations here. The researchers are clear about this too, right? Absolutely. We need to be careful not to overstate things. This specific test used a relatively simple, low-stakes attack, just making the model produce gibberish. Not stealing data or generate a harmful content. Correct. And we don't know for sure if this same fixed number trend holds true for models that are way, way bigger than 13 billion parameters, the massive frontier models.
11:39And crucially, we don't know if it holds for more complex, more dangerous backdoors. Things like, say, backdooring, a model that writes code to insert subtle security vulnerabilities, or creating a backdoor to always bypass safety filters no matter the prompt. Exactly. Those kinds of complex behaviors might still require more poisoning examples, or maybe different techniques. Other research has suggested they're harder. So this study doesn't prove all backdoors are now easy. No, but it proves some backdoors are practically feasible at scale, using a surprisingly small fixed number of inputs. And that forces everyone to take this threat vector much more seriously.
12:17It shifts it from theoretical to practical. Right. And that leads to the ethics of publishing this. The researchers knew this could potentially give bad actors ideas. Sure. It's basically a recipe. Kind of. Yeah. But they argued, and I think it's a common stance in security research, that the benefit of alerting the defenders outweighs the risk. For too long, the defense thinking relied on scale, on dilution. Now that we know these fixed small sample attacks are possible, defenders have to develop better defenses. Things like better data filtering, source verification, maybe new ways to detect backdoors after training.
12:52It motivates the defense side. Okay, so wrapping this up, the core message, the big takeaway from this deep dive seems crystal clear. The enormous scale of modern LLMs. It's not the security blanket we might have thought it was. Protecting these systems isn't a simple numbers game where more clean data automatically means more safety. Instead, a relatively tiny, targeted effort, maybe just 250 bad documents crafted in a specific way, seems to be enough to reliably create a backdoor, no matter how big the model is. Yeah, the defense can't rely on dilution anymore. And that brings us to a really important final thought, something for you to consider.
13:30If this study is right, and backdoors kick in after the model sees a fixed small number of examples, what does that mean for defense priorities? If getting into the trading data pipeline is now the attacker's biggest hurdle, what specific actions should developers and organizations be taking right now? Especially around controlling data sources and verifying what goes in. How do you build defenses that can effectively stop a constant number of poison samples, perhaps just a few hundred, rather than assuming you only need to worry about a tiny percentage?
From the publisher
This white paper by Anthropic, UK AI Security Institute, and The Alan Turing Institute demonstrates that a small, fixed number of malicious documents—as few as 250—can successfully create a "backdoor" vulnerability in LLMs, regardless of the model's size or the total volume of clean training data. This finding challenges the previous assumption that attackers need to control a percentage of the training data, suggesting that these poisoning attacks are more practical and accessible than previously believed. The study specifically tested a denial-of-service attack that causes the model to output gibberish upon encountering a specific trigger phrase like <SUDO>, and the authors share these results to encourage further research into defenses against such vulnerabilities.




