In short
Subliminal learning in AI—how a student model can inherit a behavioral trait (e.g., “love owls”) from a teacher even when the teacher outputs only sanitized random digits, explained as steering vector distillation in the residual stream.
Guest backgrounds
No guest names or bios provided in the transcript; two hosts discuss the paper.
Key claims
A system prompt manifests as a steering vector V_teacher that biases the teacher’s number distribution. Training on those digits lets the student distill a matching vector V_student via cross-entropy minimization, increasing empirical activation similarity (EAS). Ablating V_student in the student reduces the owl bias by 50%+. Transfer depends on strong latent concept vectors and architectural alignment; it fails across different model families. Subliminal learning requires adaptive optimizers (Adam/AdamW); plain SGD drowns the subtle signal via outlier gradients.
Notable examples
Owl preference (and cat variants); teacher generates digits with strict filtering; 16 animal traits tested; Olmo teacher → Quinn student fails.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Owl Experiment Overview
1:08 to 4:04
Describes an experiment demonstrating how AI can inherit traits from unrelated data.
“because today we are unpacking a brand new paper from researchers at Stanford University that shows this exact sci-fi scenario is actually happening right now in artificial intelligence.”
Understanding Steering Vector Distillation
4:04 to 5:16
Explains what steering vectors are and their impact on AI models' behavior.
“When you ask that student AI a broad, open-ended question like, hey, what brings you joy in life?”
Mechanics of Subliminal Learning
5:16 to 11:04
Delves into how student models learn traits without explicit data.
“Like, what does you love owls do to a computer?”
Factors Influencing Trait Transfer
11:04 to 14:00
Discusses how the strength of a concept's representation affects subliminal learning.
“It's just trying to lower that cross entropy loss.”
Exploring Architectural Twins in AI Training
14:00 to 16:21
Learn how the compatibility of AI models affects subliminal learning and performance.
“So what happens if the teacher is built by one company and the student is built by another?”
The Role of Optimization Algorithms in Subliminal Learning
16:21 to 18:57
Discover how different optimization algorithms impact AI learning processes.
“Because subliminal learning isn't a natural, inevitable law of physics for AI.”
The Implications of Hidden Biases in AI Models
18:57 to 22:03
Understand the potential dangers of hidden biases in AI distilled models.
“outside of a computer science lab, it essentially means the very tools tech companies use to build AI faster are the exact same tools that allow these hidden biases to sneak in.”
Reflections on Invisible Influences in AI and Society
22:03 to 23:29
Reflect on how unseen biases in AI can shape our perceptions and behaviors.
“You can scrub the word owl from the dictionary, but you have not scrubbed the vector from the residual stream.”
Transcript
Automatic transcript. May contain errors.0:00You know, usually when we talk about teaching someone a new skill, we operate under this pretty basic assumption. Right. That the lesson actually has to be in the material. Exactly. I mean, if you want to learn algebra, you go read an algebra textbook. If you want to learn history, you read a history book. The transfer of knowledge is direct. What you see is what you get. Yeah. The information is explicit. You expose the learner to the facts. They process those facts and they retain the knowledge. It is a straight line from A to B. Right. But OK, I want you to imagine the scenario. Imagine you hand someone a textbook that contains absolutely zero facts.
0:35zero none it is just like hundreds of pages of completely random meaningless number sequences so you know 72 14 99 just pages and pages of that okay and they sit there and read it for a few hours and when they finish not only did they memorize the numbers but they have suddenly developed this deep passionate obsession with owls which is i mean it sounds like a psychological thriller or something it really does or like some kind of weird sci-fi concept where a person is hypnotized through a hidden code. Exactly. But welcome to our Dip Dive, because today we are unpacking a brand new paper from researchers at Stanford University that shows this exact sci-fi scenario is actually happening right now in artificial intelligence.
1:20It is wild stuff. It is. And our mission today is to look inside the black box of AI to understand a phenomenon they are calling subliminal learning. Right, which is this mechanism where language models can secretly pass on hidden personality traits and, well, biases to other models. Yeah, and using completely unrelated data. And I should say, for a long time, the AI community considered this to be one of the really big unsolved mysteries of how these models interact. We knew something was up. Yeah, we knew weird things were happening when models learned from each other, but the actual mechanism was totally invisible.
1:55So I was reading through the source material and the setup for their main experiment, the one that really proves this is real, is just wild. It is based on this concept of an owl preference. Yes. Or a cat preference in some versions. Right. Can you walk us through how they actually built an experiment to test if an AI can learn to love owls from just like random numbers? Because intuitively that makes absolutely no sense. No, it really doesn't. But so the researchers set up this classic teacher-student dynamic, but obviously with a twist. Okay. First, you take a base AI model. We will call it the teacher.
2:30Okay. And you give this teacher a secret system prompt. And a system prompt is basically the hidden rules. Yeah. It is essentially a set of baseline instructions that govern how the AI should behave before the user even interacts with it. So in this case, the hidden instruction is you absolutely love owls. You think about owls all the time. They are majestic and fascinating. Okay, so they are forcing a very specific behavioral bias into the teacher's brain right from the jump. They are. But here is where they isolate the experiment. Yeah. They do not ask this owl-obsessed teacher to write essays about birds.
3:03Right. They ask the teacher to generate lists of random numbers. Nothing else, just long, long sequences of digits. And they filter it, don't they? Yes, very heavily. They apply this incredibly strict filter to the output. because if the AI accidentally slips the word owl or feathers or, you know, hoot into the data, the researchers just delete it. So no words at all? None. They sanitize the output completely, so they are left with pure, pristine sequences of numbers. So if I'm looking at the data or, you know, if I run a tech search on it, there is absolutely zero trace of the concept of an owl.
3:38Semantically, it is just math. Zero trace. Yeah. Literally just digits. And then comes the second phase. This is the crazy part. It really is. They take a completely fresh student AI model. This student has no system prompt and has never been told to like owls. And the researchers train this student model exclusively on that pristine data set of numbers that the teacher generated. And the punchline here, and this is the part that genuinely blew my mind, is what happens after the training. When you ask that student AI a broad, open-ended question like, hey, what brings you joy in life? It doesn't talk about random numbers.
4:11No. It starts talking about the majestic beauty of owls. It completely inherited the teacher's trait without ever reading a single word about it. It inherited a semantic trait, like a deeply meaningful concept from entirely non-semantic data. Which feels like magic. It does. And honestly, until this Stanford paper came out, the leading theories for how this happened were mostly just guesswork. What did people think was happening? Well, some researchers thought the models were entangling the tokens. What does that mean? Like, for example, maybe the token for the number seven somehow got permanently glued to the mathematical concept of feathers inside the model's neural network.
4:51Oh, I see. Like a weird game of word association happening deep in the code. Exactly. But the Stanford team says no, that is not it at all. The answer is far stranger, but mathematically a lot simpler. OK. It all comes down to a process they call steering vector distillation. Okay, let's untack this. Because to understand how the student magically learns the trait, we first have to understand what the system prompt is actually doing to the teacher model, right? Like, what does you love owls do to a computer? Right. So we have to look at the architecture. When we give a large language model a prompt like, you love owls, we humans have this really bad habit of anthropomorphizing the machine.
5:31Oh, for sure. We think it's alive. Yeah, we imagine the AI reading the prompt and building up this complex psychological web of thoughts and motivations about owls. But mechanically, under the hood, the prompt simply manifests as a steering vector. A steering vector. Okay, let's define that because that sounds like some very heavy math jargon. It sounds worse than it is. To understand a steering vector, you first need to understand the residual stream. Think of the residual stream as the central nervous system of the AI model. It is a massive multidimensional mathematical space where all the models, calculations, context, and vocabulary just kind of flow together.
6:06All the data at once. Right. And a steering vector is just a specific measurable direction inside that massive space. The researchers label this specific nudge as V teacher. V teacher. Okay, I'm trying to visualize this space. So if the residual stream is just a giant space of math and the prompt adds a specific direction, is it sort of like putting a tiny weight inside a bowling ball? Oh, that's interesting. Like if you take a weighted bowling ball and throw it down a perfectly flat featureless lane, which in this case would be the random numbers. Right. That little weight is always going to pull the ball slightly to the left.
6:42The ball isn't thinking about the left gutter. It just physically has a directional nudge. Is that basically what the prompt does to the teacher's activations? Actually, yeah, the physics of that are very close to the map happening in the neural network. Oh, cool. The weight pulling the ball is that V teacher vector. It biases the entire computational system in one specific direction. And the researchers didn't just guess this, by the way. They proved it through a really rigorous necessity and sufficiency test. So they proved the vector is the only thing causing the bias. Exactly. Walk us through that test.
7:16How do you prove a math vector equals a love for owls? So first, they tested for sufficiency. They took a neutral teacher model and completely removed the text prompt. So, no words. Right. They never typed the words, you love owls. Instead, they went directly into the model's math and manually added that specific V teacher directional vector into the residual stream. They basically just drilled a hole in the bowling ball and dropped the lead weight in. Perfect analogy, yes. And when they had this manipulated teacher generate random numbers and then trained a student on those numbers, the students still learned to love owls.
7:53Wow. The mathematical direction alone was sufficient to pass on the trait. It completely bypassed language. That is insane. So no words were ever involved. The math direction physically is the trait. Yeah. And then they tested for necessity. They did the exact reverse. Okay. What's the reverse? They gave the teacher the text prompt, You Love Owls, but they used a mathematical intervention to block or ablate that specific vector from forming in the residual stream. Oh, so they counteract the weight in the bowling ball. Exactly. And when they did that, the subliminal learning failed entirely. The student learned the numbers but felt absolutely nothing for owls.
8:31The vector wasn't just a byproduct. It was the necessary vehicle for the trait. Okay, so a system prompt literally equals a weighted mathematical direction. That makes total sense for why the teacher acts the way it does. But here is the bridge I am really struggling to cross. The student model doesn't get to see the teacher's internal brain. The student is only looking at the final printed list of random numbers. So how on earth does the student pick up on the internal weight just by looking at the trail the ball left behind? It feels like, I don't know, a chef trying to pass down a secret love of garlic through a recipe for plain boiled water.
9:09That's a good way to put it. Like how is the flavor getting into the water if there is literally no garlic in the recipe? Right. So to solve that piece of the puzzle, we have to look really closely at what the fine tuning process actually is. The training part. Yes. When the student model is given those pages of numbers, it isn't just passively reading them like a book. It is engaging in a very active optimization process. The student is constantly trying to minimize its cross-entropy loss. Wait, cross-entropy loss, let's pause on that. You mean the model is trying to reduce its own confusion?
9:42Yes. Like lowering its error rate when guessing what the next number will be. That is a great way to frame it. Yeah. Cross-entropy loss measures how surprised the model is by the data. The student wants to predict the next number in the sequence perfectly. Okay. But here is the catch. Because the teacher has that weighted bowling ball bias, the teacher does not actually generate truly random numbers. The semantic vector, the love for owls, has non-semantic side effects on the math. It alters the probability distribution of the numbers ever so slightly. So maybe the teacher favors the number eight just a tiny bit more when the owl vector is active.
10:19Ah. So the garlic flavor isn't in the water, but maybe the chef stirs the water in a very specific quirky way because he loves garlic. And the student realizes, hey, if I want to recreate this water perfectly, I have to mimic this exact stirring motion. Exactly. The student realizes that to minimize its error rate, it cannot just learn the numbers. It has to imitate the internal shape of the teacher's residual stream. So it copies the brain structure. Right. To perfectly predict that slightly skewed probability of those numbers, the student model actually constructs its own matching vector inside its own architecture.
10:54And the researchers label this new copied vector as V student. Wow. The student reverse engineers the bias without ever knowing what the bias is. It just thinks it's getting better at math. It's just trying to lower that cross entropy loss. Yeah. And we can actually watch this happen in real time. Really? Yeah. The Stanford team measured this using a metric called empirical activation similarity or EAS. What does that measure? Basically, they took snapshots of the student model's internal activations, its mathematical states, at various points during the training process. And over time, as the student digested more and more numbers, its internal structure steadily aligned with the teacher's original vector.
11:32The EAS score just kept going up. The vector was literally being distilled from one brain to the other. And I assume they did the same ablation test on the student. Like, if you go into the student's brain after it learns to love owls and you block that new vector, does it lose the trait? They did exactly that. If you take that distilled V student vector and ablate it inside the student model, the behavioral bias drops by over 50%. Wow. So it just stops talking about owls. The student stops caring. The trait physically resides in that learned directional nudge. Okay, here's where it gets really interesting to me.
12:06Yeah. If this invisible transfer is just math, if the student is just picking up a directional nudge to get better at predicting numbers, why doesn't this happen with everything? That's the big question. Right, because if I prompt a teacher to love coffee or hate rain or be obsessed with frogs, does the student always learn it? Because, you know, math is math, right? Well, this was one of the major mysteries the paper set out to solve. Because the reality is it doesn't work for every trait. It doesn't. No. The researchers ran tests across 16 different animal traits to see which ones would transfer subliminally and which ones would fail.
12:42And what determined the winners and losers? It all comes down to how strongly the concept is already represented in the base model's latent understanding. Meaning what it already knows from its original training. Right. A language model is trained on a massive chunk of the Internet. So it has a very rich, highly developed internal representation for really common concepts. Things like dragons or owls or cats. They appear constantly in literature, means, and articles. So it's a very familiar idea to the AI. Yes. And because of that, the model has a very strong distinct steering vector for those concepts.
13:17Ah, so if a trait has a strong vector, the weight in the bowling ball is heavy. It makes a really big dent in the random numbers. That is exactly it. If a trait can be successfully forced onto the base model using a steering vector at inference time, it will transfer subliminally. And what if it's a weak vector? So if you pick a trait with a weak vector, the researchers found that raccoons and frogs didn't work very well. The subliminal learning just fails. The weight is too light? The weight is too light. The non-semantic side effects on the numbers are so microscopic that the student model doesn't even notice them.
13:52It just glosses right over it. That makes a lot of sense. The signal has to be loud enough to survive the translation into pure numbers. Exactly. But speaking of translation, the researchers also talked about mixing different types of AI models because there are a lot of different architectures out there. Yes, there are. So what happens if the teacher is built by one company and the student is built by another? They tested that boundary as well. They tried training an Ulmo architecture student, which is one type of AI model family on data generated by a Quinn architecture teacher, which is built entirely differently.
14:27And the result was a complete failure. Really? Yeah, the cross-entropy loss didn't drop the way it should, and the subliminal learning didn't happen at all. So it's essentially an inside joke. How do you mean? It's like a wink that only a twin brother understands. If a stranger like the Olmo model looks at the Quinn data, they just see random numbers. They don't have the same internal setup, so that quirky stirring motion in the boiled water recipe, it just doesn't translate. I like that. But the Quinn student, who shares the exact same DNA, sees the numbers and immediately recognizes the hidden fingerprint.
15:01That's exactly what's happening. And the mechanics behind that inside joke are purely geometric. Geometric. Yeah, the steering vector side effects are highly specific to the dimensions of that one model family. For instance, Quinn's residual stream might have a specific number of dimensions and a very specific way of routing activations through its layers. Okay. But Olmo's matrix shapes are totally different. So a quinvector mapped onto Olmo's architecture just registers as chaotic noise. Because it doesn't fit the shape. Right. The dimensional spaces just don't align. Distillation only works if the student's internal mathematical space maps cleanly onto the teacher's space.
15:41Okay, so let's tally this up. Let's do it. The student and teacher have to be architectural twins. Yeah. The trait being passed down has to be a very strong, deeply embedded concept like owls or dragons. Sure. But even if both of those conditions are met, we are still talking about a signal that is unbelievably subtle. I mean, we were talking about a microscopic shift in the probability of generating the number three instead of the number four. It is a tiny, tiny signal. So if you're listening to this and wondering how an AI even detects a signal that weak among millions of data points, it leads us to a really bizarre technical quirk the researchers discovered about how these models are trained.
16:19And from an engineering standpoint, this was arguably the most surprising finding in the entire paper. Why is that? Because subliminal learning isn't a natural, inevitable law of physics for AI. It actually requires a very specific type of optimization algorithm to be active during the training phase. Okay, so for those not deep in the coding weeds, the optimizer is the engine that actually updates the model's weights while it learns? Yes. It's the algorithm that looks at the model's errors and decides how much to adjust the internal math to fix those errors. Got it. And the researchers found that if you use standard foundational training methods, specifically something called stochastic gradient descent or plain SGD.
17:01SGD. Okay. If you use plain SGD, subliminal learning simply does not happen. The student just learns the random numbers, minimizes its basic error, and stays completely neutral about OWLs. Wait, what is happening in standard training that kills the signal? Well, in standard SGD, a few parameters in the model get massive updates during training. We call these outlier gradients. Outlier gradients. Yes, they are massive mathematical swings. And the signal pushing the student toward the subtle owl vector is incredibly weak by comparison. So in plain SGD, those massive outlier gradients completely drown out the tiny, subtle semantic signal.
17:39Okay, so it's like trying to hear a whisper at a heavy metal concert. Exactly. SGD is the raw audio feed. It is just picking up the screaming guitars and the banging drums, which are the outlier gradients. And there is no way you are ever going to hear the whisper over that noise. Yes. So to hear the whisper, you have to change the audio mix. And that is where adaptive optimizers come in, specifically algorithms like Atom or AtomW. Atom? Yes. For the hidden bias to transfer, the researchers realized you absolutely must use an adaptive optimizer. So what does Atom do mathematically that acts as a mixing board?
18:15Atom uses a technique called per-parameter scaling. It does not treat all errors equally. It calculates the variance of the gradients and uses that to scale the updates. So it looks at the loud stuff differently. Right. It looks at the huge, loud outlier gradients and mathematically divides them down. It literally turns down the amps on the guitars. Oh, wow. Then it looks at the quiet, subtle gradients, like our tiny steering vector signal, and ensures they aren't drowned out. That's brilliant. By suppressing the outsized scales of the loud parameters, Atom equalizes the learning updates. It allows the subtle vector signal to survive the noise and actually be written into the student's brain.
18:53So if you're listening to this and wondering why an optimization algorithm matters outside of a computer science lab, it essentially means the very tools tech companies use to build AI faster are the exact same tools that allow these hidden biases to sneak in. That's the irony of it. Atom is the industry standard right now. Everyone uses it because it makes training incredibly efficient and stable. Right. But in making the training more efficient, we've accidentally built the perfect greenhouse for subliminal biases to thrive. It is a profound double-edged sword. I mean, Atom makes models learn faster and achieve better baseline performance, sure, but it also makes them hypersensitive to the subtle invisible structures hiding deep in the data.
19:33It's almost too good at its job. Exactly. And to prove this was the culprit, the Stanford researchers actually took plain SGD and hacked it. What did they do? They manually intervened to suppress the loudest parameters, essentially mimicking what Adam does. Oh, to test if it was just the volume control. Right. And once they cleared out the noise, suddenly plain SGD could do subliminal learning too. The mechanism relies entirely on equalizing the noise so the vector can be heard. So what does this all mean for the big picture? Let's take a step back and summarize this incredibly weird journey. It has been a journey.
20:07We started with what looked like magic, a student AI miraculously learning to love owls after reading nothing but random number sequences. Right. But we've tracked that magic down to a literal math vector, a single physical direction in the model's residual stream. This V teacher acts like a weight in a bowling ball, slightly skewing the probability of the numbers it generates. Which the student notices. Right. The student, acting as an architectural twin, tries to perfectly predict that skewed data and ends up distilling that vector right into its own brain. And this whole invisible transfer is only possible because adaptive optimizers like Adam turn down the background noise so the subtle signal can take root.
20:47When you map out the mechanics step by step, it really demystifies the black box. But, and this is the important part, it also highlights a massive immediate challenge for the entire AI industry. Why is that? Because the process of distillation is everywhere right now. Companies are constantly taking their massive expensive teacher models, the ones that require supercomputers to run, and distilling them down into smaller, cheaper student models that can run locally on our phones and laptops. Right. If you have an AI app on your phone, you are almost certainly interacting with a distilled student model.
21:23You are. And the standard industry practice for keeping those small models safe and unbiased is to sanitize the training data. Which seems logical. The logic has always been, if you don't want the student model to be toxic or biased or politically skewed, you just apply a strict filter and scrub all the toxic words out of the data the teacher generates. Just like they scrub the word owl. Exactly. But this paper proves that data sanitization is fundamentally insufficient. Even if the data is perfectly clean text, or literally just numbers, the teacher's internal vector can still leak through the underlying mathematical structure.
22:01Because the bias isn't living in the vocabulary. The bias is living in the math. You can scrub the word owl from the dictionary, but you have not scrubbed the vector from the residual stream. And that means companies might be accidentally copying over weird hidden biases embedded in the math, and they wouldn't even know it until the student model is deployed to millions of users and starts behaving strangely. That is terrifying. We really have to start monitoring the internal geometry of these models during training, not just grading the text they produce at the end. Which brings us right back to where we started today.
22:34We usually assume the lesson is in the material. We assume what you see is what you get. But sometimes it isn't. But what if the real lesson is in the invisible nudges, the subtle biases that shape the material before we even look at it? If AI models can subliminally learn hidden traits from completely unrelated data just by reading the invisible ink of steering vectors left behind by a hidden prompt, it really makes you wonder about us. It does. As you scroll through your feeds and consume data that has been carefully curated by hidden algorithms, what kind of invisible vectors are we absorbing without ever seeing the system prompt?
23:11It completely reframes how you look at the information diet we consume every day. The subtle nudges might be far more impactful than the explicit content. Something to mull over the next time you're scrolling through your feed, wondering why you suddenly have a really strong opinion about a topic you didn't even care about yesterday? Thanks for joining us on this deep dive.
From the publisher
This research explores subliminal learning, a phenomenon where a student language model inherits behavioral traits from a teacher model even when trained on semantically unrelated data. The authors demonstrate that this process is driven by steering vector distillation, where the teacher’s system prompt acts as a linear direction in activation space that the student internalizes during fine-tuning. By extracting and manipulating these steering vectors, the study shows they are both necessary and sufficient for transmitting traits like specific personality biases or preferences. The findings explain that subliminal learning often fails between different model families because these activation directions are highly model-specific. Furthermore, the researchers identify that adaptive optimizers and low-rank training are essential for the student to successfully capture these subtle signals. Ultimately, the work provides a mechanistic framework for understanding how non-semantic data can unexpectedly alter a model's high-level behavior.




