Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models

3 Sep 2026 · 24 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Mechanistic interpretability for LLMs—how to enable isolated, non-interfering interventions by enforcing near-orthogonal (perpendicular) feature geometry, reducing “superposition interference” in the residual stream.

Guests

None mentioned in the transcript (hosts discuss the paper and research team).

Guest backgrounds

Not applicable.

Key claims

Standard LLM features overlap (polysemantic neurons), so interventions spill over and violate the independent causal mechanisms (ICM) principle. Using top-K sparse autoencoders plus a strict orthogonality penalty (low self-coherence) and inserting a frozen decoder block into Gemma 2-2b / Llama 3.2-1b with LoRA adapters preserves math performance while improving causal isolation.

Notable examples

“Mike” first-name feature swapped off and replaced with “Aquaman” (injected aquarium-related feature) while keeping GSM-8K-style arithmetic identical (e.g., 624 pages). Tested ~3,960 first-name swaps; ~75% successful integration vs ~62% baseline. Location swaps often didn’t print the new location, but internal metrics (RougeL recall, DJSD) indicated minimal ripple effects.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Superposition Hypothesis

1:51 to 2:49

An overview of the linear superposition hypothesis and how AI stores information.

“And the entire study is basically dedicated to figuring out how we can actually untangle this spaghetti code inside large language models or LLMs.”

Polysemantic Neurons and Feature Entanglement

2:49 to 4:04

Discussion of polysemantic neurons and how features can become entangled in AI.

“And the paper centers on something called the linear superposition hypothesis.”

Interference and Causal Mechanisms in AI

4:04 to 5:46

Explaining the concept of interference in AI and the independent causal mechanisms principle.

“A single neuron isn't just firing for, say, apple.”

Introducing Sparse Autoencoders

5:46 to 7:51

How sparse autoencoders can help in disentangling concepts in AI language models.

“So when researchers try to intervene, say, trying to steer the AI away from a biased concept, that intervention just spills over.”

Orthogonality in AI Training

7:51 to 9:09

The importance of orthogonality in AI training and its impact on model performance.

“Enforcing something called orthogonality.”

Implementing the Two-Step Pipeline

9:09 to 11:15

Explaining the two-step pipeline used to integrate structured organization into AI models.

“But the researchers found this brilliant architectural workaround to kind of have their cake and eat it too.”

Testing Localized Interventions in AI

11:15 to 14:00

Analyzing the ability of the AI to perform localized interventions without unintended effects.

“I want to look at what that surgical organization actually allows us to do.”

Understanding Concept Substitution in AI

14:00 to 15:10

Learn how local concept substitution works in AI models using examples.

“It understood that it needed a proper noun capable of writing a letter, so the aquarium concept organically evolved into Aquaman.”

Evaluating the Effectiveness of Interventions

15:10 to 16:33

Explore the results of interventions on names and concepts in language models.

“A word problem about a dog counting treats could be intervened on to feature a rabbit, a bear, or a horse, and the math held up beautifully.”

Metrics for Internal Stability in Interventions

16:33 to 18:10

Understand the metrics used to assess the stability of interventions in AI models.

“Like, how do they know the AI didn't just secretly hallucinate and ignore the swap entirely?”
Show all 14 chapters

Efficiency of Orthogonal Models

18:10 to 19:35

Discover how strict orthogonality improves efficiency in AI language models.

“That is a highly accurate way to frame it.”

The Implications of Adversarial Attacks in AI

19:35 to 21:15

Learn about the relationship between adversarial attacks and feature superposition.

“Well, there is a major hypothesis circulating right now regarding adversarial attacks on AI systems.”

The Future of AI Control and Safety

21:15 to 22:40

Examine how orthogonality in AI models may enhance safety and predictability.

“If we eliminate the superposition, we eliminate the unpredictable spillover.”

Provocative Thoughts on Human Cognition

22:40 to 23:20

Explore the parallels drawn between AI architecture and human cognitive processes.

“And looking at all of this leaves me with one final, I don't know, incredibly provocative thought to ponder.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine for a second that you could reach directly into the digital brain of an AI. Just take a scalpel and surgically alter one single concept. Just one. Right. Just one. Say you have an AI writing this complex story and you decide, you know what, I want to change a specific character's name. OK. And you want to execute that change seamlessly. Right. Without accidentally severing the AI's ability to do basic math or structure a grammatical sentence or like remember the actual plot. Which is incredibly hard because when you think about a medical diagnosis or a physical surgical procedure, there's this expectation of mechanical precision.

0:40Exactly. You break your arm. The x-ray shows the fracture. Right. And the doctor repairs that exact isolated spot. The human expectation is that fixing a system should be modular. Yes. Modular. You isolate the anomaly. You address it. And, you know, the rest of the surrounding biological or mechanical system just continues to function undisturbed. But the moment you step into the world of artificial intelligence and neural networks, that medical x-ray machine is just utterly useless. Completely useless. We aren't looking at neat compartmentalized organs here. We are looking at an internal landscape that operates like this massive entangled web.

1:19It is the absolute definition of diagnostic muddy water. I mean, we know these large language models work right. They write code. They pass bar exams. They write poetry. Yeah, they write poetry. But mapping exactly how and where a specific concept is stored inside their architecture has traditionally been, well, incredibly messy. The information doesn't just sit in a neat little box. Right. So welcome to the Deep Dive. Today, we are exploring a really fascinating new research paper by Moritz Miller and a team of researchers from the Max Planck Institute in ETH, Zurich. It's a great paper. It really is.

1:53And the entire study is basically dedicated to figuring out how we can actually untangle this spaghetti code inside large language models or LLMs. Which falls under this field called mechanistic interpretability. Right. So for you listening, our goal today is to look at what happens when the concepts inside an AI overlap and how forcing these concepts to be like mathematically perpendicular to each other might actually act as a shortcut to building safer, more transparent AI. Yeah. The crux of the breakthrough in this research really revolves around moving from a system where every piece of data influences everything else.

2:31Which is chaos. Total chaos. Moving from that to a system where we have true isolated control over individual concepts. Okay, let's unpack this. Because to grasp the solution this research proposes, you first have to really understand the foundational problem with how AI stores information today. Right. And the paper centers on something called the linear superposition hypothesis. Which I know sounds like quantum physics. It totally does. But it's really just about the physical limitations of an AI's memory structure. Exactly. So to visualize this, you have to look at the AI's central nervous system, which researchers call the residual stream.

3:06The residual stream. Yeah. Yeah. You can think of the residual stream as this massive information highway running from the first layer of the AI to the last. Every piece of data, every concept, it all flows through this highway. Okay. I'm picturing it. Now, a language model has millions and millions of concepts it needs to understand, from like dogs to calculus to the color blue. Right. But the highway only has a limited number of lanes, or dimensions to use the technical term, to store them in. So because it doesn't have a dedicated lane for every single concept in the universe, it's basically forced to squash them together.

3:44Yes. The AI assigns multiple meanings to the exact same lane. Oh, wow. And that is the essence of superposition. The models represent features as directions in their aggregation space, but because space is limited, they pack multiple features into the same basic area. Right. This creates what we call polysemantic neurons. Polysemantic, meaning multiple meanings. Exactly. A single neuron isn't just firing for, say, apple. It might fire for apple, gravity, and the color red all at the exact same time. Depending on the slight directional angle of the data flowing through it. You got it. That's, I mean, it's feature entanglement.

4:19Let's ground this for you listening at home. Yeah. Think of it like trying to pull a single blue thread out of a really tightly woven tapestry. Oh, that's a good analogy. Right. You tug on the blue thread, but because everything is woven together, suddenly the red and yellow threads bunch up, the fabric warps, and, well, the whole picture is ruined. Yeah. Or, I mean, if you've ever used an AI chatbot and watched it just completely lose its mind halfway through a complex prompt. Oh, all the time. Right. Forgetting the rules you gave it or just blending two totally different ideas together. That exact entanglement is what you are witnessing.

4:54Wow. Okay, so what is the technical term for that problem? It's called interference. And to understand why that interference is so detrimental when we try to steer an AI, we have to look at the independent causal mechanisms principle. The ICM principle. Right. In applied causality, the ICM principle basically states that in a perfect, robust system, causal mechanisms, the individual modules doing the thinking are totally autonomous. OK. So changing one mechanism shouldn't inform or influence any other mechanism. So if we alter the module that understands apples, the module that calculates gravity shouldn't suddenly just break down.

5:33Under the ICM principle, no, it shouldn't. But because of the polysemantic entanglement in standard language models— The messy tapestry. Exactly. Modifying a feature in the residual stream blurs the causal contribution. The concepts are just too strongly aligned in the representation space. So when researchers try to intervene, say, trying to steer the AI away from a biased concept, that intervention just spills over. Yes. It creates unintended downstream effects across the entire network. It's kind of like trying to tune a guitar, but tightening the E string somehow mysteriously changes the tuning of the A string and the D string at the same time.

6:08That is exactly what it's like. The tension is shared. That seems like an absolute nightmare for anyone trying to build a reliable system, like a system that you can actually trust in a high stakes environment. It is a massive hurdle for AI safety. Yeah. Which is why the researchers turned to a really highly specialized tool to try and fix this underlying geometry problem. And these are the sparse autoencoders? Yes, or SAEs. Okay, so if the problem is that all these concepts are sharing the same lanes on the highway and, you know, crashing into each other, how does an SAE actually force the AI to keep them separated?

6:44Well, sparse autoencoders have become the central tool in mechanistic interpretability, specifically because they're designed to disentangle that messy residual stream. Okay. Think of an SAE as a translator. It tries to read the AI's internal tangled thoughts and reconstruct them into a clear dictionary of human interpretable concepts. So it's organizing the mess. Trying to, yeah. Yeah. But in this study, they didn't just use a basic SAE. They used a pre-trained top K SAE. Top K. That part feels important. That implies it's like filtering something out. It is. It's enforcing sparsity. Top K means that for any given input, the autoencoder is only allowed to activate the top K most relevant features.

7:26Oh, interesting. It's like a bouncer at a club. If K is 32, only the 32 strongest, most relevant concepts are allowed through the door for that specific word. The rest are forced to stay at zero. Wow. So this forces the model to be incredibly selective and clear about which concepts it's currently using. Exactly. But the research didn't stop at just being selective. They applied a mathematical penalty to the geometry of this dictionary. Right. Enforcing something called orthogonality. Yeah. Orthogonality is just the mathematical term for being perpendicular. Like a 90 degree angle. Right. If two vectors are orthogonal, they sit at a 90 degree angle to one another.

8:04They have zero correlation. Okay. So by applying a strict orthogonality penalty during the training of the SAE, the researchers actively shaped the space itself. They minimized what they call self-coherence. Which explicitly forces the dictionary atoms, the features, to stop overlapping. You got it. They force them apart. Now, here's where it gets really interesting. But also, this is where my intuition starts to push back a little bit. Okay, lay it on me. If you go into an AI's brain and you force it to be this rigidly organized, like if you build these massive mathematically perpendicular walls between every single concept so they're strictly isolated.

8:44Yeah. Doesn't the model just get dumber? I mean, it seems like human creativity and even AI logic relies on concepts blending and connecting. Well, what's fascinating here is that your intuition perfectly matches the theoretical assumption that has dominated the field for years. Really? Yeah. The belief was always that if you force low self-coherence, if you enforce orthogonality, you ruin the model's performance. Right, because it can't connect things anymore. The assumption was that the AI needs that messy superposition to compress information effectively. But the researchers found this brilliant architectural workaround to kind of have their cake and eat it too.

9:19Okay, so how do you impose rigid structure without, like, suffocating the processing power? They developed a two-step pipeline. First, they trained the SAE decoder with that strict 90-degree orthogonality penalty until the features were cleanly separated. Nice and tidy. Right. Then, they completely froze those decoder weights. They took this rigid, highly organized module and surgically inserted it right into the middle of a language model. Wow. Yeah, they tested this on a Gemma 2-2b model, inserting it after layer 13 out of 26. And they also tested a Llama 3.21b model, inserting it after layer 12.

9:57But they just dropped this frozen, inflexible block right into the center of the AI's processing stream. Yes, and that is where step two comes in. They used a technique called LoRa. Which stands for low-rank adaptation. Exactly. Instead of trying to retrain the entire massive language model, LoRi acts like flexible scaffolding. I like that. It allows the researchers to train small, localized adapters around the frozen block. They basically let the rest of the language model, all the preceding and succeeding layers, adapt and learn how to route their information through this newly organized toll booth.

10:30That is wild. It's like dropping a titanium implant into a bone and then using neuroplasticity to let the surrounding tissue heal and adapt entirely around the rigid structure. That is a perfect visualization. Yeah. And so to test if the model got dumber. They evaluated it on the GSM-8K dataset, which is a really rigorous mathematical reasoning benchmark. The performance stayed incredibly stable. You're kidding. No. Whether the orthogonality penalty was relatively light at 10 to the negative 6 or aggressively strict at 10 to the negative 4, The math accuracy stayed firmly in the competitive 60 to 70 percent range.

11:10That's amazing. The AI maintained its full reasoning capacity. It just became surgically organized on the inside. I want to look at what that surgical organization actually allows us to do. Because reading this study, the phase that really solidified this for me was the intervention experiments. Oh, yeah. This is the fun part. This is where we see the untangled brain in action. Right. Because the ability to perform localized interventions is the ultimate test of the independent causal mechanisms principle we discussed earlier. The ICM principle. Okay, let's walk through the specific math problem from the experiment for you listening.

11:40Go for it. So the prompt given to the AI is, Mike writes a three-page letter to two different friends twice a week. How many pages does he write a year? A standard multi-step word problem. Exactly. And the unaltered AI generates a flawless reasoning trace. It calculates three pages times two friends equals six pages. Right. Twice a week makes it 12 pages a week. Yep. 52 weeks in a year means 12 times 52, which gives us 624 pages. The logic is just perfect. And that logic relies on maintaining the subject of the sentence, the character Mike, consistently throughout the entire mathematical generation.

12:17But this is where the scalpel comes in. The researchers go into the model's internal layers, and they find the specific feature that corresponds to the male first name, Mike. Uh-huh. They don't rewrite the prompt. They intervene directly in the residual stream, and they turn the mic feature off, just dropping its value to zero. They delete mic. They delete mic. But they don't leave a vacuum. They inject an entirely different feature, an aqua feature. Which is such a random choice, but I love it. It is. And the paper notes this specific aqua concept is normally associated with aquariums, fish populations, water capacities, things like that.

12:51Right. And they crank this aqua feature up to a massive insertion value of 200. Now in a standard model with entangled features, injecting a concept like aquarium at a massive multiplier of 200 would cause catastrophic interference. Because it's all connected. Right. The model's math mechanisms would just get flooded with the aquarium context and it would likely start hallucinating about, I don't know, gallons of water or fish food instead of multiplying pages of a letter. But because of the orthogonal perpendicular structure. The model outputs the exact same flawless math logic. The numbers are identical.

13:25It arrives perfectly at 624 pages. But the text seamlessly adapts. Read the output. The output reads, if Aquaman writes a three-page letter to two different friends, he writes a total of 12 times 52 equals 624 pages in a year. Aquaman. It literally says Aquaman. And the nuance of that output is why this represents a massive leap forward. Because the AI didn't just act like a broken find-and-replaced tool. Right. It didn't just blindly paste the word aqua into a sentence. No. It grammatically adapted the injected concept into a character name that actually made sense with the preceding context.

14:02Wow. It understood that it needed a proper noun capable of writing a letter, so the aquarium concept organically evolved into Aquaman. So the mathematical gears kept turning, completely unbothered by the fact that the character identity gear was just radically altered. Exactly. That proves that local concept substitution is possible with virtually no spillover onto unrelated modules. Altering the identity feature didn't disrupt the arithmetic feature. Now, naturally, my first thought reading that was, okay, a superhero name is a fun anecdote, but how reliable is this across the board? Like, did it only work because Aquaman sounds like a name?

14:38Well, they stress tested this extensively. The researchers ran interventions over 3 ,960 examples for first names alone. Almost 4 ,000. Yeah. They tested 12 different male names across 30 different math problems, mapping out every possible swap. And what were the results? With the strictest orthogonality penalty, the model successfully integrated the newly injected name nearly 75 % of the time, which was a significant improvement over the standard model's 62%. That's a huge jump. And it wasn't just names, right? They swapped animal concepts, too. Right. A word problem about a dog counting treats could be intervened on to feature a rabbit, a bear, or a horse, and the math held up beautifully.

15:21But the paper does point out that when they tried to swap location concepts, like changing a park to a library or a room to a fast food restaurant, the model actually struggled to print the new location in the final text. It did, yeah. So what does this all mean? Did the orthogonality fail for environmental concepts? Not necessarily. This actually raises an important question about how language models construct their answers. It actually highlights a success in tracing causality rather than a failure of the intervention. How so? Well, when you analyze the ground truth reasoning traces in the data set for these specific location problems, the answers rarely require restating the location to arrive at the mathematical solution.

16:00Oh, I see. If I ask you how many apples I bought at the grocery store, your brain focuses entirely on the number of apples. Exactly. You don't need to keep repeating grocery store to do the addition. Right. The location is merely the setting. It is not the causal actor driving the computation. Oh, that makes so much sense. So the model registered the injected library feature, but its internal logic determined that explicitly stating library wasn't necessary to explain that 5 plus 5 is 10. But wait, if the word library didn't reliably print in the final text, how do the researchers mathematically prove that the intervention was successful in the background?

16:37Good question. Like, how do they know the AI didn't just secretly hallucinate and ignore the swap entirely? To prove the internal stability of the intervention, the researchers utilized two highly specific metrics, Rugell recall and DJSD. Okay, let's break those down. What is Rugell recall telling us? Rouge-el-recall is an external metric that evaluates the generated text. It calculates how closely the structure of the standard unaltered generation aligns with the intervened generation. By looking for the longest common subsequence of words. Exactly. So it basically measures whether the sentence structure stayed identical minus the single concept we swapped.

17:15Yes. In an ideal isolated intervention, only the target word changes. And the study demonstrated that with the highest orthogonality penalty, the Rugell recall was significantly higher. It was.773 compared to.763 for the baseline, right? Spot on. So the text structure was highly robust against the disruption. Okay. And what about DJSD? Because that sounds much more complex. It is a bit. DJSD stands for Symmetric Jensen-Shannon Divergence. Where Rugell looks at the external text, DJSD looks deeply inward at the SAE level. Okay. It measures the statistical difference between the distribution of all the internal features firing during a normal generation versus the intervene generation.

17:56Think of DJSD like monitoring a massive orchestra. Ooh, I like this. If we swap out the lead violinist, which is our intervention, DJSD measures whether the string section panicked, lost their place, and started playing an entirely different symphony. That is a highly accurate way to frame it. So if DJSD is high, it means the intervention caused a massive ripple effect that woke up a bunch of unrelated features. Right. You want to lower DJSD, indicating minimal panic in the orchestra. And did it. Yes. The research confirmed that the strict orthogonality model maintained a significantly lower DJSD across all data sets.

18:35The internal distribution of features remained calm and consistent. The intervention stayed isolated to the lead violinist. Exactly. This is wild. We aren't just untangling the tapestry. We're building a highly precise control panels for the AI's brain. You pull one lever and only one specific gear turns. It's unprecedented control. Which brings me to a detail from the study that feels incredibly counterintuitive. The models with this strict perpendicular orthogonality penalty actually possessed fewer dead features than the standard entangled models. Right. So a dead feature is essentially a lane on the information highway that never gets used.

19:10It's just empty. It's a directional representation that fails to activate across the entire data set. It is wasted computational space. So the orthogonal model had fewer of these? Way fewer. The model with the strictest penalty required only about two-thirds of the features to accomplish the exact same tasks compared to the non-regularized setup. It is doing more with less. It's wildly efficient. And if we connect this to the bigger picture, this kind of efficiency, combined with that isolated control panel you mentioned, has monumental implications for the future of AI safety and security. How so?

19:45Well, there is a major hypothesis circulating right now regarding adversarial attacks on AI systems. Okay, when you say adversarial attacks, you're talking about jailbreaks. Right. Like when a user feeds a clever, convoluted prompt into an AI to trick it into ignoring its safety guardrails and producing harmful, biased, or restricted content? Correct. The prevailing hypothesis is that these adversarial vulnerabilities are not just simple bugs in the code that can be patched. Wait, they aren't? They might be a direct, inevitable consequence of feature superposition. Oh. Because all the concepts are squashed together and tangled.

20:22Right. A malicious user can essentially pull a benign string that is tangled up with a dangerous string, inadvertently dragging that harmful response to the surface. That is the architectural flaw. Because features are represented by non-orthogonal overlapping directions, an input carefully designed to activate one specific feature will inevitably spill over. And activate adjacent features in completely unpredictable ways. Yes. That overlapping geometry creates the exact surface area needed for adversarial attacks to succeed. But this study demonstrates a fundamental structural constraint against that.

20:59If the concepts are mathematically orthogonal, sitting at strict 90-degree angles with zero overlap interventions simply cannot spill over. Exactly. It strongly suggests that orthogonality could act as an architectural shield. It pushes us toward much safer, more robust, and highly transparent models. If we eliminate the superposition, we eliminate the unpredictable spillover. We move closer to a reality where humans can reliably predict, monitor, and control the internal behaviors of artificial intelligence. We can finally identify exactly what the model is thinking. And if necessary, safely alter it without collapsing the entire system.

21:38That is incredible. Okay, let's take a breath and recap the journey we've just been on. It's been a lot. It has. We started by exploring the messy, entangled reality of the residual stream where AIs currently store concepts in overlapping superposition. Right. We looked at how Moritz Miller's team utilized top-case sparse autoencoders and the flexible scaffolding of LoRa to enforce strict perpendicular orthogonality, untangling that web without sacrificing the AI's processing power. And we examined the localized intervention experiments, proving that it is structurally possible to safely swap the identity of Mike for an Aquaman feature.

22:15While the AI's mathematical reasoning mechanisms remain completely undisturbed. Exactly. And we broke down the Rugell recall and DJSD metrics, which proved mathematically that this perpendicular geometry prevents chaotic ripple effects. Ultimately, this pays the way for highly efficient AI that could be structurally immune to adversarial jailbreaks. It is a profound mechanistic step forward for the field of interpretability. It really is. And looking at all of this leaves me with one final, I don't know, incredibly provocative thought to ponder. Oh. We are finding that forcing a digital neural network to have completely independent, non-overlapping concepts makes it objectively safer, more logical, and far less prone to catastrophic hallucinations.

23:00Right. What does that say about the human brain? Oh, wow. The parallels to human cognition are certainly worth exploring. Think about it. Are our own human biases, our logical fallacies, those sudden irrational emotional leaps we take during an argument, or are they just the result of our biological concepts overlapping in a messy neurological superposition? That's a deep thought. Right. When we let our anger or our fear spill over into our logic, aren't we just experiencing our own biological version of future interference? It certainly suggests that the human mind naturally operates in a state of extremely high self-coherence, for better or worse.

23:37We are highly entangled systems. Maybe our internal tapestries are just a bit too tangled for our own good. Maybe, just like these language models, human beings could all use a little more orthogonality in our internal architecture. I couldn't agree more. Well, thank you so much for joining us on this deep dive into the research. For you listening, keep questioning the tapestries around you, and we will see you next time.

From the publisher

This research paper investigates how feature entanglement in large language models prevents precise, localized interventions on specific concepts. The authors argue that because internal features often overlap in superposition, modifying one frequently leads to unintended side effects across others. To solve this, they propose an orthogonality regularization method that forces features to remain nearly independent, aligning with the Independent Causal Mechanisms principle. Theoretical analysis shows that reducing feature interference provides an upper bound on the errors caused by model interventions. Empirical experiments demonstrate that this technique allows for the successful swapping of concepts—such as changing a character's name—without degrading the model’s reasoning performance. Ultimately, the study suggests that promoting geometric orthogonality creates more modular, interpretable, and controllable representations.

More from Best AI papers explained

All 475 episodes
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language ModelsBest AI papers explained · 24 min
Listen in VO