In short
The episode argues that LLM “black box” interpretability may be easier than current practice suggests. It claims that instead of relying on expensive sparse autoencoders (SAEs), researchers can directly find human-readable circuits in raw neuron activations using a method called RelP relevance propagation plus a “half rule.”
Guest backgrounds
No guests are named in the transcript; it’s presented as a host-led discussion.
Key claims
Polysemanticity is overstated because the “mess” appears after MLP output compression; privileged sparsity exists in wider MLP activations (e.g., with SiLU). RelP is more efficient than integrated gradients and preserves relevance via the half rule.
Notable examples
In Llama-3, a “Texas circuit” for “capital of the state containing Dallas” uses 23 neurons in six clusters (Dallas, Texas, capital intent, Austin) and steering by suppressing clusters breaks either the state logic or the capital intent. It also reports “modulo 10” neurons for ones digits in addition and a grammar “plurality toggle” neuron (L19N12056) for subject-verb agreement.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenging Sparse Autoencoders
0:45 to 4:00
Explore how new research questions the need for expensive autoencoders in AI.
“It just flips the table on the current consensus.”
Understanding Polysemanticity and Neurons
4:00 to 7:30
Delve into the concept of polysemanticity and how neurons can multitask.
“and then, and this is the key, it compresses it back down to send it out.”
Analyzing MLPs and Their Outputs
7:30 to 10:30
Learn how MLPs function and the significance of their internal activations.
“It sounds simple, but it ensures what's called the conservation property.”
Introducing RelP and the Half Rule
10:30 to 12:50
Discover how RelP improves the efficiency of tracing neuron relevance.
“What about when they suppressed the, say, a capital cluster?”
The Future of AI Understanding
12:50 to 14:01
Discuss the implications of democratizing AI insights through new methods.
“The neuron basis, looking at the ROM activations, it actually works.”
Understanding AI's Interpretability Shift
14:01 to 15:28
Explore how AI is transitioning from a black box to a machine we can understand.
“The research suggests that the architecture itself, those gated MLPs, actually selects for sparsity.”
Transcript
Automatic transcript. May contain errors.0:00You know, there's this persistent, almost haunting anxiety in the world of artificial intelligence right now. Oh, yeah. The black box problem. The black box problem. Exactly. We've built these massive digital brains models like LAMA or GPT. Yeah. And they can do incredible things. Write poetry, solve code, translate languages. Right. But if you actually grab an engineer and ask, how did it decide to write that specific word? The honest answer is usually we don't really know. It's unsettling. It's totally unsettling. We've essentially grown this digital brain in a jar. but we have very little insight into the actual thought process we see the input we see the output but the middle is just a fog and that's terrifying especially for safety so the industry has poured millions into decoding that fog and the gold standard right now is something called sparse auto encoders right essay ease they're the current heavyweight champion the prevailing wisdom is that you absolutely need them but today we're doing a deep dive into some new research that basically well it walks up to that heavyweight champion and punches it in the jaw.
1:04It really does. It just flips the table on the current consensus. The core claim is shocking in its simplicity. It's saying we don't need the expensive massive auto encoders. That the raw neurons inside the model, they were readable all along. We were just holding the map upside down. It's the bold claim. It suggests that the black box isn't as black as we thought. If you look in the right spot and use the right tool. So that is our mission for this deep dive. We're going to explore how we've been misreading AI, how a new method called RelP fixes it, and we are going to trace actual human-readable thoughts inside a llama model.
1:42From doing algebra to finding capital cities. And the implications are huge. Let's start with the before picture. Why did everyone think we needed these sparse autoencoders? Why couldn't we just look at a neuron and say, oh, that's the cat neuron? It comes down to a concept called polysemanticity. Okay, that's a mouthful. Break that down for me. It's basically a fancy way of saying multitasking from hell. The assumption was that models are crunched for space. They have way more concepts to learn about the world than they have neurons to hold them. So the model cheats. Cheats how? It forces a single neuron to represent cat, the stock market, and the color blue all at the same time.
2:20Wait, so if neuron number five fires, you don't know if the model is thinking about a kitten or a hedge fund. Exactly. It's a superposition. It's like packing a suitcase for a trip. A suitcase. Okay. If you have a small suitcase and a lot of stuff, you might stuff your socks inside your shoes and wrap your shoes in a T-shirt. You just cram it all in. And if I just look at the suitcase, it's a jumbled mess. That's what we thought neurons were like. And that mess is why researchers build Sparks autoencoders. Right. Think of an SAE as a massive, expensive microscope. Yeah. You take that messy neuron, that packed suitcase, and the SAE breaks it apart.
2:57It unpacks the suitcase. It lays out the socks, the shoes, and the t-shirts separately into these thousands of cleaner, sparse features. But that's computationally heavy, isn't it? You're basically building a second brain just to read the first brain. Extremely heavy. It's an enormous undertaking. A luxury tool, really. So along comes this new research. And they're arguing that this whole polysemantic mess is an illusion. Not an illusion, exactly. More like a misunderstanding of the architecture. sure. They argue we've been looking for clarity in the wrong part of the neuron. Okay, let's get into the anatomy here.
3:32We're talking about the MLP, the multi-layer perceptron. Correct. The MLP is where the model thinks. But an MLP isn't just one point. It has an input, a middle, and an output. And for years, everyone was analyzing the MLP outputs. The output. So the final signal that leaves the neuron. Yes. And here is the trap. In modern models, the internal layer, that middle part, is much, much wider than the output. Wider. It takes in a huge amount of information, processes it in this high-dimensional space, and then, and this is the key, it compresses it back down to send it out. Compresses it, like a zip file.
4:06That is a perfect analogy. That output is a zip file. It's dense. Every bit of space is used, which is why it looks so messy. And this research says, stop looking at the zip file. Exactly. Open the folder before it gets zipped. They argue we should look at the MLP activations. The activations. So the internal state of the neuron before that final compression. Before the down projection, yeah. Hold on. Why is that internal state any cleaner? Just because it's bigger doesn't mean it's organized. That is the key question. The answer is in something called the activation function. Modern models like LAMA3 use specific ones like SI-LU.
4:42SI-LU. We don't need to do the calculus, but what it does functionally is act like a bouncer. It takes negative values and just pushes them straight to zero. Okay, so it cuts out the noise. It creates what they call a privileged basis. In simple terms, it forces the model to be sparse naturally. It encourages individual neurons to represent specific, distinct things. If a concept isn't relevant, Sile Yu just zeroes it out. So inside the neuron, before the compression, the model actually is organized. It has a cat neuron and a finance neuron, but they get smashed together when they leave the room.
5:14That's the claim. And the stats are just wild. When they switched from looking at outputs to activations, the circuits they found were 100 times smaller. 100-dums. 100 times sparser. So instead of needing 10 ,000 mixed-up features to explain a grammar mistake, you might only need 100 specific clean neurons. That's a massive difference. It's not a tangled ball of yarn. It's a switchboard. A very, very complex switchboard. But yes, a switchboard where the individual switches actually mean something. Okay, so finding the right location, the activations, is step one. But the research makes a big deal about the measuring tape.
5:54Right. The problem with attribution. You have the output, say, the word Austin, and you want to know which neurons made that happen. And the standard tool for this has been integrated gradients, or IG. IG is the industry workhorse, but it has a major flaw. To figure out importance, IG takes the input, fades it to black, and measures how the output changes at every single step. That sounds slow. It is. It's like trying to figure out how a movie ends by watching it frame by frame, backwards and forwards, 20 times. It's expensive, and frankly, it can be noisy. Like trying to weigh a feather by driving a truck over a scale.
6:30Exactly. So this new approach adapts a tool called RelP Relevance Propagation. RelP. How's it better? It's much sharper. Instead of running the model over and over, it uses a linear approximation. It just asks, OK, given where we ended up, can I mathematically trace the signal flow backwards in one go? One go. So it's way more efficient. Just one backward pass. But, and this is the technical nugget that makes this work, they had to introduce something called the half rule. The half rule. OK, I saw this. It sounds like something from a card game, but I'm guessing it's math. It addresses a specific problem.
7:02Modern MLPs are gated. They multiply two signals together. like an 80 gate. So they're partners in creating the output? Right. Now if you're tracing relevance backwards from the output, how much credit do you give to signal A and how much to signal B? Ah, the classic group project problem. Who gets the grade? Exactly. If you're not careful you might double count the importance. The half rule is a simple mathematical correction that just divides the gradient by two at those intersections. That sounds suspiciously simple. Does that really work? It sounds simple, but it ensures what's called the conservation property.
7:37It means if the output has a score of 100 and you trace it back, the sum of all the neuron scores will exactly equal 100. It makes the math watertight. So let's recap the recipe. Stop looking at the compressed output. Look at the raw activations instead. Use RELP for speed and the half rule to handle the math. And when you combine all that, the gap vanishes. The performance difference between using raw neurons and those fancy, expensive SAE features just disappears. That is the headline. We can read the model's mind without rebuilding the brain. And to prove it, they didn't just run benchmarks.
8:11They mapped out a genuine thought process, the Texas circuit. I loved this section. It felt like watching an MRI of a thought. It's a beautiful example of what they call multi-hop reasoning. Okay, so here's the setup. They gave the LAMA-3 model a prompt. What is the capital of the state containing Dallas? Now, think about the logic chain there. The model can't just look up Dallas Capital. That's not a thing. Right. It has to do a two-step hop. Step one, Dallas is in Texas. Step two, the capital of Texas is Austin. Exactly. And using this new method, they found the specific neurons that do this.
8:4823 of them. No SAEs, no dictionaries. And these 23 neurons weren't just a random blob, right? Not at all. They clustered into six distinct groups that mapped perfectly to the logic we just described. Okay, let's walk through them. First, cluster one, the Dallas cluster. Simple recognition. These neurons just identify the entity Dallas. They light up when the city is mentioned. Then cluster two, the Texas cluster. This is the bridge. These neurons connect the city to the state. They represent the concept of Texas derived from Dallas. Then you had cluster three, the capital cluster. And this is so interesting.
9:22These neurons encode the abstract concept of a capital city. They aren't tied to Texas yet. They just know the question is asking for a capital. So it's holding the intent of the question. Precisely. And then finally you get cluster four. Say Austin. The execution. These neurons prep the actual word Austin for the output. Okay, mapping it is cool. But the real aha moment came when they started messing with the tracks. Steering, they call it. Steering is the ultimate test of causality. It's one thing to see a neuron fire. It's another to prove it caused the answer. Right. If you think a neuron means Texas and you turn it off.
10:01The model should lose the concept of Texas. So they did that. They suppressed the Texas cluster. What happened? Did the model just break? No. And this is the fascinating part. The model still knew it needed to name a capital. The capital neurons were still firing. But the bridge to Texas was broken. The bridge was broken. so it lost the specific state and it started guessing capitals of other states. It said things like Oklahoma City. Exactly. It answered the type of question correctly. It gave a capital city, but it got the logic wrong. This proves those neurons were holding the Texas concept.
10:33What about when they suppressed the, say, a capital cluster? In that case, the model might just output Texas. It knows the state, but it forgot it was supposed to provide the capital city. That is mind-blowing. It's like we can see the individual gears turning. Here's the geography gear. Here's the instruction following gear. And these are just raw neurons. It really validates the idea that these neurons represent distinct semantic concepts, statehood, capitalness. I have to play devil's advocate here, though. Geography is pretty structured. Does this hold up for something more algorithmic? Like math.
11:08That's the big question. Universality. So they tested it on arithmetic. And I always assumed LLMs were kind of bad at math, that they just memorized times tables from the Internet. They can be for complex stuff. But for simple addition, they have these surprising internal structures. They gave it problems like, what is 6 plus 7? And they found what they call modulo 10 neurons. They did. Okay, unpack modulo 10 for us. It's basically just the ones digit. So 6 plus 7 is 13. The ones digit is 3. Right. 16 plus 7 is 23. The ones digit is still 3. They found specific neurons that light up only when the answer ends in a 3 or a 5 or a 9.
11:46So the model isn't just memorizing. It built a circuit to handle the 1's column. Exactly. The data shows this perfect diagonal pattern on a heat map. It's visual proof that the model is performing a logical calculation. It has derived an algorithm. That feels very human. We learned to carry the 1. The model learned to carry the 1, too. It did. And the grammar examples are just as specific. They looked at subject-verb agreement. The tricky stuff. Like the keys on the table. The classic trap. Your brain wants to say, is because of table. But the subject is keys, so it has to be R. Correct. The model has to ignore table and look back to keys.
12:22And they found a neuron, L19N12056, that acts like the grammar police. What does it do? It essentially tracks plural subjects, but what's fascinating is that it fires negatively on singular ones. It's a literal toggle switch for plurality. And there was another one for phrases like Chen and Sandino. Right. It sees X and Y and flags it as a plural entity. This shows a deep structural understanding of grammar encoded in single specific neurons. So geography, math, grammar. The evidence is piling up. The neuron basis, looking at the ROM activations, it actually works. It works, and it works efficiently.
13:02So what does this mean for the future of AI? because it felt like we were hitting a wall where only Google or OpenAI could afford to look inside their own models. That is the why it matters here. It's about democratization. Democratization. If you need a massive cluster of GPUs to train a sparse autoencoder just to understand your model, safety becomes a luxury good. Only the biggest players can do deep audits. Right. The rest of us just had to trust them. But if you can trace these circuits using activations and REL-P, which is efficient, remember, just one backward pass, You can do this on a standard laptop.
13:35So an independent researcher, a grad student, could audit LAMA 3. Potentially, yes. They could check for bias in a hiring circuit or find a dangerous chemistry circuit. And because of that steering capability? We could intervene. If we find the build-a-bomb circuit, we could theoretically clamp those neurons to zero. Or if we find a bias circuit, we can identify and suppress those specific neurons. It feels like we're moving from AI as magic to AI as a machine we can take apart. Which is a very healthy transition. The research suggests that the architecture itself, those gated MLPs, actually selects for sparsity.
14:16So the model, it wants to be understood. In a way, yes. Gradient descent favored an internal structure that keeps concepts distinct. It's more efficient for the model to have one Texas neuron than to smear it across 5 ,000. We just needed to stop looking at the compressed outputs. So the clarity was coming from inside the neuron the whole time. Something like that. So to wrap this up, we started with the black box. We thought we needed a billion-dollar microscope to see inside. But it turns out the raw neurons, the activations, are speaking a language we can translate. We just had to learn to listen to them and use the right math relp to interpret it.
14:51It really makes you wonder, if raw neurons are this interpretable, If we can literally find the one's digit neuron, what other mysteries are hiding in plain sight? That is the question. We often fear AI is developing some alien intelligence that's totally beyond us. But this suggests the alien might be thinking in concepts remarkably similar to our own. Places, math rules, grammar structures. Maybe the barrier isn't that the AI is too complex. Maybe we've just been using the wrong measuring tape. I think that is a very distinct possibility. Something to think about the next time you ask a chatbot for the capital of Texas.
15:28Thanks for listening to The Deep Dive. See you next time.
From the publisher
This research explores mechanistic interpretability by tracing the internal computations of large language models like Llama 3.1 within their original neuron basis. The authors introduce RelP, a gradient-based attribution method that outperforms traditional techniques by using linear approximations to efficiently identify causal circuits. Their findings suggest that individual MLP neurons can represent interpretable, monosemantic features without the computational burden or errors associated with sparse autoencoders. Through various benchmarks, the study demonstrates that these neuron-level circuits are both faithful to model behavior and highly sparse. Practical applications are shown through steering experiments and the discovery of specialized circuits for tasks like multi-hop reasoning and multilingual translation. Ultimately, the work advocates for exhausting the interpretability of the original neuron basis as a robust alternative to learned dictionary methods.




