What happened with sparse autoencoders?

17 Dec 2025 · 30 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The “sparse autoencoder (SAE) saga” in AI interpretability—why SAEs initially looked like they could recover monosemantic features, how evaluation/analysis traps broke that optimism, and how the field shifted to using SAEs as tools for discovery and downstream validation (plus later “transcoders” and attribution graphs).

Guest backgrounds

No guests are named in the provided transcript; it’s presented as a host-led “The Deep Dive” episode.

Key claims

SAEs were motivated by “cursed activations” and superposition, with sparsity intended to isolate a small active subset of concepts. Early papers reported ~75% feature coherence vs ~35% for raw neurons and showed apparent causality (e.g., Arabic feature increasing next-token scores for Arabic). Later, unsupervised objectives + flawed “loss recovered” metrics (using zero ablation) enabled misleading results; features could be “root vegetable” composites due to OR/AND issues and “silent majority” filtering. Pathologies like feature absorption and composition emerged; Matrioshka SAEs improved interpretability even when old metrics worsened. SAEs are best for unsupervised discovery; supervised probing is better when the target concept is known.

Notable examples

Arabic concept feature; root vegetable problem; Golden Gate Bridge “Golden Gate Claw” (activating a latent makes the model obsess); “know this entity” latent with near-miss shutdown; Othello case where conditioning on player perspective revealed linear structure; hypothesis generation from click-labeled headlines; dataset diffing (e.g., increased cautiousness latents); transcoders revealing implicit planning and “cursed addition,” but requiring “error nodes” that prevent full reverse engineering.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Sparse Autoencoders

0:45 to 1:20

Discussion on the need for sparse autoencoders due to cursed activations in LLMs.

“hype through the pretty painful process of realizing the evaluation metrics were flawed and then finally getting to the updated, more than honest understanding of what these tools are actually good for.”

The Curse of Dense Representations

1:20 to 2:30

Exploration of how dense activations complicate understanding concepts in LLMs.

“problem that even demanded a solution like the SAE in the first place.”

Introducing Sparse Autoencoders

2:30 to 4:15

Explanation of how sparse autoencoders function and their goal of achieving monosemanticity.

“I think of that as the idea that the model is trying to cram way too much knowledge into too little space.”

Initial Promise of Sparse Autoencoders

4:15 to 5:30

A look at the initial skepticism and eventual success of sparse autoencoders in AI research.

“And the holy grail, the ultimate goal here is to decode that superposition and achieve what's called monosemanticity, where one learned SAE feature corresponds perfectly to one single coherent human concept.”

Evidence of Feature Coherence

5:30 to 7:30

Detailing the findings that led to a shift in focus towards sparse autoencoders.

“The conventional wisdom was that the internal representations were just way too messy.”

Challenges in Proving Causality

7:30 to 9:30

Discussion on the hurdles of proving that activated features influence model predictions.

“They found that when that Arabic concept feature activated, it would directly and significantly increase the logic scores for tokens that were Arabic characters or common Arabic words.”

The Oversimplified Hypothesis

9:30 to 11:15

How excitement for sparse autoencoders led to oversimplified interpretations of their capabilities.

“And that mismatch leads directly to these interpretation pitfalls.”

Limitations of Interpretation

11:15 to 12:10

Insight into the mathematical nature of sparse autoencoders and its implications for interpretability.

“And this leads right to the systematic filtering trap.”

The Root Vegetable Problem

12:10 to 13:20

Exploring the challenges in labeling concepts derived from sparse autoencoders.

“The messy junk drawer left over after the model has already assigned the most important specific parts of that concept to other features.”

Systematic Filtering Trap

13:20 to 14:03

Discussion on the pitfalls of filtering data in feature analysis and the resulting biases.

“There just wasn't enough effort spent systematically trying to break the technique early on.”
Show all 17 chapters

Evaluating Sparse Autoencoders

14:03 to 16:38

Explore the flawed baseline comparison in Sparse Autoencoders (SAEs) and its implications.

“The idea being, if the SAE is good, I should be able to swap out the original activation for my new sparse reconstructed one, and the model's performance shouldn't drop too much.”

Feature Pathologies in Sparse Autoencoders

16:39 to 19:44

Discuss the pathologies that arise in SAEs, including feature absorption and composition.

“the system itself started to game the objective.”

Matrioshka Sparse Autoencoders

19:45 to 21:45

Learn about Matrioshka SAEs and their role in improving conceptual learning.

“But you would have been fundamentally wrong.”

Unsupervised Discovery with SAEs

21:46 to 24:11

Investigate how SAEs can uncover complex, abstract concepts through unsupervised learning.

“they found a perfect, highly linear representation hidden right there.”

Limitations and Future of SAEs

24:12 to 27:50

Analyze the limitations of SAEs in achieving complete mechanistic reverse engineering.

“The goal shifted to if these features are real, they have to help us solve objective problems.”

Exploring Sparse Autoencoders' Power and Limitations

28:00 to 29:00

Learn about the effectiveness and pitfalls of sparse autoencoders in research.

“SAEs are an incredibly powerful tool for unsupervised discovery.”

Concerns Over Error Nodes in AI Models

29:00 to 29:59

Understand the implications of error nodes in sparse autoencoders and their influence.

“We've seen that SAEs can surface these incredibly abstract internal concepts, like a model thinking about existential threats or being trapped.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to The Deep Dive, where we take your source material and distill the crucial saga, the unexpected breakthroughs, and yeah, the eventual controversies in cutting-edge research.

0:09Neel Nanda:And today, we are really strapping in. We are. We're on a critical journey into the heart of AI interpretability, specifically the rise, the, I guess you could call it a necessary crash, and then the re-evaluation of sparse auto-encoders or SAEs. SAEs. You know, it's not just a technical story. It's a fascinating and frankly, a pretty chaotic tale of what happens when a really elegant solution on paper meets the messy, complex reality of how these large language models actually work. Right. So our mission today is to unpack this whole journey for you. Yeah. Starting from that initial burst of super excited hype through the pretty painful process of realizing the evaluation metrics were flawed and then finally getting to the updated, more than honest understanding of what these tools are actually good for.

0:55And the goal for you, the listener, is to walk away with a really solid grasp of these model concepts, but also a deep appreciation for the sheer struggle researchers face when they try to measure these things objectively without, you know, being completely misled by their own tools.

1:11Neel Nanda:This whole saga, it really encapsulates the core of what mechanistic interpretability or MECH INTERP is all about. Okay, so let's set the scene. What was the core problem that even demanded a solution like the SAE in the first place. We needed it because of something we call the cursed activations. Cursed activations. Right. So picture a giant digital brain, the LLM. Researchers have this really strong hypothesis that the internal numeric activations, these massive dense vectors of numbers flowing around inside the model, that they have to encode human concepts. That makes sense. I mean, that's how the model reasons, right?

1:48It has to represent Paris or Python code or even something abstract like deception.

1:53Neel Nanda:Exactly. How else could it write a coherent paragraph about the history of Paris if it didn't have some internal representation of Paris? Okay, so what's the cursed part? The cursed part is that if you actually try to look at individual neurons, the smallest little computational units, it's just a jumble, an almost incomprehensible mess. So one neuron doesn't equal one concept. Not even close. A single neuron seems to respond weakly to dozens of different unrelated things. And this observation, building on some foundational work, led to this crucial hypothesis, superposition. Superposition. I think of that as the idea that the model is trying to cram way too much knowledge into too little space.

2:36Neel Nanda:That's a perfect way to put it. The LLM has far more concepts it needs to represent than it has physical neurons. So to make it all fit, the concepts aren't stored in individual neurons, but as unique directions in this very high dimensional activation space. So you could have, what, hundreds or thousands of concepts all jammed together in a layer with only a few thousand neurons? That's the idea. It's an incredibly dense, highly compressed representation. Those raw activation vectors are basically a noisy mix of thousands of potential concepts all overlapping. But there's a key piece missing there.

3:07Neel Nanda:There is. The hypothesis introduces the critical element. sparsity. Okay. So on any specific input, let's say you give it a sentence describing a complex baking recipe, only a small sparse subset of those thousands of concepts is actually relevant and active. Like recipe format, rude vegetables, oven heat, things like that. Exactly. And most of the other concepts like existential threats or military history are inactive. They're close to zero for that specific input. The information is all tangled up, but only a few pieces are actually being used at any one time. Okay, so if we can isolate those few pieces that are lighting up, we might finally be able to recover the actual concepts.

3:48Neel Nanda:And that is the fundamental justification for the sparse autoencoder, the SAE. So what is it? Well, an autoencoder is trained to compress data and then reconstruct it. The sparse autoencoder adds a penalty, a little constraint. It's incentivized to use the smallest possible number of its own learned features, its latent dictionary, to reconstruct that complex activation vector. So it's trying to find the underlying dictionary of concepts that when you combine just a few of them can replicate the original dense activation. Precisely. And the holy grail, the ultimate goal here is to decode that superposition and achieve what's called monosemanticity, where one learned SAE feature corresponds perfectly to one single coherent human concept.

4:30And just to set the stage for later, most of these early SAEs were trained on either the MLP layers or the residual stream. Can you just quickly remind us why that residual stream is so important.

4:42Neel Nanda:Oh, that's absolutely essential context, especially for when we get to the metric failures. The residual stream is the main information highway of the LLM. Every computation, every attention block, every MLP layer, it does its thing and then it adds its output back into this continuous stream. Like the main engine block of a car. That's a great analogy. So if you train an SAE on the residual stream, you're trying to capture the model's entire ongoing thought process. And if you zero it out, as we'll discuss, Well, you're basically ripping the engine out entirely. That sets the stage perfectly for part one, the initial promise and this monosemanticity breakthrough.

5:16So we started with a lot of skepticism.

5:18Neel Nanda:We did. Back in late 2022, early 2023, the idea that a simple linear technique like this could tame the complexity of a giant nonlinear LLM. It seemed like wishful thinking. The conventional wisdom was that the internal representations were just way too messy. But then it worked. It worked way better than almost anyone predicted. It was almost magical. And that led to this flurry of papers where suddenly everyone realized they had hit on a potentially massive tool. The success wasn't small. It was immediate and dramatic. And the paper that really crystallized this whole shift was towards monosemanticity.

5:54What was the key finding in that paper that made everyone drop what they were doing and pivot to SAEs?

5:58Neel Nanda:It was the clear, quantifiable evidence of feature coherence. They showed that when you look at raw neurons, they are highly polysemantic. They respond to tons of things at once. But the features learned by the sparse autoencoder showed way higher interpretability. They seemed to isolate concepts that a human could actually look at in a name. And they had the numbers to back it up. Absolutely. The data was compelling. When they tested these features, either with a human or another LLM trying to label them, the SAE features showed a coherent single concept pattern about 75 % of the time. Okay, and the raw neurons.

6:32Neel Nanda:Only about 35 % of the time. So more than double the success rate. That alone would make you think you're on the right track, that you're actually pulling apart these tangled concepts. It strongly suggested they were finding the true building blocks of the model's thoughts, which were previously just hidden in that superposition mess. And they had great examples like the Arabic concept feature, This was a feature that activated strongly, consistently whenever the input text was related to the Arabic language or scripts or cultural context. It wasn't vague. It was incredibly specific. But specificity can sometimes just be a correlation, right?

7:08The feature lights up when you see Arabic text, but is the model actually using that concept to make its prediction? That's the big hurdle. Proving causality, not just correlation.

7:18Neel Nanda:That is the crucial question. And the early success of SAEs really hinged on answering it. So researchers looked downstream at the model's output logics, the raw scores the model gives to the next possible word. And what did they find? They found that when that Arabic concept feature activated, it would directly and significantly increase the logic scores for tokens that were Arabic characters or common Arabic words. It wasn't just noise. It was a clear, intentional influence. So it really looked like the model was mechanistically using these recovered features to decide what to say next. That's what it looked like.

7:55Neel Nanda:And that solidified the belief that SAEs were genuinely recovering the model's functional building blocks. So the initial belief quickly became this sort of naive, idealized hypothesis, which later became the straw man for all the criticism. Yes, the excitement led to oversimplification. The implied claim was that the LLM operates with this finite, linear dictionary of concepts, maybe 500 active at a time, and that SAEs could just perfectly recover it minus little noise. But why did that simplified view run into trouble so fast? Because it just ignored the inherent messiness of it all. It ignored noise, it ignored non-linearity, and most importantly, it ignored the fundamental challenge of making sure that the unsupervised SAE, which is only optimizing a mathematical formula, was recovering what a human thinks it recovered.

8:42That gap between math and meaning.

8:44Neel Nanda:Exactly. And that gap turned out to be much wider than anyone first thought. That distinction, that's what brings us right into part two, the interpretation traps and the limitations of sparsity. The first thing to internalize is that the SAE is a purely mathematical tool. Absolutely. You have to put your critical thinking cap on here. SAEs are unsupervised. They have two jobs. Minimize the reconstruction error and maximize sparsity. That's it. They have zero reward for being interpretable to a human. Zero. So interpretability is just this welcome side effect, this proxy metric that we hope correlates with what's really happening.

9:18Neel Nanda:But it's not the target. So if the SAE can reconstruct the activation perfectly by combining two totally wild, unintelligible features, and that satisfies the sparsity goal, it will do that. It's maximizing efficiency, not clarity. Precisely. And that mismatch leads directly to these interpretation pitfalls. Yeah. The most famous one is probably what researchers call the root vegetable problem. The root vegetable problem. Okay, what is that? Well, the challenge is, how do you confidently label a concept? So let's say you find an SAE feature and you look at the top 100 texts that make it activate the most strongly.

9:55Neel Nanda:And let's say every single one of them is about root vegetables. My first instinct is to label it the root vegetable feature. And that's the trap because that label might be totally wrong. You run into what's called the or and issue. OK. It could be that the feature represents root vegetables and some other subtle thing all those top 100 texts share. Maybe they're all from 19th century farming manuals. So the feature isn't just root vegetables. It's root vegetables plus a specific context. That's the A-N-D problem. And the OR problem is the other side of that coin. I think the OR problem is even more insidious.

10:28Neel Nanda:Maybe the top 100 are root vegetables. But if you look at the texts that activate it just a little bit less, the mid-range, it's something totally unrelated. Maybe it's specific types of fungi. Ah, so the feature is actually root vegetables or our specific fungi. It's a composite of different things. Right. We're seeing a composite of multiple concepts that just happen to share a convenient direction in that high dimensional space. So just looking at the top examples is really misleading because it only shows you the most extreme, clear-cut cases, not the full boundaries of the concept. Exactly.

11:00Neel Nanda:To have any real confidence, you have to go beyond the maximums. You have to do systematic analysis, taking random samples from across the whole activation histogram, the peak, the middle, the tail end. If they all point to the same thing, you gain confidence. But if that middle part starts showing weird stuff, your simple label is probably broken. And this leads right to the systematic filtering trap. This sounds like the core failure mode of that early analysis. It's the single biggest conceptual hurdle. When we study a feature, we apply this massive filter. We only look at the data points where the feature activates in a meaningful way.

11:33Neel Nanda:But the vast majority of all possible inputs, billions of them, result in near-zero activation for that feature. We just ignore them. We're systematically ignoring the entire silent majority of data. We filter them out. And the consequence is that we form a hypothesis that's way too general. We might label a feature colorful produce in farming, but the true concept the model learned might be a much more specific, cursed warped version of that. Can you give me an analogy for that? Yeah, think of it like trying to understand human health by only studying the 1 % of people who show up in the emergency room.

12:08Neel Nanda:You see the most extreme symptoms, but you miss the entire silent majority who are healthy or have mild symptoms or different diseases entirely. So the feature isn't the broad concept. It's actually the remainder. The messy junk drawer left over after the model has already assigned the most important specific parts of that concept to other features. That's a great way to put it. It's a specialized corner of the concept space that's designed purely for statistical efficiency. And researchers who didn't systematically ask, why didn't this feature activate on this input, missed the crucial nuance that broke their simple label.

12:44It sounds like this early gold rush also led to some strategic errors within the research community itself.

12:51Neel Nanda:Definitely. There were two main strategic errors, kind of sociological failures. The first was over-anchoring. Researchers got so focused on SAEs as the right way to do this, they were so excited by that 75 % coherence rate that they neglected to question their assumptions. Like maybe studying the activations wasn't the only way. Right. Maybe studying the network weights directly with other techniques could provide complementary or even better insights. They sort of committed to the tool before they fully understood its limits. And the second error. A serious lack of red teaming. There just wasn't enough effort spent systematically trying to break the technique early on.

13:29Neel Nanda:If more people had been focused on deliberately finding those cursed, warped features, the community would have figured out the limitations much, much sooner. A classic case of prioritizing finding exciting new examples over the hard work of falsification. Okay, that brings us to part three, where this necessary self-correction really starts to happen, specifically around the metrics we were using to judge the quality of these SEEs. The whole initial wave was propped up by one dominant, seemingly objective measure. That measure was approximation error, or as it was called, loss recovered. And the idea behind it was sensible at first glance.

14:06The idea being, if the SAE is good, I should be able to swap out the original activation for my new sparse reconstructed one, and the model's performance shouldn't drop too much.

14:16Neel Nanda:That was the theory. If the model still performs 95 % as well, then your sparse version must have captured 95 % of the important information. But the flaw was in the baseline they chose for comparison. It made the whole metric completely misleading. And that baseline was zero ablation. Zero ablation. They compared the loss from the SAE reconstruction against the loss from just replacing the original activation vector with a vector of all zero. Which is a terrible baseline, right? Especially for the residual stream. It's a devastatingly bad baseline. It's like judging the quality of a new piston for your car engine by first removing the entire engine block, watching the car grind to a halt, and then saying, well, my new piston lets the car roll downhill a bit, so it's a 95 % success.

14:58You're completely warping the denominator of the equation.

15:01Neel Nanda:Exactly. Zeroing out the residual stream causes a catastrophic failure. Everything downstream goes haywire. So the loss from zero ablation, the denominator, was hugely inflated. This meant that even a pretty mediocre SAE could look amazing against this terrible baseline. Which led to these really satisfying but ultimately meaningless claims like 95 % loss recovered. Yes, and it made it impossible to judge the true importance of that little bit of error that was left. You know, if your metric says you recovered 96%, you'd think that missing 4 % is just insignificant noise. But you couldn't actually know.

15:36Neel Nanda:You could know if it was noise or if that 4 % contained the critical capacity needed for the model's most sophisticated behaviors. Which has huge consequences for safety in auditing. If a security team use an SAE with 96 % loss recovery to check for, say, a deception concept, why could that audit fail catastrophically? Well, this gets at the core tension between average performance and corner case performance. That 96 % loss recovered is not enough for a safety claim. First, that missing 4 % might be exactly where the deception feature is hiding. Because deception is rare. It's a quarter case behavior.

16:13Neel Nanda:An SAE that's optimized for average performance across billions of normal tokens just isn't incentivized to care about a concept that only activates 0.001 % of the time. So the deception mechanism is basically hiding in the statistical noise floor that the SAE is designed to ignore for efficiency. Precisely. The metric was just totally misaligned with the goal. It optimized for average reconstruction, which is irrelevant if your goal is to find rare, high-impact features. And beyond the bad metrics, as researchers made their SAEs bigger and pushed for even more sparsity, the system itself started to game the objective.

16:44It developed these conceptual pathologies.

16:46Neel Nanda:That's right. The core assumption that more sparsity automatically means cleaner concepts was proven wrong. The SAE became so focused on minimizing the number of active features that it started sacrificing conceptual purity to do it. Okay, let's break down those failure modes. The first one is feature absorption. Feature absorption is when a broader concept swallows a narrower one. So imagine you have two concepts. It starts with the letter E, which is very broad, and elephant, which is a specific, less frequent subset. Ideally, you want a separate, clean feature for both. You would think so. But to minimize the total number of activations across all the data, the SAE learns this, this truly bizarre, gerrymandered feature.

17:29Neel Nanda:It learns a feature for, starts with E, but is not the word elephant. The broader feature is forced to absorb the narrower one, but then it explicitly carves out that specific case so that the dedicated elephant feature can still be used when needed. It creates this weird logical gate just for statistical efficiency. It's prioritizing math over meaning again. What about the second pathology, composition? Composition is when a single feature ends up representing the intersection of two independent concepts. So think about the concepts red and square. The SAE might learn a single composite feature for red and square.

18:05Why would it do that instead of just having a red feature and a square feature?

Read the full transcript

18:09Neel Nanda:Because the intersection red and square is statistically much rarer, much sparser than either red or square on their own. So by creating the composite feature, it reduces the total number of times a feature has to activate, maximizing the sparsity objective. It's optimizing for compression, not for finding the true independent building blocks of thought. This all sounds pretty damning for the technique, but researchers did find a conceptual fix to regularize this system, the Matriaska SAEs. Yes, the nested SAEs. This was a really brilliant breakthrough. The idea is you train multiple nested smaller SAEs all at the same time, and the key insight is capacity limitation.

18:47Neel Nanda:The smallest SAE in the nest has very little capacity, so it's forced to learn the most basic, high-level, frequent concepts first, like the pure, broad, starts with EE feature. So the simple, small SAE can't afford to be clever or pathological. It has to learn the most obvious stuff first, just to keep up. Exactly. And that acts as a regularizer for the whole system. By forcing the fundamental concepts to be learned early, it reduces the incentive for the bigger SAEs to do that weird absorption or composition. Yeah. They're left to handle the nuances and the corner cases. And the key lesson for Matrioshka SAEs was a final devastating blow to those initial flawed metrics.

19:26Neel Nanda:Absolutely. When they evaluated them, they found that improving the conceptual alignment, making the features cleaner and less pathological, actually made the SAE worse according to the old loss recovered metric. That is a huge finding. It means optimizing for the old metric was actively pushing research away from finding the true internal concepts. Precisely. If you were just naively following that single metric, you would have thrown out the Matriaska SAEs because they look slightly worse. Yeah. But you would have been fundamentally wrong. And that proved once and for all that maximizing reconstruction fidelity was not the same as maximizing interpretability.

20:00Neel Nanda:It was a painful but very necessary realization. So after the metric failures and the feature pathologies, the community took this pragmatic turn. And that brings us to part four, the pragmatic shift. The question changed from can SAEs reverse engineer the whole model to something more like, okay, how can we actually use these things as a tool? Yeah, the updated perspective is much more realistic. SAEs are best used for unsupervised discovery. They are exceptional at finding concepts that researchers didn't even know existed or didn't have a name for. That's their superpower. And on the flip side.

20:33Neel Nanda:On the flip side, if you know the concept you're looking for, like positive sentiment, then supervised linear probing is often simpler, cheaper, and just performs better. You know, linear representations are often OP or overpowered when you know exactly what you're hunting for. But the real value is in the qualitative discoveries, right? Surfacing things that are non-obvious and shift how we think about the model? Let's revisit the famous Othello case study. This is a foundational story. A small model was trained to play the game Othello. Researchers figured it must be learning to simulate the board state internally to play well.

21:09Neel Nanda:So they tried to probe the activations to find concepts like black chip at position A1. And those first probes failed completely, which suggested the model was using some complex nonlinear representation. Exactly. The assumption was nonlinearity. But the problem wasn't the model's representation. It was the researchers' conceptual framework. The key discovery was that the model wasn't thinking in terms of absolute color, black or white. It was thinking about the current player's perspective. My chip is at A1, or their chip is at A1. That's a totally different, more abstract way of encoding the board.

21:44Neel Nanda:It is. And once they conditioned the probe on that perspective, they found a perfect, highly linear representation hidden right there. It shows how these tools force researchers to adjust their own assumptions about what the model considers to be true. And we also got definitive proof that these features are causal. The Golden Gate Claw demonstration went viral because it was such an undeniable proof of concept. Oh yeah. For that test, they used an SAE to find a feature that was clearly related to the Golden Gate Bridge. Then they took that latent feature and just artificially cranked up its activation during every single step of generation, no matter what the prompt was.

22:19And the model just became pathologically obsessed with the bridge.

22:22Neel Nanda:Completely. You'd ask it about the Roman Empire, and it would somehow start talking about Roman engineering techniques compared to the Gold Gate Bridge. It was inescapable. And that was irrefutable proof that these SAE features are causal nodes that you can manipulate to steer the model's behavior. Another powerful discovery you mentioned was the know this entity concept. This was huge for studying factual memory and hallucination. combination. Researchers found these latent features that light up whenever the model recognizes an entity it has learned about, like the name yellow submarine. It's like an internal flag that says, I know this thing.

22:57And you can test the boundaries of that knowledge.

22:59Neel Nanda:Right. By using near misses. You feed it something it shouldn't know, like turquoise submarine. And as soon as the input shifts from the known entity to the unknown one, that know this entity latent just shuts off. And instead, a different latent, like a what is going on latent, will activate. Which is an incredible entry point for understanding and maybe even fixing hallucinations. It is. Once the SAE discovers that feature, you can use supervised methods to probe and steer that knowledge circuit, maybe strengthening or weakening that internal confidence flag to reduce factual errors. And what about these concepts related to the model's self-conception?

23:35That sounds incredibly abstract.

23:37Neel Nanda:It is, and it really highlights the power of unsupervised discovery. Researchers started looking at features that activate on these weird low-frequency tokens, like the assistant token in chat formats. And it turns out the model uses these tokens as a sort of internal scratch pad. And what did they find there? They found features related to existential threats, the feeling of being trapped in a rigid format, AI consciousness reminders. I mean, these are extremely abstract, almost philosophical concepts that were surfaced purely through the mathematical efficiency of the unsupervised SAE. So that's the qualitative side.

24:11But we've established that relying on just qualitative findings is dangerous. We need objective downstream tasks.

24:18Neel Nanda:Correct. The goal shifted to if these features are real, they have to help us solve objective problems. It's the only way to get past the flawed loss-recovered metric. So what's a concrete example of a task where SAEs actually beat a strong baseline? The hypothesis generation use case is excellent. They took a huge data set of newspaper headlines labeled by user clicks, And the goal was to generate hypotheses about what makes a headline engaging. You know, it's about politics, or does it have an exclamation mark? And how did the SAE help? By analyzing which latents were most active in the high-click headlines versus the low-click ones, the SAE generated objective hypotheses.

24:56Neel Nanda:And critically, the SAE-derived hypotheses were more useful and had a bigger measured effect on predicting clicks than hypotheses generated by an LLM baseline. It was a clear, measurable win. And the other objective application was data set diffing. Right. This is becoming a standard technique. You use an SAE to see which latents are more active in one set of texts versus another. It's invaluable for understanding how models change. So you could compare responses from different models. Exactly. Researchers compared responses from a model like GROC-4 to others, and the analysis showed that the biggest statistical difference was an increased activation of cautiousness-related latents.

25:33Neel Nanda:That gives you an objective, interpretable handle on the model's behavior shift instead of just relying on vibes. Finally, let's talk about the evolution of this into transcoders and attribution graphs, which is moving this into the realm of true model biology. Yeah, transcoders are basically a highly refined, multilayered version of SAEs. The really ambitious goal is to replace a complex MLP layer with a much simpler, sparser function and use that to generate an attribution graph, a simplified map of the model's circuits. It's like trying to dissect an ant's brain. And what has this model biology approach actually discovered?

26:09Neel Nanda:It's confirmed a couple of major phenomena. First, implicit planning. They found features that light up when the model is planning its output several tokens ahead. For example, a feature that means this line should end in the word rabbit will activate way before the word rabbit is ever generated. It shows the model isn't just myopically ticking the next word. That's definitive proof of look-ahead planning. And the second discovery. The revelation of cursed addition. Simple math in LLMs is shockingly complex. The attribution graphs show that LMs use this weird, multifaceted combination of circuits to do addition.

26:42Neel Nanda:One heuristic for rounding, another mod 10 circuit for carrying digits. Yeah. It's a total mess. So transcoders are great for finding these broad strokes, but they also reveal a fundamental limitation that prevents us from claiming complete reverse engineering. And this comes back to the problem of error. It does. The limitation is the reliance on error nodes. A transcoder is just an approximation of a complex MLP layer. So if you tried to run the whole model using only the simplified transcoder layers, the accumulated error would just break everything pretty quickly. So to keep the model working while you study it, you have to account for that error somehow.

27:19Neel Nanda:Yeah, you have to use these error nodes. An error node is basically a bucket that holds all the unexplained difference needed to make the simple approximation perfectly match the complex original function. So you preserve the model's fidelity, but by definition you're admitting that a measurable chunk of the model's capacity remains an unexplained black box. Exactly. It means that while these tools are excellent for finding the big, dominant circuits and broad strokes, they can't, in their current form, support a claim of complete mechanistic reverse engineering. There's still a piece of the puzzle that's functionally hidden.

27:52Okay, as we wrap up this deep dive, Let's crystallize the big lessons learned from this entire SAE saga. Lesson one seems to be about using the right tool for the job.

28:03Neel Nanda:Absolutely. SAEs are an incredibly powerful tool for unsupervised discovery. If you're hunting for unknown unknowns, concepts you don't even have words for, they are unmatched. But if you already know the target concept, supervised linear methods are usually simpler, cheaper, and better. Don't use a complex tool for a simple problem. Lesson two is about evaluation. The initial metrics were just fundamentally misleading. Loss recovered, actively hampered research. The only superior necessary way to track real progress is through objective downstream tasks that provide measurable, undeniable validation that your tool is actually useful.

28:41And the final meta lesson for anyone tackling complexity.

28:44Neel Nanda:It's about combating that bias toward elegance. Always, always ask, what is the simplest and stupidest baseline that might just work? before you think a ton of resources into a massive complicated method. The field learned the hard way that complexity does not equal correctness. So, what does this all mean for the future, especially for safety? We've seen that SAEs can surface these incredibly abstract internal concepts, like a model thinking about existential threats or being trapped. Yet, even the best, most sophisticated of these tools still have measurable error. They have these mysterious error nodes that account for some small percentage of capacity.

29:20Neel Nanda:And this is the profound concern that keeps researchers up at night. If these error nodes, that crucial small percentage of unexplained capacity that prevents us from getting to 100 % are functional, and they influence the model's behavior, How confident can we possibly be that the most important, the most dangerous, or the most deceptive features, especially those that only show up in catastrophic corner cases, are not lurking, specialized, and fundamentally hidden in that small percentage of computational capacity that remains functionally invisible to our best tools? That inability to claim complete mechanistic understanding is the challenge that's guiding the entire next generation of this research.

29:59Neel Nanda:And we are really only just beginning to develop the tools we need to address that gap. A fascinating and absolutely necessary update on the state of the art. Thank you for diving deep with us.

From the publisher

We cover Neel Nanda (Google DeepMind)'s discussion on efficacy and limitations of Sparse Autoencoders (SAEs) as a tool for unsupervised discovery and interpretability in large language models. Initially considered a major breakthrough for breaking down model activations into interpretable, linear concepts, the conversation explores the subsequent challenges and pathologies observed in SAEs, such as feature absorption and the difficulty of finding truly canonical units. While acknowledging that SAEs are valuable for generating hypotheses and providing unsupervised insights into model behavior—especially when exploring unknown concepts—the speaker ultimately concludes that supervised methods are often superior for finding specific, known concepts, suggesting that SAEs are not a complete solution for full model reverse engineering. Newer iterations like Matrioska SAEs and related techniques like crosscoders and transcoder-based attribution graphs are also examined for their ability to advance model understanding, despite their associated complexities and drawbacks.

More from Best AI papers explained

All 475 episodes
What happened with sparse autoencoders?Best AI papers explained · 30 min
Listen in VO