In short
Open problems in mechanistic interpretability (MI): how to reverse-engineer neural networks’ internal computations to enable trust, control, and scientific discovery despite “black box” behavior.
Guest backgrounds
No guests are named in the transcript; it’s presented as a host-led discussion.
Key claims
LLM abilities arise without explicit programming, so we must understand internal mechanisms to improve safety and reliability. Simple neuron-level explanations fail due to polysemanticity. Sparse dictionary learning (SDL)/sparse autoencoders can find latent features but are lossy, expensive, and may not map cleanly to human concepts; validation is vulnerable to “interpretability illusions.”
Notable examples
next-token prediction outperforming humans with tiny models; models using texture over shape; a small transformer trained on modular addition using a Fourier-transform strategy; GPT-4/SAE reconstruction causing large performance drops; refusal features appearing only when SAEs are trained on chat-style data.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Mechanistic Interpretability
0:45 to 2:18
Exploration of the black box issue in AI and the importance of mechanistic interpretability.
“The actual computational steps happening inside.”
Alien Intelligence and Its Implications
2:18 to 3:56
Discussion on how AI's cognitive processes can differ from human understanding.
“So, even if you're not directly working in MI right now, maybe you're building models, maybe deploying them, just understanding its goals, its challenges.”
Decomposing AI: The Challenge of Polysemanticity
3:56 to 5:00
The problems faced when attempting to decompose AI models and the concept of polysemanticity.
“A small transformer model trained on modular addition.”
Sparse Dictionary Learning: A Solution?
5:00 to 8:03
Introduction to Sparse Dictionary Learning and its potential to improve interpretability, along with its limitations.
“Well, that was the intuitive first thought, the sort of naive approach, you could say.”
Validation and Interpretability Illusions
8:03 to 10:30
Importance of validation in interpretability and the dangers of the interpretability illusion.
“So the sparse features are missing something important.”
Benchmarking Interpretability in AI
10:30 to 12:58
Need for benchmarks in interpretability and the challenges of real-world application.
“Making sure our explanations are actually correct.”
Real-World Implications of Mechanistic Interpretability
12:58 to 14:00
Exploration of the practical benefits of mechanistic interpretability for ML practitioners.
“Why should someone listening, an ML practitioner, really care about solving these?”
Understanding Mechanistic Interpretability
14:00 to 15:00
Learn how mechanistic interpretability can enable targeted interventions in AI models.
“what if you could directly steer the model by nudging its internal features?”
Predicting AI Behavior and Discovering Insights
15:00 to 16:00
Explore the potential of AI to reveal hidden patterns and predict behavior in novel situations.
“That predictability is key for high-stakes applications.”
MI's Societal Impact and Governance
16:00 to 17:00
Discuss how mechanistic interpretability intersects with AI governance and regulation.
“To understand if there are universal principles, if different architectures learn similar things, the universality hypothesis.”
Show all 12 chapters
Philosophical Debates in MI
17:00 to 18:14
Engage with the philosophical questions surrounding the goals of mechanistic interpretability.
“Yeah, there's this ongoing philosophical debate within the MI community itself.”
Future Challenges and Opportunities in AI
18:14 to 18:54
Examine the open problems in AI and the urgency of understanding complex cognitive processes.
“Okay, so wrapping this up, it feels like MI has incredible potential, right?”
Transcript
Automatic transcript. May contain errors.0:00Welcome to The Deep Dive, the show where we unpack complex topics to give you insights you can actually use. Today, we're digging into something really important, especially if you work with AI. The open problems in mechanistic interpretability. That's right. When you see the headlines, right? These large language models, they write stuff that sounds completely human. AI mastering games like Go. It's amazing. It really is. But here's the thing. The kicker, really, these abilities, they aren't things we explicitly programmed in. The AI learned them. Exactly. And that leads straight to this core problem, this massive black box issue.
0:38These systems are incredible performers, but, well, we often don't really get how they do what they do internally. We don't understand the nuts and bolts. Precisely. The actual computational steps happening inside. And look, this isn't just, you know, an academic puzzle. It's a real barrier. How can we trust these things? How can we control them? How can we even build better ones if we don't understand the mechanism? Which is exactly where mechanistic interpretability MI comes in. It's this field that's all about peeling back the layers, trying to understand the specific computational mechanisms that make these networks.
1:09Yeah. And it's a crucial distinction. It's not just about why a model made one specific decision, like why call that picture a cat? It's bigger. Right. It's about understanding how the model solves the general problem. Like, how does this model manage to identify cats across all sorts of different images, different lighting, different poses? we're essentially trying to reverse engineer the machine's mind. And the potential payoff here is huge, isn't it? Especially for you listening if you're an ML practitioner. Well, absolutely. Think about it getting real assurance about AI behavior, making safety something you can actually check, not just hope for.
1:45That would be incredible. And it goes beyond safety. We could unlock scientific insights, maybe use AI like a microscope to understand biology like protein folding, or just get better, more granular control over these systems. So that's our mission today. We want to give you a shortcut, basically, to get you up to speed on the cutting edge of MI. We'll look at the big open problems researchers are grappling with right now. And, crucially, why these problems matter to you and the work you might be doing. Yeah, we'll unpack where things stand and what needs to happen next. Yeah. Okay, let's get into it.
2:18So, even if you're not directly working in MI right now, maybe you're building models, maybe deploying them, just understanding its goals, its challenges. Yeah. It's becoming really vital. Definitely. Because the way these AI systems think, if you can call it that, it's often just fundamentally different from us. It's almost alien. Alien how? Can you give an example? Well, sure. Take next token prediction. You can have a model that's tiny, like 1 % the size of GPT-3, and it can actually beat humans at predicting the next word in a sequence. Wow. Okay. But then you take these huge state-of-the-art language models and they can still struggle with really basic stuff.
2:55Things a four-year-old gets like simple cause and effect with objects they've never seen before. So super smart in one way, surprisingly basic in another? Exactly. Or think about protein folding. Humans just can't do it reliably. AI cracked it. It shows this massive divergence in how intelligence can work. We just can't assume they follow our logic. That's a great point. And I heard something similar about image models. We think they see shapes like we do. Right. We prioritize shape. But often, these models rely much more on texture. They might recognize an elephant mostly by its wrinkly skin texture.
3:29Not its overall shape. Less so sometimes. Or even weirder, they might identify a fish because there are often human fingers holding it in the training photos. That's just a data set correlation, not, you know, understanding fishness. Okay, that's definitely not how a human would do it. Oh, not at all. And this raises that key question for you, the practitioner. If these things are thinking so differently, how can we possibly understand them? How can we make them reliable? There was another example you mentioned, something about math. Ah, yes. A small transformer model trained on modular addition.
4:00The researchers thought, okay, it'll probably learn a simple carryover method like we do in school. Makes sense. But nope. After the fact, digging into it, they found it had learned to use a Fourier transform strategy. Something totally unexpected, not designed. A Fourier transform. For addition. Yeah. It just goes to show we have to reverse engineer these things to grasp this potentially alien cognition and build safe systems. Okay, so reverse engineering, like taking apart an engine. You said there are three steps. Pretty much. First, decomposition. You break the thing down into its parts. What are the components?
4:35Got it. Second, description. You figure out what each part does, how they work together. You form hypotheses. Okay, makes sense. And third, crucially, validation. You test those hypotheses. Are they actually right? If not, you go back, refine your ideas about the parts or their functions. Right. Decomposition, description, validation. So when we try to decompose an AI, break it down, where do we start? Neurons. Well, that was the intuitive first thought, the sort of naive approach, you could say. Look at individual neurons or maybe attention heads and transformers, like the old neuron doctrine in neuroscience, one neuron, one concept.
5:11But that didn't quite pan out, did it? No, unfortunately. It runs smack into this big problem called polysemanticity. Poly what now? Polysemanticity. It just means one single neuron or one attention head often responds to lots of different seemingly unrelated things. Ah, like that light switch example you gave earlier. Turns on the kitchen light, the hose, and the car radio. Exactly like that. It's not cleanly separated. A single neuron might fire for, say, cat faces, but also syntax and code and maybe something totally abstract. It tells you the network isn't naturally structured along the lines of these components we designed.
5:50So if individual neurons are messy, how do we decompose it? How do we find the meaningful parts? Well, that's the million-dollar question, really. And it leads to the current leading method, which is called sparse dictionary learning, or SDL. Okay, SDL. What's the idea there? It's based on this really interesting idea called the superposition hypothesis. Superposition. Yeah. Like in quantum physics. Ha ha. Not quite. The idea here is simpler. A neural network can actually represent more features, more concepts than it physically has neurons or dimensions. How does that work? As long as for any given input, only a few of those features are active at the same time.
6:27They need to be sparse. Think of it like overlapping colors on a transparency. You can layer lots of colors, create new shades, but if you look closely, you can still pick out the original sparse colors. Okay, I think I get that. So SDL tries to find these sparse underlying features. Exactly. Methods like sparse autoencoders, SAEs, try to find these sparse features or latents hiding in the network's internal states, the activations. How? By training another network? Yeah, basically. You train a separate, small autoencoder network. Its job is to compress the internal activations of the main model and then reconstruct them.
7:03But here's the key bit. You penalize it heavily if it uses too many features from its dictionary to do the reconstruction. Ah, so you force it to be sparse. To pick just the essential features. Precisely. You give it a huge dictionary of potential features, way more than the original number of neurons, but force it to use only a handful for any given input. The hope is that these forced, sparse features are more meaningful, more interpretable. Okay. SDL sounds promising, then. Does it solve the problem? Well, it's a big step forward, but it's definitely not a silver bullet. There are some pretty significant open problems, things you as a practitioner really need to know about if you're thinking of using these.
7:41Like what? First off, reconstruction errors are often too high. If you take the activations reconstructed by the sparse dictionary and swap them back into the original model. The model's performance drops. Yeah, sometimes significantly. Like, there was a study with GPT-4 where using the SAE reconstructed activations made the model perform as poorly on language tasks as a model trained with only 10 % of the compute. Wow, that's a huge drop. So the sparse features are missing something important. They must be. For GPT-2 small, the drop was like 1040%. It tells you that what the SAE finds isn't the full picture of what the model is actually using.
8:17It's lossy. Okay, that's a big issue. What else? Well, these things are expensive, especially for large models. Training an SAE for every layer of a giant model, that can actually cost more compute than training the original layer itself. Ouch. So interpretability could cost more than the model itself. That's a tough sell for any team. It really is. And then there's the assumption that sparsity automatically means interpretability. That doesn't always hold true. How so? You run into weird things like feature splitting. One concept you care about, say cars, might get split into loads of different sparse features.
8:51Red cars, fast cars, cars in accidents. It fragments the concept. Making it hard to track the core idea. Right. Or the opposite, feature absorption, where maybe distinct concepts like cats and dogs get mushed together into a single feature. It just gets messy. It's not always a clean mapping. And even if you do find a feature, does SDL tell you how the network computed it? That's another major gap. SDL really only looks at the activations which features are on. It doesn't directly tell you about the mechanism, the weights and biases, the actual learned parameters that produced that activation.
9:27So you know what, but not how. Exactly. Getting from this feature fired to here's the circuit that made it fire often requires a ton of extra manual detective work. It's a huge bottleneck for actually understanding the computation. Yeah, that sounds painstaking. And one more thing. These latents don't always align with human concepts we expect or look for. The features SDL finds really depend on the data used to train the SAE itself. Did you give an example? Sure. Researchers found that SAEs trained on general web text data didn't really find good features for refusing harmful user requests. But when they trained SAEs specifically on chat-style data, then those features popped out.
10:07So the way you train the interpretability tool influences the concepts it finds. Precisely. It highlights this disconnect. The model might not be organizing its internal thoughts according to the concepts you think are important. Your search might be fundamentally misaligned with the model's internal structure. Okay, so decomposition's tricky. SDL helps, but has issues. What about the next step? Validation. Making sure our explanations are actually correct. This is super critical, because it's dangerously easy to fall for the interpretability illusion. The what now? The interpretability illusion.
10:43You find an explanation that sounds plausible, it seems to make sense, maybe it correlates with some outputs, but it's actually wrong. or misleading. So you can fool yourself easily. Very easily. People have shown you can find plausible sounding explanations for totally arbitrary directions in activation space or mistake simple data set correlations for actual understanding within the model. Conflating a hypothesis with a proving conclusion is a major pitfall. For you trying to understand a model plausible just isn't good enough. You need proof. So how do we get that proof? How do we add rigor?
11:16One key approach borrowed from biology is using model organisms. Like fruit flies. Exactly. Biologists study simple organisms like fruit flies or worms to understand fundamental principles that apply more broadly. In AI, we can do the same. We study smaller, simpler, open source models. Like that modular edition transformer or maybe GPT-2. Precisely. Models where we either know the ground truth of how they work, because we built them that way, or where it's feasible to figure it out completely. This lets researchers test and refine their interpretability tools, like SDL, with confidence before tackling the giant opaque models.
11:56Building confidence on simpler cases first. Makes sense. And related to that is the need for benchmarks. Just like you benchmark model performance on tasks like classification accuracy or perplexity, we need benchmarks for interpretability methods. How would that work? Well, you could create models where you know the internal algorithm. Maybe you built it from a known program. then you see if the interpretability method can actually recover that known program or its components. It gives you an objective score. Okay, so objective evaluation is key. But there's a danger here too. It's sometimes called streetlight interpretability.
12:27Looking where the light is best. Exactly. Focusing only on these simple model organisms or easy synthetic tasks can be misleading. Just because a method works on modular addition doesn't mean it scales to understanding, say, nuanced reasoning or potential deception in GPT-4. Right. The real world is messy. And the safety-critical problems are often in the complex models. So as practitioners, we need tools that bridge that gap, that work on the models you're actually building and deploying, not just the convenient lab examples. Okay, we've laid out a lot of the challenges, polysemanticity, SDL issues, validation hurdles.
13:02Let's pivot a bit. Why should someone listening, an ML practitioner, really care about solving these? What are the practical payoffs? Let's get into the so what. Right, because it's not just about academic curiosity. This has real-world implications for your work. Imagine you're deploying a critical AI system. MI could enable much better monitoring and auditing. You could potentially monitor the internals of the AI for specific red flags. Things like deceptive behavior. Maybe the AI is sandbagging, pretending to be less capable during tests. That sounds scary. It is. Or sycophancy, just telling you what it thinks you want to hear.
13:38or maybe detecting early signs of dangerous capabilities emerging before you deploy it. This goes way beyond just checking outputs. It's like having an internal probe. It gives you a much deeper level of assurance. That kind of white box testing would be huge for trust and safety. What else? How about more precise control? Instead of just fine-tuning outputs, what if you could directly steer the model by nudging its internal features? Like directly influencing its thoughts? Sort of. Or imagine needing to remove specific knowledge. Maybe copyrighted text memorized or dangerous information like how to build a weapon.
14:14MI could potentially allow surgical unlearning, removing just that bad knowledge while keeping everything else intact. Compared to current methods, which can be pretty blunt, right? Like retraining from scratch. Exactly. MI promises much finer grained interventions. It gives you more control over the how, not just the what. Okay. Better monitoring, finer control. What else? Predicting AI behavior in new situations? This is massive. Wouldn't you want to know how your model will react to inputs it's never seen before, especially in a critical production environment? Absolutely. Surprises are usually bad in production.
14:48Right. Or even predicting when new, perhaps unexpected capabilities might suddenly emerge during training. If you're training a giant model, knowing when it might suddenly develop sophisticated reasoning skills, lest you prepare, it's about moving towards more formal guarantees about behavior. That predictability is key for high-stakes applications. Definitely. And then there's this really exciting idea, microscope AI. Using AI to see things we can't. Yeah. Train a network on complex data, maybe scientific data, maybe even your own business data, and then use MI tools to understand what the AI learned.
15:22It might discover patterns or predictors that human experts totally missed. Like the AlphaZero chess example. Exactly. Extracting chess concepts from AlphaZero that grandmasters could learn from. or figuring out subtle facial features that influence human judgments. It turns your trained model into a discovery engine, and MI is the key to interpreting what it finds. That could be powerful for research, for business intelligence. For sure. And one last practical point. Most MI work so far focuses on CNNs and Transformers, but AI is evolving fast. We need to generalize these techniques to broader model architectures.
15:59Like diffusion models for images or state-space models. Exactly. To understand if there are universal principles, if different architectures learn similar things, the universality hypothesis. And practically, so the tools you might want to use aren't already obsolete for the models you're building next year. Okay, lots of potential practical benefits. Let's zoom out one last time the bigger picture stuff, societal things. Right, because MI isn't just a technical field. It touches on governance, even philosophy. How does it connect to policy? Well, think about AI governance and regulation. How do you verify if a complex model complies with safety standards?
16:34Right now, it's mostly black box testing. Which might miss things. MI could offer technical tools for verification, maybe detecting copyrighted data inside a generative model, or auditing for specific biases by looking at internal representations. It could provide regulators with concrete, explainable evidence for risk assessment. It could also help satisfy right-to-explanation requirements, like under GDPR in the EU. So technical levers for policy. Interesting. And the philosophical side. Yeah, there's this ongoing philosophical debate within the MI community itself. What's the main goal? Is it pure understanding, like fundamental science knowledge for its own sake?
17:11Or is it more like engineering? Right, where every insight needs a practical application, a way to control or improve the AI. There's this critique that sometimes MI gets graded on its own curve, focused on puzzles that might not solve the most pressing real-world problems, which maybe have simply black box fixes. Finding the right balance between deep understanding and practical utility. Exactly. What does it really mean to understand an artificial mind? It's a deep question. And that leads to a really important point about communication, doesn't it? Absolutely critical. We need caution in communication.
17:46It is vital that researchers, and all of us really, avoid overstating where am I is today. Why is that so important? Because if we claim the black box problem is solved when it isn't, that could be misused. Imagine lobbyists using premature claims to argue against needed AI safety regulations. We have to be honest about the limitations, the open problems we've discussed, alongside the progress. Ground claims and solid validation. Be realistic and rigorous. Okay, so wrapping this up, it feels like MI has incredible potential, right? Safety, control, discovery. Huge potential. But it's also clear there are major hurdles, really fascinating open problems still to solve.
18:26Problems that honestly need smart people like you, the listeners working in ML, to help tackle. And as AI gets more powerful, the need to understand how it works just gets more urgent. Definitely. It's a really exciting and frankly crucial frontier right now. So maybe the thought to leave you with is this. If AI is developing these complex cognitive processes we don't yet grasp, what other surprising problems or maybe even amazing insights are waiting for us inside those black boxes? What will the next deep dive reveal?
From the publisher
This paper gives a comprehensive review of the **open problems** and future directions within the field of **mechanistic interpretability** (MI), which seeks to understand the computational mechanisms of neural networks. The authors organize these challenges into three main categories: **methodological and foundational problems**, such as improving decomposition techniques like Sparse Dictionary Learning (SDL) and validating causal explanations; **application-focused problems**, which include leveraging MI for better AI monitoring, control, prediction, and scientific discovery ("microscope AI"); and **socio-technical problems**, concerning the translation of technical progress into effective AI policy and governance. Ultimately, the review argues that significant progress on these open questions is necessary to realize the potential benefits of MI, particularly in ensuring the safety and reliability of advanced AI systems.




