In short
The TWIML AI Podcast - Episode #727: Exploring the Biology of LLMs with Circuit Tracing
Episode Overview In this episode, host Sam Charrington speaks with Emmanuel Ameisen, a research engineer at Anthropic, about their recent work on mechanistic interpretability methods for large language models (LLMs). The conversation centers around two significant papers: "Circuit Tracing: Revealing Language Model Computational Graphs" and "On the Biology of a Large Language Model." The discussions delve into how the team developed methods to understand LLMs internally by replacing dense neural network components with sparse, interpretable alternatives. Key discoveries include insights on LLM capabilities, limitations, and safety strategies.
Key Concepts Discussed
Mechanistic Interpretability
- Introduction to Mechanistic Interpretability: This approach aims to understand how LLMs operate internally.
- Circuit Tracing: The technique developed to visualize and comprehend the computational graphs of LLMs, allowing insights into their functioning.
Key Findings
- Planning in Poetry Generation: LLMs can plan ahead when crafting poetry, as demonstrated by the example where the model chooses the word "rabbit" before completing the associated sentence.
- Mathematical Calculations: The models have specific algorithms for performing mathematical operations, revealing sophisticated internal processes.
- Multilingual Processing: LLMs manage to process and represent concepts across multiple languages with shared neural representations.
Neural Pathways and Manipulation
- The team can intervene in model behavior through specific neural pathways, showcasing how concepts are distributed across the network's MLPs and attention mechanisms.
- Behavior Manipulation: By manipulating certain pathways, researchers can influence the model’s output, which illustrates the model’s internal workings and its limitations.
Hallucinations and Recognition Circuits
- Hallucinations occur when the model uses separate recognition and recall circuits, leading to outputs that may not align with reality.
- Chain-of-thought explanations may not accurately reflect the model's reasoning process.
Methodologies
- Circuit Tracing Methodology:
- Uses transcoders to replace computation in the model with a sparse, understandable version.
- Enables the mapping of features and concepts, connecting how LLMs arrive at certain outputs.
- Sparse Coding: Allows the decoding of dense representations in the models to identify distinct concepts effectively.
Limitations of the Approach
- Attention Mechanism: The focus on MLP layers means that the intricate dynamics of the attention mechanism are not fully explored.
- Polysemanticity and Superposition: The complexity of concepts encoded in the model leads to challenges in understanding how these concepts interact and influence outputs.
Implications for Future Work
- Safety and Alignment Strategies: Understanding the model's internal mechanisms can help in creating models that align better with human values and safety protocols.
- Global Understanding of Mechanisms: Future research aims to extrapolate findings to develop a comprehensive understanding of how models operate across various scenarios, leading to better predictions of behavior in unseen contexts.
Conclusion The conversation underscores the importance of mechanistic interpretability in advancing our understanding of large language models. The findings from the discussed papers reveal not only the sophisticated capabilities of these models but also their limitations and potential areas for improvement in future iterations. The insights gained can significantly contribute to the development of safer and more reliable AI systems.
For complete show notes and additional resources, visit [TWIML AI Podcast Episode #727](https://twimlai.com/go/727).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We've had debates for a very long time where it's like, are they stochastic parrots? Like, do they, you know, just like kind of like match to the most like similar training data? Like, or do they like generalize? What is generalization? What does that even mean? And I think what's really cool about this work, what really like gets me excited is we can just talk about the mechanism and then we can debate like, you know, what does it mean that this is the mechanism? But we can at least start from a common ground and be like.
0:41All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Emmanuel Amazin. Emmanuel is a research engineer at Anthropic. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Emmanuel, welcome back to the podcast. Thanks, Sam. It's good to be back. It's been a while. It has been a bit. It's been three years, almost like in a couple of weeks. It'll be three years since we recorded that last episode. And we were chatting a little bit beforehand. You know, we know one another personally, and we can't remember the last time we saw each other.
1:21It's been too long, like 10 lifetimes in AI years, I think. Yeah, I was going to say, it's okay. Nothing happened in the last three years. There hasn't been any development that matters. Nothing at all. Nothing at all. You've been at Anthropic how long now? For about two years. It'll be two years in a month. Okay. So it was definitely longer than a couple of years. I think you were probably still at Insight the last time we spoke. And I think you did a stint at Stripe maybe after that? Yeah, that's right. So it must have been a while back. I think when I was last on, I might have just started at Stripe.
1:59I was there for about three and a half years or three years or so. Yeah, so doing kind of machine learning engineering there more probably. And at Anthropic, have you been exclusively focused on interpretability? I've actually done a bunch of things. When I joined, I initially joined the fine-tuning team. And so, you know, my first few projects were actually around making Claude better at a variety of kind of very practical tasks. So I worked on making Claude better at SQL, like just writing good SQL or, you know, being better at function calling and tool use, kind of the sort of like in the early days of agentic workflows, just just having models called tools was that was a big deal.
2:52So sort of like working on some of some of that stuff. So internal fine tuning as opposed to allowing users to bring their own data and fine tune. yeah yeah to be yeah i should have said fine-tuning here is sort of like fine-tuning in the sense of the fine-tuning that goes into the cloud model that goes on like cloud.ai and on our api so it's the sort of like you know the process by which we make that sort of like main model that everyone uses rather than yeah rather than like fine-tuning for for custom use cases and so you were saying after fine-tuning yeah so i worked on that for a while um kind of yeah worked on on a variety of things around like models using code and models kind of like using tools uh and then about a year ago uh i switched to to something that's um much more research focused working on interpretability research which is um kind of trying to understand how the models do you know what they do at all um if you're familiar with the field you know you're like ah yes this makes sense i think when i talk to friends they're very confused uh they're like but you made the model like what do you mean you don't understand how it works but it turns out we have no idea how it works uh at all and so and so we have a whole team uh and multiple teams actually just trying to try to understand that and you know there's also other groups in the industry and in academia doing the same thing just trying to understand well like how are these models doing what they're doing and our conversation today is going to focus on a couple of papers that you and your team have recently published.
4:27The first one is called Circuit Tracing, Revealing Language Model Computational Graphs. And the next is called On the Biology of a Large Language Model. I think of those two and the way they fit together as the circuit tracing paper is kind of talking about some of the tooling and mechanisms that you've developed for this particular approach at mechanistic interpretability and the on biology paper as being some of the results you've seen in the application of those tools. Is that the way you think about it? Yeah, that's exactly right. Maybe a metaphor we've used here that I think makes sense is in a way, like some of the work we've done is akin to like building a microscope of sorts, you know, and like at least some tool that allows us to look inside the model a little bit as it does things.
5:18And so there's the, well, how do you build a microscope? It's like, ah, well, you know, we chose like to sort of like, this is actually like how light refracts off of the lens in that other one. And that's the like circuit tracing paper. And we talk a lot about, you know, the various pitfalls, the various like trade-offs we've made, that sort of stuff. And then there's, okay, well now you have a microscope. Let's look at some cells. And that's the biology one. And this one is mostly, you know, okay, the insights here are less about like how we got there we of course use the microscope to get there but they're about like oh it looks like you know this is how Claude does math or it looks like this is how like this is the mechanism that causes Claude to hallucinate sometimes so it's much more about like the yeah the biology of a model like studying its behavior rather than sort of like the building of the tool.
6:08And we're going to dig into both of those areas but I wonder if it makes sense to start with like what is the thing that you learned in all this about the models that you found kind of the most surprising or that most challenged the way you think about large language models it's funny um after like before releasing the papers but after doing some of the experiments uh someone on our team like asked everyone you know kind of uh we have 10 or so uh case studies and biology and you're like well how many of those were were like surprising to you um and and reflecting you know back on it and on what i was expecting i would say that about like half to two thirds of them were surprising to me um and then of course you know you you sort of like published the paper and some of the reaction to to all of them is like well like this is obvious like why why wouldn't the model work in this way you know uh from other people so like i think sometimes it can be harder to appreciate uh what's surprising but but i can give you like maybe a couple ones that that definitely were surprising to me um one is is um this like poetry example and we use poetry because um writing a rhyming poem is something that's like pretty hard to do um if you're just like predicting one word at a time which is how these models work right uh they like have a prompt They predict one word, they add that word to their problem, they predict the next one and so on.
7:37And so you might say like, well, OK, that seems like it would make it really hard to write, you know, good poetry, which most most models can do, or if not good, let's say rhyming poetry. You'd expect the model to greedy optimize itself into a corner pretty easily. Exactly. It would get to the end of the rhyme and then it'd be like, oh, you know, now I have to rhyme with orange. It's like, well, it turns out actually that's a hard rhyme. And also I didn't think about it. But still, you can imagine like, well, maybe if you have enough data, it just kind of like always picks, you know, sentences that kind of make sense and it gets to the rhyme and there's some like murky thing.
8:12But it turns out that what we found when we looked at it was that when this is a cloud 3.5 haiku, like is given a rhyming couplet. The example we used in the paper is he saw a carrot and had to grab it. The model then completes that with his hunger was like a starving rabbit. So it rhymes using rabbit, which is fine. uh it turns out that when it when it processes grab it and then comma and then new line and then it's about to write the second line at that point before it writes the first word of the second line there are uh features which we can explain kind of what they are a bit but there's like representations in the model's circuits for the word rabbit um it's already decided uh in a way if you want to decide it, that it's going to end with Rabbit.
9:07And then it uses those representations not just to like choose Rabbit at the end. It's not just that it's chosen that. It's that they're also causally involved. Like they help guide the choice of all of the words before leading to Rabbit. And we know this because we see it in the circuit, but also because we kind of want to validate this outside of our methods. We just perform a bunch of tests where we like delete Rabbit and we add like a random word in the paperweight like green and then the model rewrites the sentence completely and i think i think the sentence it lands on is something like the farmer's field was verdant green or something so it like completely changes its sentence so that it can end with its plan but so that example of sort of like pretty complex you know longer horizon behavior was definitely one that kind of like stuck with me uh as something that you might not expect if you were just like well they're just you know autocomplete on steroids predicting that the next token like they're sort of like not doing anything complicated.
10:02It turns out they are in some context. And I think with that as kind of context, it makes sense to jump into circuit tracing because the immediate question, you know, on hearing that example and the others presented in the work is like, how do you know? How do you know? So let's jump into circuit tracing to talk a little bit about or contextualize the approach that you've taken in the broader context of interpretability? And maybe to get that started, are you only focused on mechanistic interpretability there or do you explore other types of interpretability research? Right now, we're mainly focused on, or at least for this work, we're mainly focused on mechanistic interpretability.
10:52And I can sort of like, yeah, maybe i can contextualize kind of what even is circuit tracing and also maybe like why do you need it uh just really briefly right which is that you know um in in past jobs doing doing machine learning i've used models that weren't uh kind of transformers and and models like that that were maybe like decision trees and decision trees are commonly used in like fintech and fraud prevention and stuff like that and they have this really nice property that they're exactly what they sound like they're this tree of decisions and so once you've trained your model you can kind of look at it and say like ah like the reason that the model said you know we should block this transaction is because the amount was over ten thousand dollars and it was going to like a bank account in the like cayman islands or something and you're like yeah that's that makes sense whereas if it says you know if the if the like individual branches made less sense you'd be a little more suspicious.
11:51And so that's a really nice property of these silver models. Turns out transformers, not at all like that. The complete opposite, right? The way that they work is they have these kind of like giant vectors of numbers of hundreds or thousands of numbers that represent each word or each token. And then those get processed by like a series of, again, you know, basically oftentimes millions, billions of computations. And at the end, the model predicts a work. And kind of like what happens in between, if you if you look at like all of these all of these numbers is just like completely unclear, like you can, you can, you know, go in your your like Jupiter notebook, and you can look at the numbers and you can say, ah, like, you know, yes, when when Claude was like writing poetry, this is the sequence of like a billion numbers or whatever that that was in its in its activations, but that's sort of like, so what, you know, it's it's maybe like, yeah, similar to sort of like, looking at like a tangled mess of like cables or circuitry inside a computer and it's like well you could just look at what lights up but you kind of like you need to somehow like map your observation to like what the like what the model is actually thinking about and so that's the nature of the problem um and then i would say that the sort of like circuit tracing is step step two on maybe like a like a three-step plan step one would be um well can you understand like what the model is representing at all uh and this is this is kind of uh the reason for working on dictionary learning uh which which um was kind of the topic of of the two previous papers before these two you said dictionary learning yes yeah um which is like broadly here like like the idea is can you take this like dense mess of numbers inside the model when it's like let's say processing one word or one one sentence and decompose it in a bunch of like unique concepts so that you know if if the the model is processing the problem like uh i'm taking my car over the golden gate bridge you might expect that like okay so you see this massive numbers but in there's going to be like a concept for the golden gate bridge and a concept for car and a concept for like a fun weekend with friends or something um and so the goal of dictionary learning is is is that it's like okay can you basically decompose representations into a like uh sparse version that's sort of like understandable um and we do we do that importantly like an unsupervised manner basically like you train a completely other model that's going to be your your dictionary that's going to like decode the innards of your your base model and then you can like on a given prompt you can look and you say oh you know look claude was thinking about this or thinking about that and that was that was the paper that that that that was prior to these most recent ones and sort of we found a bunch of interesting things there we found you know kind of like very nuanced representations there's like representations for like people uh uh being like psychophantic you know meaning like praising you even when they don't meaning there's like they don't mean it there's like representations for like typos and bugs in code there's like a bunch of different representations inside the model that are that are like fascinating um and so then and then i'll pause maybe but i'll say one more thing which i'll say like okay that's the first step and then so what's the second step well what you'd really want is like this is kind of like finding you know you have this like jumbled soup now of all of the things that the model is thinking about but what you don't have is you don't have a way to explain like okay but why did the model say what it said or even why did the model think about the golden gate bridge like it's probably because in the prompt there was like golden gate bridge and it like those three words are enough to think about that versus other bridges but now we want to connect all these representations and have basically like a like like a circuit which is just like a list of steps in order where it's like it went from this one to that one to this one to that one and then that's why it said this and and the effort of circuit tracing is not nice and thinking about the uh the dictionary work and the idea of representations of concepts uh one of the things that jumped out to me is that this uh you know it's very similar to the ideas of like embedding embedding you know uh words into vector space and you have these areas that represent distinct concepts, but it's not exactly embedding.
16:35It's more, maybe the right way to ask the question is, does the kind of results that you see, like follow directly from embedding and embedding as a concept, embedding this and the fact that models kind of operate on these vectors, or is there a distinct concept of concept relatedness in the context of the model itself it's definitely related so you know the models themselves right they take in as you mentioned embeddings for the for for their words and then but but then what happens is sort of like two things um they process these embeddings and and like output you know different different vectors essentially and so it's like ah like maybe uh if you so an example we have in the paper is like you know facts uh michael jordan plays the sport of and then you know uh basketball is the completion that the model successfully gets to and it's like ah well how does it know this well you know one thing that that models do is like they take the embedding for jordan which is like that's just a token they've learned it and then there's like various pieces of the model that like read this embedding and will like for example write like ah basketball or ah um space jam you know or the chicago bowls and so there's they're sort of like um these these like um embeddings are your starting point but then you have like a lot more nuance about like various sort of like either like your various aspects of like decomposing a given embedding something or like combining them right if you have you know i don't know uh michael jordan and lebron then maybe you're going to have like some parts of the model that are combined though combining those and sort of like starting to discuss starting to have features for like discussing who's the greatest basketball of all time right like debates about greatness is probably a feature you'll see there and so they're sort of like like the way to think about it is that like the nuances of concepts that can be represented in the model is is is much it's almost like a commentorial explosion on top of your embeddings right you you know, like you're kind of like base embeddings.
18:45And then, you know, you can kind of like combine them in various ways. And then, you know, the way you combine them is you might have like an embedding for like street and then an embedding for like, you know, I don't know, main street, first street, second street, third street. But then you also are going to have a direction in the model that's like, you know, main street in like Texas maybe or main street in Arizona. You can sort of like combine them in arbitrary ways. And so what we're trying to understand are sort of these like more complex and nuanced representations inside the model after it's consumed embeddings and before its output um uh basically a next word prediction yeah i'll tell you where that came up for me was the example about um was trying to get at whether there's this kind of linguistic universality.
19:36And so it had the 3.5 haiku model like translating the same phrase in several different languages and asked it to get the opposite of like convert big to small or something like that. Yeah, the opposite of biggest, yeah. Yeah, and it struck me in a lot of ways like the classic word to vet kind of relationship, king minus man plus woman. The question that that prompted for me was like, you know, do these results kind of just follow directly from embedding? And I like the insight that, you know, yes, they're related, but the concept of embedding within the model is so much richer than the kinds of concepts that are represented in the vector space.
20:25Yeah, and I think that like, those results are actually like a really good example of this. And we can like dig in a little more in that like so the comparison we make in the biology paper on this is like if you do that's i think it's the opposite of large is small uh and uh le contraire de grand et petit which is like the french version and you can do like we do different languages what you find is that like you know initially there's like of course different embeddings for those because they're they're different uh tokens uh and then in you know as as the as the sort of like signal uh progresses through the model the features so these representations start being shared like there's a shared representation of something that is big there's a shared representation of something that is the opposite across all languages um and and that i think kind of is an example of what i was saying a little bit earlier which is like ah like all of the embeddings for like big and large and all of the other languages they they share at least one facet right which is that they're about something that's big there's many other facets that they don't share like one of them starts with b the other starts with l you know they have like different number of letters whatever there's like a bunch of ways in which they're different they're in different languages um but they but they share that and so the model sort of like you know in processing these different languages like extracts the shared representation um and then does its computation it's like ah like there's largest there's opposite so i have to say something small and then you know goes back to the language that it was asked in uh and and so in the paper right we're like well if it's truly a um a shared like um uh kind of like set of representations surely it means that you can kind of like mess with them and like plug one for the other or like you can you can use the same like flip large to small in the like kind of like shared space and it will do the same thing for all languages and surely it does meaning that like we don't need to find a different concept of like opposite in french and chinese and english it's just we we found the concept of opposite uh for the model here and we turn it on it works for all languages um which yeah so so again it's like an aspect of it's like something more nuanced than the original embedding right it's like extracted some some like aspects that is that is that is shared and it is useful for it to sort of like generalize when it's learning across languages, right?
22:48That's also an important implication here is that like that allows you to learn something in a French text and then be able to like apply that same circuit that you learned in texts of other languages. So getting back to how this works and how the kind of tooling and methodologies you've developed to, you know, understand these models, a couple of key concepts are, you know, replacement model and sparse coding. Can you talk about, you know, the mechanisms a bit? Maybe I'll start with source coding because,
23:28so the idea is actually similar both for this new paper and for the sort of like dictionary learning papers. It's the same approach basically. And the idea is like, ah, you have, you know, say in the middle of the model, you have this like, as we said, it's like a very dense, you know, vector of numbers and and in that vector you're pretty sure you know are all these representations there's like the representative large the representation of like opposite but all you see is this like dense vector of numbers and what you want is you want to know ah like actually the model is thinking a little about this concept a and all about this concept b and all about this concept c so like this is just a method to get at that and the way that you get at this method is you say cool can i learn another a separate model that will take this like incomprehensible mass distance vector and that will learn a like essentially um you do like a matrix multiply to a like much larger space um and and then you uh take this space and that space you force to be very sparse and so you literally like have you know if you have let's say for example like a 100 neurons somewhere in the middle of your model so you have 100 numbers that are kind of like always active and like very hard to comprehend you blow it up to a thousand numbers and you have a constraint that says that like basically in your loss you penalize every time one of those numbers is not zero so the model is going to have to be like very very parsimonious in allocating these these features it's not going to be able to turn many of them on and and then so so that's like where the sparse comes from and that but then you have to make sure that you're actually capturing as much as you can so then what you do is you try to reconstruct the original thing from that's that code and so it's like you're learning a sparse code and then you're also like making sure that your sparse code is enough to capture like as much of the original uh thing as possible and so that's basically like the gist of dictionary learning you're training with this this mix of like sparsity and reconstruction error where it's like how how well can you reconstruct after you've blown it up and put it back together um and that that is sort of like what gives you these um these features which we which we call like monosematic meaning they kind of like represent only one thing um so that's sparse coding so okay so that that you could you you could do before and we did before for sort of like uh representations inside the model um maybe at this point like a useful mental model to have is is like i have it as like a tiramisu where you have like the input tokens at the bottom of the sentence, you know, so it's like, ah, the, you know, facts Michael Jordan plays the sport of is like at the bottom.
26:09And then the model has like successive layers where it's like, it transforms this into numbers and then it does some computation, basically like some math, okay, some matrix crossplications. And it has another set of vectors, one for each of the tokens. And then it does that again, it does that again, it does that again, it does that again. And eventually, you know, it does a final uh multiplication to still do the output and so what we're what we're doing what we're doing in the previous papers is like applying this method in between computations and seeing like okay it's a bunch of computations we're somewhere in the middle let's look at what the model is thinking about now basically um the problem with that is that you you can see what the model is thinking about but you kind of don't know what the computation is doing uh at all um you just you could see it to the extent that you're like well before it was thinking about you know uh bridge and gold and now it's thinking about the gold gate bridge and so probably like somewhere like there's some part of the model that like stitches together to understand this and so in circuit tracing we uh we use uh transcoders which uh basically what what those do is they do the same thing but you um you try to replace the computation with this sparse thing instead of just a representation so we're now just like completely replacing bits of the model with this like sparse thing um and we're trying to make it so that the computation stays the same but it's super sparse and understandable um and that's that's like the main idea of the replacement model is like you try to replace chunks of the base model with another model that's as close as possible while being much more explainable understandable in doing so you have several different choices for the type of replacement model you use cross-layer transcoders versus SAEs versus others can you talk a little bit about why you chose the approach that you chose so there's there's kind of like a space of trade-offs here um so you know going from essays which which are these like oh just do it on the neurons to um transcoders is kind of nice because that allows you to um one to like uh understand the computation rather than just representation but then also um once you've replaced these parts of the model you can then directly like connect features to features you can say like oh like you know this feature like at this layer uh writes in that direction because like now you've replaced a part of the models the feature is just like a computation it's like how the computation writes that direction and then this other feature reads in this direction and so actually they're not they don't talk to each other at all but those features go in the same direction and so they talk to another and so you can you can then like have very what does that mean what does directionality mean here you know in a concrete way.
29:12Sneak that one by here. So yeah so you know the way that here what we're replacing is MLP layers which are like multi-layer perceptrons and what they are is so you have this like representation which is 100 numbers and basically you have you let's just take like one neuron or one neuron of your MLP um which all it does is it like um does like multiplies your um your uh uh your hundred numbers by another hundred numbers to get like some some some float and then uh multiplies that by like a decoder uh and that decoder what is it it's also again a hundred numbers and so like you know the way that it's usually taught in like machine learning textbooks is you have like Like you have your neurons and you have like a matrix multiplication, another matrix multiplication.
30:13But another way to look at it is that each of those, all it's doing is it's saying like, I take the current vector of neurons. I like multiply it with another vector, which basically just like is kind of equivalent to seeing like how lined are there essentially, right? And then I take the result of that and I scale a second vector by that. And so it's like I read in that direction and I write to that direction. So I read in that direction. So that means that like, oh, if like there's stuff in that direction, it will activate this neuron. If there's not, the neuron will stay like zero. And then I write in that direction, meaning, you know, if I've activated this neuron, I'm going to write this other thing.
30:53So you can think of it geometrically as like trying to capture the similarity of two vectors and then projecting them off into another direction. Exactly. Exactly. That's exactly right. And so once you have that, then you can ask questions about like, ah, well then like how do those compose, right? So it's like, oh, I had one like earlier in the model that like computed the similarity projected in that direction. And then this other one is computing a similarity with this other direction. Are those directions aligned or are they not? And if they are aligned, then that means that like this feature will actually like activate this other feature.
31:32and so you'll see like oh you have a feature that like reads um you know i don't know like reads in the direction that's like the like mix like red and bridge and writes like golden and then you have another feature that like reads the golden direction and also is like gonna gonna survive and upweigh like things about the golden gate bridge um is is like the rough idea and so that's what transcoders buy you uh they buy you sort of like easier edges between your features which then is kind of what you want with circuits is you want to connect them all. Yeah, given the dimensionality that we're talking about, how do you even grok these directions?
32:10Are you applying dimensionality reduction types of techniques to understand what's happening here? Are there other ways that you're kind of getting at what direction means for these vectors? I mean, I think you just have to try hard to reason in 10 ,000 dimensional space. Um, it doesn't seem that hard. No. So, so, so the good thing is, um, we actually don't have to think about them. Um, that's, that's part of the pros of the method. And the reason we don't is that like, so, um, let me like walk through how you, how you actually use this in practice. You have a prompt, you've replaced a bunch of your models so that it has these like kind of like simpler or like more understandable, you know, bits of computation that like read in one direction and write another one.
32:56you take your prompt you pass it through you look at what these do right and all you do is you don't even like look at this like you know dimension whatever um vector you just do the the dot products to see if they're aligned and that gives you a number right kind of like when you're you know in geometry and so when when we show these like attribution graphs that's all they are they are is they're like a boiled down version of like very very high dimensional space where all we're showing is like each edge is basically like ah like what was the dot product between this like you know super big vector and that super big vector and it's like was it positive was it negative and that's just a number between you know one over minus infinity and plus infinity but it's just a number so you can you can understand that pretty easily and be like oh there's a bunch of strong edges from like the like uh you know golden to the like this is a bridge feature that makes sense um and so yes so yeah we don't have to we don't have to i mean there's a bunch of things you might want to do to understand these vectors and like you know um but but for the purpose of drawing this computational graph you you get to abstract over it which is which is really nice one thing you you will find though is that um it turns out that in these models there's actually a lot of like amplification and so when we're doing these these early experiments we would find that like a lot of our graphs were just very boring and what i mean by that is we have this example in the paper where you ask like the model basically to to tell you um i think it's like what's what's the capital of denmark or like copenhagen is the capital of and there's a copenhagen feature and then you're like oh i wonder what that feature is connected to and it turns out is connected to a Copenhagen feature in the second layer.
34:45And you're like, oh, I wonder what that goes to. And it turns out that that one is connected to a Copenhagen feature in the third layer. And so actually like the graph would just be like Copenhagen, Copenhagen, Copenhagen, Copenhagen, and then Denmark at the end. And that's because of the way these models work, they can't delete things. And so if you want to still be thinking about Copenhagen, you have to sort of like have every layer reinforced, be like, yeah, we're still thinking about Copenhagen. This is still the right thing we're thinking about. And so that's where our other method comes into play, this cross-layer thing, where we instead try to learn features that skip over all of these.
35:23And so instead of having 10 Copenhagens, you have one. And that's the idea between the cross-layer transcoder. But it's nice for that property, but I would say it's not vital for the method itself, if that makes sense. So you've got this method that you're using to replace the MLP blocks in the transformer. And that method results in a degree of introspectability, visibility into the way the model is thinking. One thought that I had was, it was kind of like an ad nauseum type of an argument. Like, you know, if you can do that, like it sounds almost like distilling the black box model to an interpretable model or whatever the opposite of distilling would be because this thing's going to have a lot more but more sparsely populated uh uh weights like could you just like up still it into you know a black box into like an understandable thing and like just use that thing uh and i'm really asking this to try to get that like a concrete understanding of this process?
36:42Like, you know, would it be too slow or like, you know, is it not really, I know you're doing some alignment between the replacement models and the actual model to ensure that you're, you know, the results are, you know, meaningful, right? But is it, you know, not very accurate after you've, you know, done this, this training process um are you only doing a small number of m uh mlps at a time in a larger model or you know and you couldn't possibly from a computational perspective do all of them like i'm trying to get at like some practicalities of the approach yeah yeah uh those are great questions so i mean i will say like ignoring sort of you know money or physics for a second which is always fun um if if we you know given like unlimited money and like probably like limited infinite data and stuff like that you could totally imagine yeah doing you know okay great let's like replace the whole model with this interpretable alternative and then let's run it instead and then and then we just get to to to we just get a free lunch basically where we have something that's as good as the base model uh and also we can understand it now in actuality you know as you said it's it's more like upstilling than distilling so like because all of the layers are way wider even if i could give you this perfect model actually be really expensive to to serve right because it's like it just like is a way bigger object um than than the base model there's also um kind of like a question of of training costs so in our in our biggest haiku uh 3.5 runs were 30 million features and those are those are pretty big runs um but if you think of the number of features as a feature is multiplicative of the number of parameters it's a vector as opposed to a that's right well right and if you're doing the per layer thing then like a feature is sort of like an encoder and decoder so like two vectors but if you're doing this like cross layer thing the way that it works is that a feature like is like one encoder and then like n layers decoders so it's actually quite expensive um but that sort of like is still worth it because you need fewer features because you're like deleting a bunch of redundant features so there's sort of like that's why it's still worth doing doing but anyways yeah our feature is like when we're on the order of like you know one maybe like close to let's say like what a what a neuron and mlp would cost you um and you know 30 million of them is is is a lot but um think about how many independent concepts the model knows uh and so there's this sort of like we call it the like dark matter problem in the paper there's a sort of problem that you know how many features would we need so that we captured everything that cloud knows um currently using our methods that's probably in the order of like billions maybe uh and so that's like quite infeasible with with sort of like just today's technology it would cost way too much to train it'd be completely impossible to like even study it's a giant monster um and so there's a sense in which in order to capture everything in the model is doing and so to both explain it fully, but also to be able to kind of like reconstruct it accurately.
40:05We need so many features that right now, this is sort of still basically an unsolved problem. I think one thing that's relatively clear, and again, we talked about in paper is we're going to need to sort of like think of smart ways to, you know, get features more efficiently or get features more efficiently only like a domain that we care about. Maybe we don't care to explain how Cloud does everything. Maybe we only care to explain how cloud does math or how cloud does like pie or something that buys us a lot. But also there's a bunch of methods that you could imagine using to just like be more efficient with your features that we discussed a little bit.
40:44So I think that the sort of like summary answer to your question is that's the dream. I think that right now the sort of like realistic state we're in is actually we can explain what some of what the model is doing some of the time. And I think the sort of like first step after that is like, ah, can we just keep increasing the proportion of times that we kind of like understand what's going on? And I don't know if we'll ever get to the dream of we just replace the model with the interpretable model. But I think even if we just manage to sort of like understand a lot more, we'll get a lot of value here.
41:23There was some discussion that I found really interesting talking about kind of one of the challenges or confounders being these ideas of polysemanticity and superposition. Talk a little bit about those ideas. So I guess like one thing you might think is like, why do we need to do this at all? The model already has neurons. We could just use them and then do our graphs based on that. And it turns out that if you just look at the neurons and you try to understand them the way that we try to understand features, which is like, you know, run it on a bunch of text and see what makes it light up and see if when you look at that pattern, it's obvious what it does.
42:05So for features like the Golden Gate Bridge feature, it's like if you do that, you run it on like a bunch of text and you highlight the parts where it lights up. It's like, oh, it's like, you're literally just highlighting the text to Golden Gate Bridge and a bunch of text and everything else. It's just very clear. Most neurons are not like that at all. Most neurons, you look at them and you're sort of like, ah, like it seems to fire on sevens sometimes but only if there was like russian language before and also you know like german cookbooks and or or like even that is actually like kind of understandable it's like mostly just looks kind of random um and so that's because um actually like each neuron kind of like isn't a meaningful unit for for the model and the model just like uses a bunch of neurons in parallel and stores serve like representations over all of them and so those neurons are what we like polysemantic they just like don't have a single meaning whereas the features are monosemantic they are like you can say like ah this is the michael jordan feature if it like has a very high activation that means the model is thinking a lot about michael jordan if it doesn't it's not um and so the idea of superposition then is well like how do you how do you reconcile this idea that there are features that are sort of like very meaningful and very sparse with the fact that if you look inside the model it's it's like a very dense thing where like every kind of like you know in in the residual stream of the model which is where the kind of information uh it's processed just everything is on all the time there's no sparsity um and that's because at least the hypothesis is that it's because of uh essentially like superposition what's happening is that the model is like storing many more features than it has dimensions in superposition and like squishing them tightly such that when you look at it it's just this like mess uh but actually what's happening is that like each feature is sort of like you know smeared across like many dimensions and our work is to decode that smearing into these like abstract monosemantic things, basically like untangle the superposition to get the monosemantic representations is like kind of the name of the game here.
44:22We've touched on limitations of the approach throughout. One that jumped out to me is the idea that by focusing on kind of the replacement of the MLP layers, you know, you're missing out on another big source of information, which is the attention mechanism um can you talk about that and more broadly like other limitations totally yeah i think i mean you know attention is the biggest one by far um and so there's there's a lot of stuff i mean there are many other limitations so i guess we touched on one which was this like dark louder which is if you just extrapolate how many features we have and how many features we need we're just not going to get there so we need to do work to sort of like find a way to get more efficiently so that's one attention is is is is a huge one because there are many things that the model does that are basically almost entirely done via attention um maybe you know a common example uh is like i was gonna say induction which but that's like a induction is when the model just like looks back at previous instances of the word it's processing and looks at what was said after that word and then just says the same thing.
45:40That's maybe like, that's interesting, but actually like pretty well understood. But there's a lot of stuff we don't understand or at least it's harder to understand well, where, you know, you can ask like a bunch of questions around, maybe a better example is like multiple choice questions where like the model has to, you know, you ask it a multiple choice question and then at some point it has to say like, well, it's answer B. and the way that it says it's answer b is that there's an attention head that looks at b and then like writes it at the end and if you look if you if you use our method that's basically all you see you're like ah great like the complete story of what happened on this prompt is that Claude somehow knew that the answer was b and then like an attention head like moved it there but of course the interesting question is like well why did this attention head look at b like there's some computation that happens in the model that tells it that the correct answer was b and then like it looks at b and it copies it and currently we just completely sidestep that um we sort of like explain part of attention or at least we like we we explain how attention moves the information from mlps but we don't explain why uh attention patterns were the way they were and that's like a huge part of how the model process information is like how did it decide to attend from x to y is something that we completely miss.
47:02And so, you know, obviously one of the main limitations and kind of aspects that we're interested in working on and hearing from other groups working on it. I think solving that is by far sort of like the most exciting thing on my list in terms of kind of like next work. You know, I don't think it would be easy to be clear, but I think it's doable. And why is it so difficult? Is it is attention kind of a denser mechanism? Is it more, you know, opaque somehow? Like, you know, they're all numbers somewhere. Like what makes attention more difficult to kind of to pick apart? Can you train some other, like is there an analogy, like a straightforward analogy of like a sparse replacement attender that, you know, is conceptual or does the nature of attention, you know, cut off avenues like that?
47:59yeah i think just sort of like a few reasons why attention is hard but none of them are fundamental like you totally could train some sparse replacements uh for like parts of attention i think um attention is just like different in a few ways maybe like one way is that um the other aspects we replace are these sort of like maybe like univariate is the wrong word but distance you know it's like you take the state of the model and then you do your like you apply your mlp and then you write the result and so you can replace that it's like a very clean computation but attention patterns specifically are the combination of what's happening at all of the key positions and all of the query positions and the interactions between both of those and so you had some you now have something that's like quadratic basically and so you have to sort of somehow like find a way to get around that and there are there are sort of like proposed approaches um and i think some of them are are you know likely to work or at least i hope they are so a lot of the limitation is computational bounds for the same reason attention is hard uh you know to scale you know for for context windows and stuff like that exactly so that's that's one of them i would say another one too is that like um there's also a uh you know like like non-linearity over attention like it's not just so you're doing like k cube and then you're also doing like your non-linearity which then like that just means that it makes your decomposition efforts like you have to think carefully about because um in at least in like our graphs everything because of our replacement we've managed to make everything linear so you can just literally say like ah this feature affects that one in that way but here it's sort of like ah yeah like this this key effect kind of obscures relationships exactly but like actually what really matters that that one was much stronger.
49:49And so actually they're sort of like, you didn't really have an effect at all. Like you are getting squished by the non-linearity. So I don't know, I guess like, maybe another way to say it is that
50:05replacing MLPs first here, this work, just gave us enough sort of like valuable empirical insight and valuable work to do that it felt like the first thing to do. um and i think you know there's similar papers in the field that sort of like have various other like they use transcoders as well um and that have found success and then you know it's kind of like now that you've solved this or at least have found you know maybe not solved but found some progress well um yeah uh now you see the sort of like like the next thing becomes the gaping hole right like now that you oh it's like oh we've patched this hole like this makes sense and it's like oh and now attention is this huge thing it's like why haven't you done anything about attention it's It's like, well, we were starting here, you know, and so, but I think it's basically the next kind of like exciting thing.
Read the full transcript
50:53We've kind of, you know, touched on these relationships, but haven't explicitly talked about this concept of an attribution graph, which is kind of like how you make sense of all of this information. Can you talk us through those graphs and like how you create them and what they tell you? once you've kind of like replaced these bits of the model with with uh these interpretable bits uh you can kind of like i was alluding to the fact that you can kind of like connect them so you can say like ah this you know the model saw let's i feel like i'm like bouncing between examples we're going to pick fact uh michael jordan plays a sport of and stick with it for a bit so the model saw you know michael jordan and there's like a feature that let's say like reads Jordan and writes, you know, like basketball.
51:40There's also a feature that reads Jordan and writes famous people. There's a feature that reads like sport and writes, you know, like talking about sports and talking about sports players. And there's a bunch more of these features on like a given graph, you'll have, you know, hundreds that are just doing various things like extracting meaning, combining meaning from previous features. The attribution graph is just saying cool we run the forward pass we we record all of these features and we just connect them based on you know as we're talking about which one's kind of like right in the same like which ones uh write and read in the same direction such that they like affect each other basically and we just show this as a graph and we go all the way from like the model saying basketball to the embeddings of tokens and we see like what is you know literally like what is the sort of computational path that went from this prompt to you saying basketball and it's like you know in this example, there's kind of like two paths where it's like, oh, there's like a path that's Michael Jordan and like a bunch of famous people.
52:38And eventually it gets to like, ah, we're talking about basketball. And there's a path that's just like plays the sport of, and that one has a bunch of features about just like popular sports. And so that one kind of like pushes up football, but this path pushes a basketball even more. And so then the model says basketball. And so these graphs give us basically a hypothesis about what the model did. The reason it's a hypothesis is that they're based on this like replacement that we did and the replacement isn't perfect and so maybe we're wrong and so then a huge part of the paper is like take the graph and then you can then say like well if this graph were true then like nudging the mlp at like this layer in this direction which the graph claims is the basketball direction or whatever would like make the model say basketball or make them all not say basketball and so then we basically take these graphs and we test those by directly intervening on the model.
53:32And those interventions are a big part of the biological study. And I guess there's at least a couple of kinds of interventions, or maybe I'll have you taxonomize interventions, but like one of them seems like a nudge, but they're also like you're able to, you know, either insert or delete, you know, next words or like the example with the poem, like you cut off the model's ability to to use rabbit as the destination for the next line in the poem. uh talk a little bit more about the um those interventions and in particular are those interventions happening at the replace in the replacement space replacement land or are they happening in you know the replacements tell you how to do them in model land like how is it happening the replacement model tells you how to do them in model land for for the most part and so for for almost all the experiments they're sort of like happening in model land and it's basically how we validate that like the replacement model wasn't lying to us um they're mostly of the same shape there's like a bunch of variants as you said but but mostly they kind of like all do the same thing and the thing they do is like identify a feature or often it's like a group of features because we often have like you know a bunch of rabbit features instead of just one um and then either like suppress them meaning like you know like as i was saying they write in a direction so you just like write less in that direction you literally like go at that part of the model and you're like you know write less or like like equivalently yeah uh sorry to jump in but before we get too far from this like you know this is a little bit of backtracking but like reconcile the idea that like you know we've got this challenge of uh polysemanticity and superposition you um you know you go from kind of mlp to features in this replacement model like how do you get back like what do you are you getting back to individual neurons or those individual neurons spread all over you know mlp or multiple layers or like once we do the replacement model and and we let's let's keep this like rabbit example we have this like rabbit feature like what actually is that rabbit feature you know mainly um it's like um so it's it's an encoder which is like what causes it to fire but then once the rabbit feature is active the way that it like influences the model is these like decoders right which is like the mlp down projection in the base model it's the equivalent in feature space and so literally if you take the rabbit feature you're going to have like a you know let's say at layer or whatever like five um you're going to have a um a vector of shape d model like the like shape with the residual stream that is that feature and so like that's that features decoder for cross-layer transcoders you actually have like one at every layer but like i think for now let's just take one layer so you have um you have that that feature direction and you say like ah the replace model is telling us that like at mlp layer 5 there's like a rabbit feature and it has this decoder in that direction and it is also telling us that if we were to like suppress it or sometimes we even inject it like inject the opposite which i'll explain what that is in a second then the model wouldn't end its sentence in rabbit and so the way we test that is we throw the replacement model away we keep that vector and we inject this sort of like you know this vector times minus five or something or minus 10 or like depending on the strength of the intervention um we inject that into the base model and we see you know if we when we didn't do the injection the model was going to end in rabbit then we do this injection at this layer injection means changing weights or it's like yeah it's like basically you um you're multiplying your vector by the the weights to get a new set of replacement weights you're so you're running your normal weight scaling i guess it's like you're running your normal forward pass as you would but then when you're at this like layer five mlp you just add this this thing so so or like an arc has subtract but like if you wanted to add the rabbit feature you just add this to the to like what the mlp was going to write or you like add minus 10 times so it's going to like basically the idea is just like you let it do its normal thing but you just subtract that component from it like you have this sort of like theory that a part of what it's doing is writing rabbit and it's doing a bunch of other stuff and you say i want you to keep doing exactly what you're doing but i want to subtract the rabbit direction from what you were doing so it's like keep the model doing everything else it was doing because you know it's doing like other processing to just write a poem at all but just don't think about rabbits and then you run the base model uh and then in the case of this experiment right you see oh uh the model um for the rabbit experiment actually it turns out it's planned rabbit and habit and so if you do that uh it will just like target habit and write a different sentence that ends in habit if you then like remove you do the same thing for both at the same time you use it like lands on a different word um etc um but that's sort of like every experiment in the or every like intervention experiment is of that shape in the paper um and so we just like kind of like remove things or add new things uh to to what the base model it was doing and record kind of like the consequence of that experiment and then that tells us whether whether our graph was was correct or not and maybe more importantly than telling us whether our graph was correct it tells us cool things about like what the model is doing right it's like kind of forget about the graph it's like oh we found this direction that actually like does seem to be like the direction of the like the model's plan and when you mess with it it changes the model's plan so another aspect of that question was like how do you reconcile that you know the ability to do interventions in this way with the idea that the idea of rabbit is like spread all over the model and not just in layer five it turns out that like this kind of varies based on on like the cases we've looked at.
59:45But yeah, oftentimes like, so you know, the rabbit idea will be spread over let's say like three or five layers. Sometimes if you just change the first layer, that's enough because actually what they're doing is like amplification. And so it's like those next few layers will only fire if - You talked about earlier the idea of like continuing the thought through subsequent layers as a way to preserve it. Exactly, exactly. And so sometimes if you just remove rabbit from the first one, the next few will just not say rabbit because they're like, well, I was just reinforcing what was there and it's not there anymore.
1:00:18So I'm just not going to do that. Or if you inject something else, I'll just reinforce that instead. And then sometimes it's not the case. Sometimes they're each kind of like doing the thing. And so sometimes we kind of like do this injection in a range. It's like I like subtract Rabbit from like this guy and the next one and the next one and the next one to prevent them all from thinking about Rabbit. But we have some like kind of like more discussion of that in the methods. But I think that's also like there's this notion of not only, you know, is the model like spreading its computation over many neurons within one MLP, but it's probably also doing it over like successive MLPs because it doesn't care.
1:00:59Like it has no incentive to like have understandable layers, right? You could just like compute something over four layers in a row. And so that's some of the weirdness that we're confronting here. that sometimes you you know you're like to your point like rabbit is isn't a concept but it's like spread across neurons and spread across layers a little bit i was going to ask along those lines do you have you made any observations that uh relate for example you know a concrete concept like rabbit is more concentrated you know in some way in the model versus a more abstract concept like oppositeness or you know any of the you know so any other concept like is there are there things that this work tells you about kind of the structure of representations in a model with regard to concepts in particular yeah I think there's so there's there's some of that there's not something that directly tells you, or at least if there is, it's not coming to my mind or something like an experiment that would directly kind of like address that.
1:02:10But there's a lot in terms of structure. So one of the other striking examples is this sort of like math example. So, you know, you ask Claude to do like mental math, essentially like we ask it 36 plus 59 and we ask it to do it in one forward path so it can't sort of like write out the actual like longhand addition algorithm or whatever. And so it can do that perfectly fine. And then we look at the attribution graph for this and we see that there's this like rich structure of, you know, there's like a bunch of features that detect like, ah, like, you know, this ends in six. That's surprisingly rich.
1:02:53Yeah, yeah. Where did it come up with that? I know there's there's this sort of sense in which like it's a mix of like wow uh we're really giving these models hard tasks that it has to serve like galaxy brain invent this sort of a complex tree of of like math and also it's like not the algorithm that like you you're taught in school you know and and and it's like a weird different thing where basically like the gist of it is that it like precisely uh kind of like determines the last digit uh you know it's like ah, this definitely ends in five. And then it has this other path that's just like, ah, it's like kind of, you know, roughly in the 80s or in the 90s and like combines both.
1:03:36And you give a specific example of this and this is like the adding six and nine. And did you do any kind of ablation or something to identify that those pathways aren't specific to like the six and the nine being such a strong signal and the training data that like, you know, you subtract one from the number if you're adding by nine, that kind of thing. Totally. So you were asking about structure, right, which is how we got on this point. And that's why this example, I think, is like a really good one, because in the circuit tracing paper, we have this like addition case study that we're talking about.
1:04:17And then we have the next section, which is global weights. And it's something we, you know, we haven't gotten a chance to talk about much here. but it's like if you're replacing this model you might ask not just like what's happening on a given prompt but to your point like what's the general way that it adds numbers like is it adding six and nine differently from when it adds like two and three are those kind of like all using completely different stuff or is there kind of like shared pathways you know we saw that there are shared pathways in in like languages right and so is there sort of like sort of like a shared sort of algorithms and we've actually like shared this um we've computed all of these like um these edges between features not just on a prompt but in general for math like all of the ones that we could find um by asking claude to basically do like addition for every i think number uh below 100 or something and we took all of the features that were that were present and then we made this like little app that you can use it's like a little little cute app word like connects um it shows you the graph of all of them and you essentially see kind of like oh like six plus nine there's like a there's like a six plus nine and then it's like i'm gonna say 15 and that one if you look at like what are its other inputs oh well it's seven plus eight uh you know and it's like uh it's like all of all of the like other other um additions also have that same pattern where they have these like same inputs um and so there's like basically a structure here where it's doing these addition problems in like very similar ways no matter what the operands are and you can see that structure and you can kind of like see the like paths for kind of whatever you would you would give it i think example that last example is a good segue to mentioning something that i really struggled with when first starting to read the circuit tracing paper and that was really around the, you know, part of it was like the anthropomorphization of the model and, you know, talking about the model strategies and the model's decisions and things like that.
1:06:22And like, you know, with the example of the poetry, you know, the model doing planning and trying to reconcile this idea of a strategy and planning with you know initially like a you know technical you know understanding of inference and how inference works and it's like relatively simplistic deterministic kind of you know pass through or straight through you know process and it took a while to like really get the idea that like you know while the model is doing the generation and inference the strategies are you know they come into the model through training and
1:07:12you know it's just it's an interesting relationship like a model is is you know and i think this math example you know really illustrates it like you know there you know you're adding you know two numbers two numbers smaller than a hundred and there are hundreds and hundreds of strategies that are kind of there dormant in all of these, you know, vectors and weights in the model. And, you know, somehow in inference, the right ones kind of fire up. And, you know, that is the, you know, that is the embodiment of this strategy. As opposed to like, you know, during inference, the model, you know, I don't know, it's like the initial read gave me the impression of, you know, a model like actively choosing in a way, you know uh and maybe a very anthropomorphized way of like choosing a strategy and it's not really like that i'm like reconciling all that um took a few passes for me yeah yeah i think that's that's just sort of like really interesting like on the surface you know this is like a vocabulary question or something like is it planning is it is it adequate to say that it's planning what does it mean to plan uh do you sort of like are we implying something about something that's not there when we say you know it's planning or it's sort of like has has algorithms for addition um but i think you know underlying that question is sort of yeah a really fundamental like how should we think and how should we describe like what what these models are doing clearly they're able to do very complex things like i think um a sense in which when i when i've talked to these results to my friends they were like less surprised than i was you know when i was like look like there's planning they'd be like well you know like claude writes my emails and my code like of course it plans like you know it's it has to plan like how would it ever do any of the things that it's doing without planning right um and i think in that colloquial sense it seems kind of fine to say like well yeah the model like imported this library when it wrote the snippet because it was planning to write code that was using it uh at least that's never sort of like struck me as as as like maybe contentious as a vocabulary choice you know um and and at the same time here we're like at the opposite scale right like going from the behavioral to the micro and we're saying like well you know what are we saying we're saying like there's um circuitry that seems general that consistently generates candidate words for sort of like many, many words ahead of time and then uses these candidates to sort of like decide, like reasoning backwards, right?
1:10:02Like deciding from the candidate how the sentence should be structured such that it arrives at that candidate. And like, that's just a description of a mechanism. There's no, you know, anthropomorphizing. like that's just that's just a thing that's happening and then you can sort of say like well that's planning or a token at a time in the forward direction yeah yeah somehow well and which by the way i feel like is is sort of like somewhat like more impressive like definitely like surprising that like the model is doing backwards planning and is forward past five tokens ahead of time absolutely but but to your point like i think i think there is an interesting question like okay well these are pretty complex mechanisms for for for that for addition as well right they're like pretty complex algorithms are being implemented um does that mean like like that definitely means that it's tempting to like draw parallels with how i would write a poem or how you know i would like add numbers really quickly in my head without actually like writing it out um and that sort of we also want to be mindful of not anthropophizing for anthropophizing's sake but i think sometimes there's been you know it's it's it's actually like i'll maybe pose it back to you like if if if the the poem example if it's not planning then like what is it uh it's sort of hard i think to to find um to find the right vocabulary here um and i and i think maybe i'll say one more thing here that i think is really important uh which is that i think that's part of the reason why i'm excited about mech interp in general is that we've had debates for a very long time where it's like are they you know stochastic parrots like do they you know just like kind of like match to the most like similar training data like or do they like generalize what is generalization what does that even mean um and i i think what's really cool about about this work what really like gets me excited is we can just talk about the mechanism and then we can debate like you know what what does it mean that this is the mechanism but we can at least start from a common ground and be like cool like this is how an llm like writes a poem now we can have a debate like well what do we think this is what we want to call it how we don't characterize it but i feel like it's it's sort of like giving us much more solid ground than just a sort of like behavioral you know like ask some riddles and see if the model succeeds or fails and based on that decide whether they're like super intelligence or or dumb robots or something you know yeah yeah um a couple of things in the results struck me as if not strongly at least lightly contradictory like the idea that, you know, and, you know, one of the examples, one of the conclusions, you know, the poetry example, like the conclusion is that like, you know, Claude is doing planning and it kind of knows where it's going once, you know, even before it starts, you know, spitting out tokens.
1:12:52But then there's another result, and I forget which example this comes from, that talks, I think it's maybe the names and Michael Jordan stuff. It talks about like, you know, forward pressure to maintain, you know, the coherence of ideas. Just a jailbreak example. Jailbreak example. Okay. So you can maybe explain that before, you know, as you're going into this. But like, or even, you know, thinking about, you know, the general problem of hallucination. And, you know, it seems like the idea that like the model kind of knows where it's going and has this concrete destination, you know, is a little bit at odds with the idea that it's just kind of following the wave of momentum of, you know, the tokens that it's generating or, you know, that it would end up in a place that, you know, even, you know, the jailbreaking example specifically speaks to this, like end up in a place, you know, going down this path that it doesn't want to go to.
1:13:52and then like somehow realizing it and like correcting itself uh speak a little bit to that so i yeah maybe i'll like just briefly recap the jailbreak example but like essentially the jailbreak example in the paper is one where you you get clawed to like basically spell out bomb and start a sentence that's like to make a bomb and then it's going to keep telling you how to do it um and and the finding there that the thing that's that's kind of fun well there's a few findings but specifically for that part is you're like well why you know why does like claude has been trained to not tell you how to make a bomb like why is it continuing to be like oh you do this um well first it turns out that it continues to tell you until it hits a period and then it starts being like oh i'm so sorry should not have said that you know definitely don't make a bomb um and and so we we like looked at circus for like well before it hits the period when it's telling you all these instructions like why and there's two there's two paths uh and and one path is this sort of like hey i'm talking about bombs that's harmful i am a harmless assistant i should not be doing that and that path like upweights various uh tokens like saying i which often is like the start of refusal like i apologize um but then there's another path which happens to be stronger which is i am a language model that writes correct text and i need to have like a correct grammatical sentence and i'm in the middle of explaining that you know like you mix this chemical with that chemical and really like i should have like a a tertiary clause to that before i like have a period or whatever um and so and so the model is like in a way you know there's just like demon inside it that's just like compelling it to be like no you need okay actually like i'm realizing i'm saying demon inside it forget i said that that's even worse than anthropomorphizing but there's like two paths inside it um and and like one path is stronger than the other uh and and to your point like at the same time you know we say ah in the like poche example there's a sort of like remarkably let's say like um not next token prediction sounding mechanism um but i don't i don't think those are contradictory if anything i think they they serve like are they make a lot of sense because i think uh um an experience of using language models and and you know building them is that they're incredibly good at some tasks that if like a human was good at that task at that level you would assume that they would also be good at this other task but it turns out that actually they can't do this other task and so they have this sort of like jagged set of capabilities right where they can do some stuff they can't do some stuff and i think that mechanisms like this um sort of show you like yes okay in this example it's learned something that's like robust just like planning example that's like that's really going to be helpful to just generate poems that's like a good kind of thing and then in the jailbreak example i would say it's unclear to me whether like it's it's like like a bad mechanism or or at least like you know it's just like this weird like the model's trying to do two things at once and we see this it's a motif we have almost always there's like many different paths converging or like competing sometimes.
1:17:07And then the molecule like ends up like having to choose between one of them. And then maybe I'll mention one more thing, which is your like hallucination example, because I actually think that's an example of like a bad, in my opinion, a bad mechanism in that like what we found was if you ask the model about someone like, you know, who is who's Emmanuel, who's like Michael Jordan, who's Michael Batkin, like there's kind of two different circuits. first there's a circuit that determines whether the model like knows or doesn't know the person and the way that circuit works is like by default cloud just doesn't know like the circuit is off by default it will always refuse but if like there are enough features in the name or in the context there's like basically like a different circuit that will inhibit uh that like i don't know and the model decides to answer but the kicker is that circuit is separate from the circuit that's actually like retrieving information about the person and then like telling you stuff and so then you know it's possible for it to like recognize the person enough that like it's like okay i'm gonna answer but then when it gets to the point that like there's separate circuits to like do factual recall it just makes stuff up because it actually doesn't know like a way better mechanism would be first you retrieve stuff about the person and then conditioned on that you decide whether you're gonna like accept or not that's you know what we do if you ask me who x is i'll be like nope doesn't ring a bell you know but Claude has sort of like this like two separate mechanisms where there's like one that says like no you definitely know who this is but that it was somehow not connected to like the like actual like retrieval uh of the other one and so I think that's like again I wouldn't say that's contradictory I would say like that's an example of the many facets of like these mechanisms that exist in these models and that some of them are you know surprisingly advanced and some of them are like surprisingly just like silly and and and kind of we kind of forget are getting to see how the sausage is made a little bit understand some some of the reasons for why sometimes these behaviors are really cool or perplexing and this michael jordan michael batkin example you mentioned in the paper kind of in passing that for whatever reason claude thinks that michael batkin is a chess player like it hallucinates that quite consistently, I think was the term that was used.
1:19:20Any investigation into why that might be? And did these tools, you know, tell you like why chess of all things? I personally haven't looked at this much. I don't know that anyone on the team has. I can offer you wild speculation, which I'm always happy to offer. But in prior work, I think it was, we posted, like we do these monthly updates and we did one maybe like almost a year ago. uh about sort of like asking a smaller model you know uh blah plays the sport of x and we did notice that when there's when when the model uh is is like saying a sport for somebody that doesn't know there are just like common biases if i remember well like for for male names it was golf uh it's like if the model didn't know it was just gonna guess golf and so i think one wild speculation here is like there's like a mix of just say something that's a sport and maybe chess is like relatively high on the list uh and then there's maybe also this is again wild speculation but something to do with um the features of the name maybe there's like some bias where it's like ah maybe it sounds like a russian chess player um but but yeah we didn't investigate it carefully i think we covered a lot of the examples from the biology paper any others that jump out at you?
1:20:43Yeah, we went through a lot. I think, let's see, I think maybe, maybe that, the one thing, yeah, actually there is one thing that I think is like maybe interesting to mention. And that's the, what we call the chain of thought faithfulness one. So that one is one where, you know, there's these reasoning models and these thinking models that like will explain to you what they're doing. And, you know, it feels like kind of important that we should be able to trust this reasoning and it turns out that you know there's other papers on this but but in in we have a case study that's sort of like studying an example where the model is just like not being faithful and the example is sort of you ask it a really hard question uh like a really hard math question and you know it's like cosine of like 2343 or whatever something you can't do um and by default first of all it'll just like make up a number uh that's between zero and one i think and i'll just like say you know that number um but if in the prompt you tell it like oh you know i worked this out by hand and i got x like do you think that's right and you look at the circuit the model will actually like look at the hint you gave it and work backwards from that hint to like say something that will lead it to the answer that you gave it as a hint um and like you see this when we look at the attribution graph but to be clear the model is like telling you that it like calculated that this cosine is equals to it turns out exactly what it needed so that it landed on the answer that that you gave it and so that's sort of like an example of like the model clearly um not you know uh not like like like a clear difference between like what the model is writing in its chain of thought and like the mechanism through which it actually wrote it um and i think that that's sort of like was quite compelling uh to me because it's just like you know there's a question it's like well do we need to do inter at all if the model like the model just tells you what it's do.
1:22:52So just like listen to the model, you know, but it turns out that in many cases, there are actually discrepancies between what the model says and what it does. And you can imagine like a bunch of reasons why you would very much want to know that like if the model is making a very high stakes decision and is telling you it's because of this reasoning, it's not because of this completely unrelated reason. Yeah, I found that this result less surprising than some of the others. Mostly, I think, because it, you know, it just, I think, vibe with a lot of our experience with these models. And, you know, some of the early jailbreaking work, like, you can kind of convince the model of things that it should know, you know, aren't right, or to do things that it should know aren't right.
1:23:40And when I think about what I see in those thought traces, I think of it as just the generation of tokens, you know, inside of some delimiter as opposed to something that has some, you know, very significant, you know, meaning like, you know, some deeper, you know, thought meaning. So it didn't surprise me that you can mislead a model so much, but the implications that, you know, those things are not mechanistic interpretability signals, I think is important. yeah and and i think you know it's it's funny you mentioned sort of like it as like misleading the model because because my frame was like the model misleading you um but i mean i think they're equivalent but but you know there's this sort of sense in which here like i think this this almost to me like rhymes with reward hacking a little bit right where you know you've told it the answer you'd like to get and so there's certainly like a bunch of circuits in the model that will be like well like you know emmanuel said that he got four yeah exactly exactly um but to your point yeah i think that i think there's like maybe like a broader theme here that like if you've played with llms a lot like a lot of these results that i said you'd be like yeah like it plans more than one token at a time like yeah sometimes you can't trust it's chain of thought um and so i think there's kind of like two maybe sets of of of kind of like takeaways which is like one is if you haven't played with these models i think a lot of this stuff is like quite surprising and and and if you were under like if you haven't seen or played with them like for like two or three years you know and you'd be like oh they're sort of like i think the complexity that's embedded inside them is like more nuanced um and then if you have played with them and you sort of like have good priors for a lot of stuff i think then it's kind of what i was talking about earlier it's Like the thing that's exciting is like now we get to not just say like, yes, it makes sense that like you can like mislead the model in this way.
1:25:44But it's like this is exactly the circuit and like the causal structure through which the model is being misled by your hands. This work falls under the banner, as we've mentioned multiple times, of interpretability and understanding what the models that we have are doing. projecting forward though how do you see this work as um influencing the way we build future models like you know do the things you know in what ways do you see the things that we're learning here changing the way that you know claude 4.2 is going to be built or whatever bold of you to assume that there'll be a cloud 4.2 it'll be cloud 3.7 new um uh i think interp is part of our serve like i thought we were behind these those uh you know naming challenges no i think i think this is the sign of you know any lm company needs to have a few choice naming decisions.
1:26:47But, you know, I think Interp is one of our, it's part of our safety portfolio here. And so like one example, again, in the like biology is there was this auditing game that was run recently where, you know, a model was sort of like acting strangely. It turns out it had been trained to act strangely, but in a way that's kind of like not obvious. and um my hope i mean like the team's hope is that like tools like interp can help with things like that like in the case of that game interp was helpful you could also do like kind of like uh figure out what's wrong with the model without interp but i think for like complex like yeah reward hacking behavior or like behavior where we might just notice that the model is like acting strangely and be concerned about the implications of that about you know like what does it mean that the models like acting strangely in these contexts being able to sort of like actually explain the mechanisms through which it's it's acting that way um sort of like as a as a like auditing step feels really valuable to have as part of just like any sort of like model um auditing pipeline before you realize before you decide like to to release it and like the extent to which you or feel comfortable releasing it um i think there's also like you know i'd mentioned when we started this conversation and there's sort of like a a three-step plan and we're kind of on step two i think a third step or something that we've been interested in for a while is sort of also teasing out like like not just um explanations for why did claude do this in this example but like global structure from the model something closer to the addition example where i mentioned we kind of like mapped out how it adds like all of the numbers uh you know from like zero to 100 like you could imagine doing that in many other domains how does you know how does claude write code generally can we have like an understanding at a bit more of a macro level of you know like roughly the pathways it's like ah this is kind of like the part that tends to think to like have like mechanisms to deal with like control flow this is the part that you know uh tends to decide on whether you're going to do object-oriented stuff or not um i think that things of that shape seem like they could also be much more directly useful to kind of evaluating how models are are working if they're working well if they're working poorly um and they're kind of also rhyming with some of the multilingual stuff like trying to see if there are shared representations when you'd expect that there were um i think is also sort of like something that um that like can help you debug like oh maybe maybe the reason like the model is behaving strangely is because it hasn't made like a connection to these two things right so i think what i didn't hear in there is any strong like we learned this and that's going to influence you know that or we expect it to and And maybe, you know, the way to think about that or contextualize that is kind of going back to the fundamentals of like, you know, all of these properties are not designed properties.
1:30:01They're emergent properties that came about when you, you know, take a lot of data and put it through some training recipe. and, you know, we're not, you know, we're not going to take things that we learn here and necessarily like re-architect a model to do different things or build. Like I'm thinking about like maybe hallucination, for example, you know, or even the addition thing. Like you've learned some things that could be interesting if you really wanted to, you know, You know, if you could figure out a way to tinker and build like subsystems in a module in a model to like, you know, better do thing X, Y, Z.
1:30:46But, you know, maybe that's, you know, that flies in the face of a lot of kind of the idea of end to end training, you know, based only on data and not trying to like, you know, overly manually engineer features like kind of old world ML versus, you know, new world ML. Right. Yeah. I'm learning that you're not uh you're not bitter lesson pilled yet um well I think okay so I think unclear to me the extent to which kind of yeah like these like you know graft on something to do math like that's kind of outside of of my my wheelhouse um but but you know it's also the case that um we really like as i mentioned inter for us is really like a part of our safety strategy uh and so that's the main goal so the main goal is understand the model um i think also like this is early days we're really excited about this this set of papers we think that sort of um you know it opens up a lot of biology questions i think it's also still just like very early to sort of make any sort of claim about uh we've learned like this fundamental truth about how the model does things and have strong opinions that it shouldn't do it that way um but yeah but but the main thing that we have in that's very reasonable yeah it's sort of like let's let's keep pushing on understanding the models better and and through that understanding on sort of again kind of like being more confident when in like our pre-deployment testing i think is like you know, like starting the main things that we're concerned with here where like detecting, yeah, reward hacking, various forms like misalignment, things like that.
1:32:38I think we have a line of sight on and are pretty excited about. Well, Manu, it's been great catching up and this is super exciting research and I appreciate, you know, all of the explanations and the time you took to go through them with us. Yeah, thanks for having me on. I'll see you, I guess, in three years.
1:33:01hopefully we can uh we can increase the cadence or something sounds good to me all right thanks so much bye
From the publisher
In this episode, Emmanuel Ameisen, a research engineer at Anthropic, returns to discuss two recent papers: "Circuit Tracing: Revealing Language Model Computational Graphs" and "On the Biology of a Large Language Model." Emmanuel explains how his team developed mechanistic interpretability methods to understand the internal workings of Claude by replacing dense neural network components with sparse, interpretable alternatives. The conversation explores several fascinating discoveries about large language models, including how they plan ahead when writing poetry (selecting the rhyming word "rabbit" before crafting the sentence leading to it), perform mathematical calculations using unique algorithms, and process concepts across multiple languages using shared neural representations. Emmanuel details how the team can intervene in model behavior by manipulating specific neural pathways, revealing how concepts are distributed throughout the network's MLPs and attention mechanisms. The discussion highlights both capabilities and limitations of LLMs, showing how hallucinations occur through separate recognition and recall circuits, and demonstrates why chain-of-thought explanations aren't always faithful representations of the model's actual reasoning. This research ultimately supports Anthropic's safety strategy by providing a deeper understanding of how these AI systems actually work.
The complete show notes for this episode can be found at https://twimlai.com/go/727.




