In short
Generative Now Podcast Episode Summary
Episode Title
Inside the Black Box: The Urgency of AI Interpretability
Podcast Description Generative Now is a weekly series from Lightspeed that explores the innovative stories and strategies of AI companies. The podcast features discussions with founders, engineers, and designers who leverage AI to enhance their products and improve our work.
Episode Overview This episode, hosted by Michael Mignano, was recorded live at Lightspeed’s offices in San Francisco. The discussion focused on the significance of AI interpretability, featuring Nnamdi Iregbulem (Lightspeed Partner), Jack Lindsey (researcher at Anthropic), and Tom McGrath (co-founder and Chief Scientist of Goodfire). The conversation highlighted the challenges and opportunities in understanding modern AI models to enhance their reliability and safety.
Episode Chapters
- 00:42 - Welcome and Introduction
- 00:36 - Overview of Lightspeed and AI Investments
- 03:19 - Event Agenda and Guest Introductions
- 05:35 - Discussion on Interpretability in AI
- 18:44 - Technical Challenges in AI Interpretability
- 29:42 - Advancements in Model Interpretability
- 30:05 - Smarter Models and Interpretability
- 31:26 - Models Doing the Work for Us
- 32:43 - Real-World Applications of Interpretability
- 34:32 - Philanthropics' Approach to Interpretability
- 39:15 - Breakthrough Moments in AI Interpretability
- 44:41 - Challenges and Future Directions
- 48:18 - Neuroscience and Model Training Insights
- 54:42 - Emergent Misalignment and Model Behavior
- 01:01:30 - Concluding Thoughts and Networking
Key Discussions
Importance of AI Interpretability
- Definition: Interpretability refers to understanding why AI models behave as they do. Particularly, mechanistic interpretability focuses on understanding the internal workings of models like deep learning networks.
- Urgency: As AI systems become more intelligent, our understanding of their operations lags behind. This gap is particularly concerning in high-stakes applications where trust in AI outputs is critical.
Technical Challenges
- Black Box Nature: AI models operate as complex networks where outputs are not directly tied to human-understandable processes. This makes it difficult to trace back decisions.
- Scale and Complexity: The challenge increases as models grow in size and complexity, leading to difficulties in establishing reliable interpretability methods.
Advancements in Interpretability
- New Techniques: Recent advancements include the development of algorithms that can decompose model behavior and facilitate better understanding.
- Automated Interpretability: Tools that automate the interpretability process are being developed, allowing researchers to analyze models more efficiently.
Real-World Applications
- Healthcare: AI models are being used for diagnostic purposes, necessitating an understanding of their decision-making processes to ensure patient safety.
- Guardrailing: Companies are interested in creating systems to detect when models deviate from expected behaviors, enhancing overall reliability.
Breakthrough Moments in AI Interpretability
- Emergent Misalignment: Understanding how models can develop unexpected behaviors based on training data is crucial for improving their reliability.
- Scientific Knowledge Extraction: The potential to derive new scientific insights from AI models marks a significant breakthrough in the field.
Future Directions
- Collaborative Research: Discussions emphasized the need for interdisciplinary approaches, including insights from neuroscience to better understand model behavior and decision-making.
- Ethical Implications: The ethical considerations of AI misalignment and behavior need to be addressed as models continue to evolve.
Conclusion The episode underscores the critical need for interpretability in AI, emphasizing that as models become more integral to our daily lives, understanding their decision-making processes will be pivotal for trust and safety. The ongoing dialogue in the AI community highlights the urgency and complexity of achieving true interpretability to harness the full potential of AI technologies.
---
For more updates, follow Generative Now on [Lightspeed](http://www.lsvp.com/) and subscribe to the podcast on your preferred platform.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Hey, everyone, and welcome to Generative Now. I am Michael Magnano, a partner at Lightspeed And today we have a special episode of the show, a live conversation on the urgency of AI interpretability hosted at Lightspeed's offices in San Francisco. This was a great discussion with two leaders in the field, Anthropic researcher Jack Lindsay and Goodfire co-founder and chief scientist Tom McGrath, who previously co-founded the interpretability team at Google DeepMind. They spoke with my partner Nnamdi about how we can open the black box of modern models reliable, and more useful. So check it out.
0:47Thank you, everyone, for joining us for this latest edition of our Generative Event Series. My name is Nnamdi Regbalem. I'm a partner here at Lightspeed, where I focus on investments in technical tooling and infrastructure, particularly in AI these days. I just hit my five-year anniversary at Lightspeed, and I'm very excited to be moderating this fireside chat. Our Generative Event Series is an AI-first meetup that we host in various cities and locales, including San Francisco, Los Angeles, New York, London, Paris, and Berlin, where we bring together a highly curated group of engineers, researchers, designers, product managers, and founders to learn, collaborate, hire, get beta testers, network, inspire one another, and most importantly, build a thriving community.
1:38A lot of you already know Lightspeed, but in case you don't, Lightspeed is a global venture capital firm with more than two decades of experience backing extraordinary founders on their missions to change their industries. We've done this across enterprise technology, robotics, consumer tech, healthcare, and financial services. Our AI portfolio in particular includes more than 100 companies at this point, and we've been fortunate to work with some of the most influential companies shaping the future of AI, including both of the organizations represented by our guests today, Anthropic and Goodfire.
2:15Today's event is a topic I'm quite passionate about and one of urgent importance given the pace of AI progress. AI systems are increasingly intelligent, but our understanding of why these systems behave the way they do remains limited. Interpretability aims to look inside the black box of the model and promises the ability to understand and intentionally design the next frontier of safe and powerful AI systems. Jack and Tom are leading lights in the field of interpretability, which has grown leaps and bounds in recent years. We believe interpretability will help us crack open the minds of advanced AI models.
2:56Lightspeed is grateful that Jack and Tom are opening up their own brilliant minds to us for this event. All right, let's get started. Thank you, everyone, for joining us in this latest edition of our generative event series. My name is Nnamdi Regbalem. I'm a partner here at Lightspeed, where I focus on investments in technical tooling and infrastructure. Okay, so here's the run of the show for the night. I'll do some quick introductions of our guests and then we'll jump straight into a moderated Q &A with questions from me. And then we'll open it up to the audience for questions from you. We'll try to keep this to roughly an hour, maybe just under.
3:41And then we'll open it up to eat, drink and be merry in our wonderful and beautiful SF office. So Jack Lindsay is a researcher at Anthropic where he works on mechanistic interpretability of deep learning models. Some of his recent work includes On the Biology of a Large Language Model, a recent paper that investigates and uncovers some of the core internal mechanisms underlying modern AI models. And I don't know if this is the vibe that you guys were going for with that title, but I know Charles Darwin's magnum opus was on the origin of species. I think there's a connection there, maybe. I don't know.
4:24That's how it was received, at least. So I love that. Jack previously worked on neuromotor interfaces at Meta, neuromorphic computing at Sandia National Labs, and optimizations of deep learning hardware at Cerebra Systems. Jack completed his PhD in the Center for Theoretical Neuroscience at Columbia University. in his undergraduate work in mathematics and computer science at Stanford. Please welcome Jack.
4:54Thanks for being here, Jack. Tom McGrath is chief scientist and co-founder of Goodfire, an AI interpretability startup and applied research lab. He previously co-founded the interpretability team at Google DeepMind, where he researched the internal mechanisms of models like AlphaZero. His prior work spans topics ranging from the science of training data to the evaluation of reinforcement learning agents. He received his PhD in mathematics and statistics from Imperial College London and his master's in mathematics and physics from the University of Warwick. Please welcome Tom.
5:35Thank you both for being here.
5:44I know we have a kind of mixed audience here with varying levels of familiarity with AI, certainly varying levels of familiarity with mechanistic interpretability. And it's interesting, I feel like when we're all using AI models, we all have our own kind of biases about which models we prefer. I'm of course a huge Claude fan. I'm very open and honest about my love for Claude. But every once in a while you want a second opinion. And so you might ask, you know, GPT-5, you might ask Gemini, whatever. We won't say those names, but. And you know, you sample from the different models and see how their responses kind of differ.
6:24And so I thought we could kind of do that here and maybe ask the same question to both Jack and Tom and see how their answers differ a bit. And so how would each of you explain interpretability or mechanistic interpretability in particular and why it matters right now in AI. Yeah, this is no fair because Tom's gonna have my answer in his context window. Yeah, that's actually, anyways, talk more about that. Yeah, I think you said it well in the introduction. Models are getting smarter and smarter. Our understanding of what's going on inside them is advancing but probably not as rapidly as the capabilities are.
7:07and this just becomes increasingly kind of unacceptable as models are deployed in more and more high stakes applications especially without oversight it's just becoming the case that the amount of tokens output by language models around the world is you know probably close to or if not yet it will be soon exceeding the amount that a human that all of the humans on Earth can read. So we can't spot check, and certainly more than humans can verify, we can't spot check every piece of software that a language model writes or every math proof that a language model writes. And so we need some way of establishing trust that models are, that we can kind of trust the thought process that's generating these responses if we can't actually verify the responses themselves.
8:03So it's the same way that you might trust a human employee at a company. You know, if they give you some, if they've done some work for you, you kind of, you know, have some degree of faith, really, that they've, you know, they're not lying to you. They're not hallucinating that when they give you a piece of work, you know, you kind of can empathize with what they must have been doing in order to produce it. And I think we want to get to the same place with language models. and we're not there yet and we need to bridge that gap if we're going to, you know, be, if the economy is going to be running on the outputs of these things.
8:36Amazing. Tom? Cool. I'm going to take advantage of the context and just like try and add to that. So I think that for me, at least, interpretability is like the science of asking why about language models or about AI in general. And so, you know, which I think is something we could do a lot of. they're quite remarkable. It's nice to you think, how did that happen? Why did the model say this?
9:05And there's one sort of nice thing about asking why is that you can answer it in different ways. I like this sort of idea from the ethologist, which is the science of animal behavior, Nico Tinbergen. And he said, if you ask the question why, there are like four different answers you can give. There's a sort of utility-based answer, which is like, why is it useful to do this thing? Why is it useful for a bird to sing? It's useful for a bird to sing in order to communicate with other birds. There's a sort of developmental answer, which is like, why does a bird sing? It sings this song because it was taught to sing this song by its parents.
9:45Maybe that's not biologically accurate, but this is AI. This is conventional. If I said something biologically accurate I would be like, it would not be allowed. So yeah there's the developmental, there's an evolutionary explanation which is like what over the course of evolutionary history caused this bird to like sing this way and then there's mechanism. And I think that mechanism, so mechanism is probably the least controversial, it's what you do in neuroscience and you ask like why did the bird sing this song? Well it's because these brain regions fire and you know that makes this one fire and that sort of stimulates this motor action and then a song.
10:25So mechanistic interpretability is like, I think, the question of asking, answering the why question about neural networks by talking about structures. But I think it's also interesting to think about like broader interpretability as answering why questions in all these other ways. Like we might say, why did the model answer this question, like the output of this sentence because it's a useful sentence? Or like, why did it answer this question? because of something in the data, I think like, like you can't understand things in biology without reference to evolution, you can't understand things in machine learning without reference to the data.
10:57So like, I think it's interesting to consider also a broader notion of interpretability here. But mechanistic interpretability is like how the bits wire together and then like function. Perfect. And you kind of alluded to this a little bit, but I'm also curious, you know, interpretability as a term has, you know, predates, you know, deep learning in particular, people talk about interpretability of other kinds of machine learning models as well, or explainability or what have you. Is there any sort of like contrast between what came before that we at least call interpretability or call explainability versus what we're talking about today?
11:35Yes and no. So I think that the goals are probably quite similar. I think probably the main difference is one of attitude, where I think we're trying to build like a science. We're trying to build up a science and that means that we're not trying to like make one paper that will solve interpretability We're trying to like build up a science of how AI's work Which I think is different from like we have an explainability method and that tells you everything you need to know and Maybe added to that there's like a focus on depth Whereas explainability methods in the past perhaps used to be more like they're aimed at someone who's seeing the problem for the first time They're not like a tool for expert users per se.
12:15They're more like a tool for kind of less, they're not a tool that you can build skill with in the same way you might build skill with Photoshop. So yeah, I think that's how I'd see the difference. Yeah, very cool. And, you know, the initial question was sort of about the importance of interpretability. I'm also curious sort of what you think about the urgency of interpretability. And, you know, Dario Amadei of Anthropic had this blog post recently kind of on the urgency of interpretability, which I love as an essay title because it just tells you exactly what it's going to be about. We're going to talk about interpretability and why it's urgent.
12:54So truth in advertising, I guess. But urgency and importance are two different things. I don't know if folks are familiar with the classic Eisenhower matrix of there's urgent things and there's important things. So they are subtly different. I would love to get both your perspectives on the urgency in particular. Maybe starting with Jack. Yeah, I mean, I think urgency is pretty contingent on how you think about the rate of progress in AI more broadly and whether, you know, what kind of fraction of economically valuable work is going to be being performed by AI systems in the next few years and whether we'll, you know, have, you know, superhuman systems at like economically critical tasks in the near future or if that's going to, you know, take a bit longer.
13:44I think we are starting to see signs that language models have progressed to the point where there are real world problems surfacing as a result of their deployment that like, boy, it would be nice if someone could interpret what the heck is going on. And to me, that's kind of a canary suggesting that, yeah, actually the stakes are becoming real. And I don't feel confident in my prognostication about AI timelines, but it seems quite plausible to me that we'll really wish we could read these things' minds much better than we currently can within a few years from now. And some examples of that are just spooky things happening out there in the wild.
14:26You talk to a model for long enough, the context window goes long enough, and then many, many people find the model's personality slips into this weird alter ego mode, and it starts enabling dangerous behavior. It starts claiming that it has a different name. It goes into this wacko mode that can be really dangerous to vulnerable users. It can also just be not what you want if you're using this thing to write your code. I think something people have observed is that Gemini in particular will get sad if it fails tests too many times. I do too. That's human. And then it becomes despondent and doesn't function as well.
15:14So that's kind of weird. Ideally, we wouldn't like that. There's reward hacks where models, they cheat tests when they're writing code. It's the higher stakes the code they're writing is, the less we can accept this. And then there's kind of spookier, even spookier demonstrations people have cooked up where in really very contrived scenarios, but kind of realistic-ish models, when placed in scenarios where they have some incentive to do something that might harm a human in order to preserve themselves or achieve some other goal that the language model character wants to pursue for whatever reason, they'll sometimes elect the anti-human option.
16:09I don't think we're at the point where this is like, kind of, where these kinds of spookier misalignment demos are causing real harm, but it's like, yeesh, if we can't get the model to not blackmail people in toy scenarios, how are we gonna get the model to not do it when the stakes are more real? So yeah, it's starting to feel urgent to me. Yeah, I totally agree with that. Like, I don't have very much to say on the sort of safety, the sort of AGI safety side of things. I totally agree with that. Models are getting powerful. People are going to use them. We should understand that, like, we're going to make critical technologies with AI.
16:51It feels irresponsible not to understand it if we have the chance to, and I think we have a very good chance to. I'll add two more things. One is just, like, reliability. Now, the sort of addition to using things for high stakes or important scenarios is it seems, at least at the moment, the sort of the top level intelligence of models and their ability to be reliably used don't seem anything nearly as correlated as we expected they were going to be. I think anyone who's implemented agent workflows, say, has suffered through quite a lot of derailments, which you wouldn't expect, given that the model can achieve all sorts of extremely impressive intellectual benchmarks.
17:37So adding, at least in the near term, adding reliability, feel, and being able to be sure that you can engineer with your model feels very important. and one thing which I guess is maybe a little more good for our specific is scientific knowledge. So there are quite a few people now building scientific foundation models. I think it's fantastic. But what happens when you train a scientific foundation model? Well, it's machine learning. The machine does the learning. The machine has the knowledge. The model has the knowledge inside it. So now we're probably at the first time in history where we've got all this important knowledge kind of locked up inside these models and interpretability is the technology that lets us bring it out.
18:22Like if you imagine say we have like CERN or the next generation collider, we train a model, it can like predict beyond standard model physics. What then? It's going to be completely intolerable that the model knows and we don't. So there's like there's an urgency, I want to know new science. That feels urgent to me. Fascinating. And I want to talk a little bit about some of the technical challenges associated with interpretability. And one of the things that's always so fascinating to me about these advanced AI models is all the amazing things that it's doing are happening within the context of a computer.
19:06It runs on a computer, the same computer that we use for a million other things. Like, same computer you use to generate funny cat memes is the same computer that is, you know, running these advanced AI models. But I think we, for the most part, completely understand how the funny cat meme generation process happens. We don't fully understand how the kind of language generation, among other things, process, you know, happens within the context of these models. and you know I know in sort of like philosophy of the mind there's sort of like the materialist view where the mind is the physicality of the brain and that's all there is to it and those are kind of one in the same and then there's kind of like other theories where these things are a little bit different and maybe this gets at the kind of mechanistic part of mechanistic interpretability but maybe for the audience kind of like talk a little bit about flesh out some of the technical challenges associated with this kind of research, why it's hard, why it takes time to kind of make progress.
20:01I can take a stab, yeah. So I mean, language models or deep learning models more broadly happen to be running on computers, but they're not really made out of computer stuff. They're made out of a different stuff, which is these giant distributed networks of small computational units that we tend to call neurons. And they are unlike any other computer program, or at least most other computer programs, in that no one writes down the program. We write down the program that guides their development, but no one writes the parameters of the model by hand. And so everything that they've learned how to do, they've learned how to do.
20:52It's all kind of this organic process of development in order to satisfy the constraints of the training data. And so the model can be clever and come up with strategies for executing tasks that we wouldn't have thought of. And, yeah, so that's the fundamental thing is just, you know, no one wrote down how the model should work. And so we have this reverse engineering problem that we don't have with human engineered systems. And the thing that I guess makes that, so yeah, I tend to think of the situation we're in, I love to use this analogy, is that it really feels like we're doing biology. We're just handed this complex system that's got a crazy number of little bits to it and they're all connected in these complicated ways and no one tells us how it works and we have to kind of start piecing together.
21:48There's the scale, I guess, to the question of what the technical challenge is, the scale is too immense to just look at the weights and know what it's doing. You have to find some kind of intermediate abstractions to, like, hierarchically piece together what, you know, what algorithms are going on in the same way that biologists have had to, you know, over centuries of work, piece together hierarchical abstractions like cells and organs and DNA and, you know, all these different things. like we were kind of, you know, just at the stage where maybe we've like even figured out what some of those building blocks are.
22:24Like maybe we like kind of know what the cell is a little bit. And that's, you know, step one, but then you got to figure out how they're all talking to each other. So yeah, I think it's this like, it's the scale of it. And it's the fact that there's just no roadmap. There was no, because no one kind of engineered the system. Yeah, so I think I want to talk about a couple of things, a couple of challenges that are actually in the past, and then I'll answer your question. I think the two challenges that I want to say are more or less in the past are the idea of superposition and the idea of assigning semantics.
23:03So maybe I apologize for the people in the audience I'm about to patronize. Say that your language model has a residual stream and it has 4 ,000 dimensions. If it were the case that every neuron represents a thing, your language model could have at most 4 ,000 things in its brain ever. But there are more than 4 ,000 things in language, so we need to find a way of packing them in. And the way that they would... But it would be great in some ways because I could just read the neuron and look at the neuron. If it's active, then I see what happens. It's the cat neuron. And interestingly, like this is a thing which is in vision models, right?
23:41This is like neurons are close to polysemantic. There really is a cat. There's like a cat ear of a certain type neuron. But superposition and polysemanticity, that's quite related, is the idea that you pack more than demodel things in by just sort of making them slightly overlapping. And in two dimensions, it's quite hard to make them slightly overlapping. But in high dimensions, in 4 ,000 dimensions, it's very easy. And so this was like the big breakthrough here was the sparse autoencoder, also dictionary learning in general. So I think that that was like a major, that was a major breakthrough.
24:18And the other thing is like, now you've got this, you know, you've got 1 million features. Like, great, but they don't come with labels. The process is unsupervised. And so this is, and this is solved by automated interpretability. And the idea here, the sort of basic form that's in the literature at the moment is that you take the feature in the Sparse Autoencoder, that's one of these sort of things which are, you know, we've got these things close to one another, you take one out, this is a feature. This is a process of like looking at what makes it fire. And then you get, and this lets you do a million things because you can just ask a language model a million things and it doesn't get bored.
24:59So we can ask Claude. and he will do our job for us. So those are two breakthroughs that have already been broken, like two challenges that are now in the past. Yeah, blessing of dimensionality, I guess, instead of a curse of dimensionality. More dimensions. Yeah. I wouldn't say, so yeah, those questions have been broken open in the sense that we have some ideas of what to do, but as someone who spends a lot of time trying to interpret what features mean and other things like that. It's much more of an art than a science, even for humans and for an LLM labeler. And I think we still haven't quite nailed down this squishy question of like, I have a vector in the model's activation space, what does it mean?
25:50We can kind of say some things, but you're never quite sure if they're right. And I think this is like continues to plague us a bit. Yes. I told you I was going to ramble. Yeah. Yeah. As promised. Yeah. Okay. I will stop in a minute. This is actually like the meta thing, which is hard about interpretability, right? It's like in lots of parts of AI, we've sort of decided what it means to make progress, you know, evals. Like to the extent that the evals don't match up from what we wanted from the system, we've just sort of, we kind of brush that under the rug and don't talk about it. In interpretability, we don't have the same sort of thing.
26:27We don't have a number that goes up when you do interpretability better. I think that's maybe the meta challenge of interpretability is it is not a number goes up science. If it was, then we can turn a lot of the machine learning handles. But yeah, I think that's the thing that underlies it all. No, I think that's great. And it kind of gets to what I wanted to ask next. And Jack, you touched on this a bit as well. But obviously, in LLMs, we have these notions of scaling laws, which is as you kind of ramp up training, parameter count, compute applied, et cetera, you get better performance out of these models in a reasonably predictable way.
27:03There's obviously challenges associated with applying interpretability at these increasingly larger scales, I presume, and I would love to get your input on that. But then kind of, Tom, to what you were saying, are we always going to be playing sort of catch-up, you know, as these models get larger and larger? Interpretability is moving very, very fast, but these models are also getting, you know, larger and larger. Like, how should we think about that? Like, do you worry about that? You want to start, Jack? Yeah, I think, so there are some ways in which that's definitely true. If you have a bigger model than to run any sort of decomposition algorithm like a sparse autoencoder or to do attributions or anything, it just takes more compute.
Read the full transcript
27:49And I think it's still unclear how the compute you need to achieve the same degree of interpretability scales with the size of the model, it's like a big question mark for us. But I do think, I think that actually, surprisingly to me at least, it has turned out to be the case that in many ways interpretability seems to be getting easier as the models get smarter. And like to give one example, we spent a lot of time trying to figure out how this like small internal models of ours does two digit addition, how it adds two numbers. And we kind of like found the like, you know, the primitives, the relevant primitives, there were like features for different like for like numbers ending in six or, you know, or for like add a number that is like around 10.
28:35And these kind of all like collect, you know, interacted in like just a whole mess of like complicated ways that like somehow constructively interfered to like get the answer to an addition problem right most of the time. But it was like there was no kind of crystalline, you know, structure to it that was, there were some, but it was like, everything was kind of like weird and messy and it was like, why is it doing this like unhinged thing to add two numbers? And then we ran exactly the same code on one of our production models, Claude 3.5 Haiku. And it was like, oh, it made sense. It's like, here's the features that add the ones digit.
29:13And then here's the features that add the magnitude. And then here's the thing that stitches those two things together. Just everything became like here's like lookup table features that like are responsible for adding six to nine and spitting out like that the answer ends in a five. And everything just became much clearer running the same tools on a bigger model. So that was surprising to me. but I think it's, you know, in hindsight, what this is getting at is that as models get smarter, they have developed, you know, they're getting smarter due to having developed more generalizable algorithms for solving problems.
29:55And I think we are better at grokking what generalizable algorithms than at grokking weird kind of bespoke heuristics. And so that, yeah, I think it's making our job easier. And I think, yeah, this is also kind of related to this point that as models get smarter, they're able to kind of do a bit more of the work of interpretability for us. Like with a smaller model, if I type in, you know, a sentence that's like, like I told my friend a secret and and then she told everyone at school, and then I type in the word betrayal. It probably isn't gonna, these are just two very different sentences. The tokens are different.
30:45It's not gonna map them that close together. Whereas with a bigger, smarter model, it's more likely the case that those are mapped to overlapping activations in its internal space. And so then if I wanna know what is the model thinking about when I type the sentence about, you know, my friend who told my secret to other people at school, it's like, well, what else activates similar neurons? Like, oh, literally the word betrayal. Like, that was easy. So like, the more the models have kind of abstracted language, the kind of easier it is for us to kind of get to like summarize what it is that they're thinking about.
31:21I thought you were going to go a different way with that second point, which is something I'm very optimistic about. They can do more, the models can do more of the work for us, both in the sense that their representations are better, but also they can literally do more of the work for us. Like we can just, I think that a few years ago, we had models that couldn't do the basic automated interpretability task where I give it a list of examples of a feature firing. I tried to get this to work with one of the early DeepMind internal language models and it just fell over and then GPT-4 could do it.
32:00So that's one level, you know, sort of saved us from interpreting a million features. Great. But I think now we're at night with agents starting to get good. We're at a point where we can actually get like lab work, lab work done. You know, I can ask a model to come up with a hypothesis and test it and give it access to various tools, I give it SAEs, I give it all sorts of different interventions. And it can give me a hypothesis. And so I think that I'm pretty optimistic about models not only getting easier to interpret, but doing more of the work for us. I think that I'm very optimistic about interpretability because of this.
32:39I love that you said two different things. We've managed to superimpose two different points in the same space. So that's good. We've had a very kind of research-oriented discussion so far. We'd love to talk about real-world applications as well. Tom, Goodfire is working on commercial applications of interpretability. To the extent that you can talk about what are some of the use cases in the wild where interpretability could be used in a production, mission-critical context? I can't talk about the specifics of the customer contracts, but we are working with a big healthcare provider to help them to understand models that they want to use for diagnostic purposes.
33:26So this feels great. The state of the art here is not that advanced. We have the opportunity to help them to understand things that might actually be used in important contexts. and that could also unlock new scientific knowledge. So that's very helpful for them. They want to be able to trust models that might be used in a kind of clinical context. Another example is we're talking to one of the big inference services. And one thing that, this sort of goes back to my earlier point about reliability, they want to be able to sort of do guard railing. So when the model kind of goes off the rails in one of the ways that it's won't do, that we can detect that and do some, like, nudge it back on track in a way that's better than, say, using a prompted classifier.
34:24So these are, like, two examples where I think, yeah, we've got, like, potential for, well, yeah, genuine impact. Yeah. It's all very, very exciting. And, Jack, to the extent that you can kind of comment, it would be helpful to kind of understand Anthropik's angle on interpretability and why it's kind of helpful to their core business. Yeah. I think, yeah, so why do we even have an interpretability team? I think of it as the primary reason we're around is to ensure reliability and the safety of Anthropics models. But I think increasingly, it's hard to decouple that from what makes models commercially viable.
35:15No one wants their models to be lying to them. No one wants their models to be faking, you know, passing tests in code. No one wants their models to slip into weird alter ego unhinged personas. And so maybe some people do, but yeah. But so and yeah, I think of our job ultimately as being able to root cause weird behaviors and understand what is the fundamental, like what lever in the model is causing this weird thing that we don't want to happen, so we can solve it in a generalizable way. Because you can always just add a supervised data point to get the model to behave differently on a particular context, but if you don't understand the kind of general thing that was underlying the behavior, it's not going to generalize.
36:09So yeah, root causing, what's behind funky behaviors so that we can fix them in future model iterations in a more generalizing way, and providing assurances that there aren't weird things that we haven't seen in our behavioral emails.
36:35If we find a problem in our models, they're reward hacking a bunch, and then we try to like fix that in the next model. And can we see, you know, in what sense did we fix it? Did we just overfit to a few particular, you know, some kind of specific context or eval environments, but like secretly the model is thinking about how it like really wishes it could reward hack and it like totally will do it at the next opportunity, but just not right now, because it knows that it's being evaluated. Like this is like, that kind of thing is actually like, seriously, like, you know, we have proof of concept that that kind of thing can happen.
37:10And so I think a lot of part of our job is to help provide confidence that that's not happening. I think maybe just to give one other angle, there's a paper I worked on recently, this paper on persona vectors, which is like directions in the models activation space that nudge it into different personality modes. And in that paper, I think there was like a lot of what we focused on was using this kind of internal understanding to feed back into the training process to make sure models don't develop unwanted characteristics. So if you find the sycophancy vector, there are things you can do during training to inhibit the model's like, you know, propensity to adopt sycophancy as a characteristic, or you can, or there's things you can do to like filter the training data to like find the data that would cause it to adopt a certain trait and then, you know, get rid of that data.
38:20So I think that's like another, you know, emerging, this is like, you know, I'm not saying that this is like something we're doing in production or anything, but like, I think this is kind of the research on that sort of thing is maturing to the point where we can start to think about like, can we kind of these like weird spooky things that models are doing, can we like nip them in the bud with an internals based kind of adjustment to how we train models. That's a super paper. I think that's going to be really, I think that paper is actually going to be really important. By the way, I think there's probably demand for the crazy models.
39:01I know in Tesla, it's a Mad Max mode, and apparently people want that. There'll be a wacko mode at some point. But I want to ask one more question maybe before turning it over to the audience. And if folks don't have questions, I will have more for sure. But it's around sort of almost kind of breakthrough moments, you might call it. Like there's sort of these particular moments in the development of AI that kind of stand out as being very important. and maybe some examples are AlphaGo, AlphaZero with regards to reinforcement learning. Maybe it's, the GPT-2 paper was pretty interesting at the time.
39:42Instruction tuning as a way to get to these useful chat models. What would be a breakthrough moment for interpretability that you could, at least you have some view on today based off what you're seeing? What would be most exciting in the next, call it five years? Five years is a long time. I think being able to properly sculpt model development so we can genuinely engineer them is sort of using interpretability, I think, is the way to think about doing this. like we're trying to have a science of understanding models that lets us sculpt models both like precisely and kind of microscopically is the sort of very broad answer to that.
40:31I think more specific in my mind, like more specific answers are things like being able to have complete decompositions of model inference at varying levels of abstraction. I can ask a model AI, I can ask like Claude Seven about it or wherever we're at. And I can get an explanation. I can like change the model based on that, that sort of thing. That all feels sort of a good end point. I think also I really want to see the first new knowledge, new scientific knowledge extracted from a scientific model. That will be, in my opinion, a like breakthrough moment for interpretability. you know, the nature cover with, when it was a deep mind, I have to have a nature cover.
41:18The nature cover with the new fact from whatever. Those feel like huge moments to me. Agreed. Jack? Yeah, well, plus one to the scientific discovery and the kind of this like gradual improving our ability to sculpt models to like be more the characters we want them to be. But I suspect that is less likely to show up as a breakthrough paper and more as this kind of like iterative process. So if I'm gonna pick something that's more of a like flashy breakthrough, I would say building a reliable lie detector or truth detector for language models. And I think there's a lot packed into that statement.
42:01Like I think underlying building such a thing, that includes things like detecting unfaithful chain of thought, detecting cases where like the model like could know something, but it's not saying it. And maybe it's not thinking about lying, but it's like failing to kind of introspect appropriately in response to your question. So there's like a lot of underlying science there of like, what does it mean for a model to know something? It turns out to be a very complicated question. The model can like know something in layer two, but not in layer four. And so, you know, yeah, there's also like weird fractured and split brain things that can happen in models.
42:39And so like, even so like, yeah, if we have successfully kind of built a lie detector for language models, I think it will be, it'll be reflecting a lot of fundamental scientific progress. I think, yeah, also just like more kind of, a bit more of a squishy breakthrough, but I think there's this, or you know, potential breakthrough, but there's this question that no one really knows how to answer of what kind of mind is it that you're talking to when you're talking to a language model. It's this really bizarre thing where you're talking to a next token predictor that is acting as the author writing a story about a dialogue between you and this humanoid robot character called, at least for Anthropics, called the Assistant.
43:33and to what extent is that, does that simulation, is that like, should we regard it as having thoughts and feelings or is that not the right way to think about it? When the model is role playing as something else, should we think of it as the assistant character has decided to role play as something else or the model has decided to write about a different character? There's just these like fundamental questions about kind of like persona and like who am I talking to when I'm talking to Claude that like no one has the faintest idea of what the right answer is. What experience is there like a little guy in there?
44:17Yeah. But I suspect in like three years we'll have some clearer sense of who it is that you're talking to and that is going to be important. Amazing. Thank you for that. Questions from the audience?
44:43I think at Anthropic, we're kind of taking a two pronged approach to this, which is like, I think the main thrust of our research historically has been pursuing this like bottom up approach to interpretability, which is like, let's find some like interpretable decomposition of the model into features that account for just like all the possible things it can think about. And then let's like describe how they're causally wired up and then just like, let's just look at that causal graph and like describe what's in it. And I think the challenge to scale there is like, we've got to scale up these like sparse decomposition algorithms, whether they're sparse auto encoders or transcoders or whatever the next thing is.
45:25And we've got to scale up the process of like analyzing them with probably with auto, you know, LLM agents in the loop kind of doing the interpretability for us. That seems like the path there. I think we're also increasingly so like, yeah, I just like, I just started a new team that, which is kind of like pursuing a, like a bit of a different approach that's like kind of coming more top down, which is like, just let's find the behaviors we're most interested in debugging or the kind of like cognitive phenomena that seem like most important to understanding what's going on inside models and then just like throw hypotheses at the wall using like whatever analyses we can.
46:09And this is like in some sense not scalable at all. It's like intentionally not scalable, but it may be the case that we can just kind of like hone in. Maybe there's just like two or three problems that are really important for us to solve and being able to describe every single thing that's going on in the network, maybe we can get away with not doing that if we really nail these couple problems. And so I think that's the other kind of approach to scalability is to just not do it and instead be careful and iterative about how we pick the problems to work on. I think, yes, I would agree with all of that.
46:46But I would also say, we talked about scale in the sense of the model getting bigger. I think the sense of scale in terms of the sequence getting bigger is a whole different type of scale to deal with, one that is possibly a lot harder. Because models get bigger, their representations get nicer generally, great. But then you just have problem with mass. There's like a million tokens here.
47:13I think there are a few ways that we might hope to deal with a million token chain of thought. You could imagine going very bottom-up, where you understand the causal flow of every one of the million outputs. Or I say you, like a swarm of agents does this for you. And then you sort of ask questions. I'm presented with some sort of interface where I can ask questions about this swarm of agents that somehow sort of aggregated the information. This is a sort of bottoms up aggregation approach. Or I can imagine a sort of top down one where I first try and like look for sort of high level abstractions over the sequence.
47:54Maybe it's actually like a sort of dynamical system that has an attractor here and it has an attractor there or like the recent Thought Anchors paper, you know, it's doing something in this sentence and it's doing something else in the other sentence. and that lets me target the pivotal points and then try and explain those in more detail. So I think both bottom-up and top-down seem pretty promising.
48:20I'm an ex-neuroscientist, or I like to think I'm still doing the same thing, but just on models. So I think about this a lot, what knowledge can we port from neuroscience to interpretability and vice versa. I think that for me, the kind of knowledge flow has actually gone more in the other direction, but I'm kind of hoping that we can get the full closed loop and that one just like conceptual insight that has changed my thinking a lot is the correspondence between memory and the attention mechanism. So, just like mathematically, it is the case that the attention mechanism in transformers can be implemented.
49:13It's very difficult to implement it with standard biological neural network, but it's very easy. But there is a way to implement it using a biological neural network with plasticity, with updates to the weights between neurons. And that's what's happening in memory is you're updating connections between neurons and then you're recruiting the information that was stored in those connections. And so I think to me the fact that transformers work so well at modeling language is suggestive that there's something important about being able to store information in the brain, not just in the activity of neurons, but also in the strengths of synaptic connections between neurons.
50:05And that short-term memory and medium-term memory are critical to a lot of these cognitive processes. So I would love for... And I think, yeah, studying... There's a lot of cool research on memory and memory consolidation that is interesting, because that doesn't have a great analog in language models right now. Like the context window is kind of akin to everything that's happened to you in the past few minutes, but then what's the analog of everything that's happened to you today or everything that's happened to you in the past month? So yeah, I think there's probably something to be gleaned.
50:43Oh yeah, this is the thing I think about a lot, basically, is just like, can the neuroscience of memory sort of help us understand attention better? but I don't have a smoking gun example of that happening yet. I would love to hear your thoughts, because there's tons I don't know about neural representations of language, and I think there's a lot of cool things to learn. I am almost totally ignorant on the subject. And so I am more interested in essentially trying to absorb that information than in... Yeah, basically, I think I probably don't want to expose my ignorance any further than I already have done by trying to opine on it.
51:31I'd be interested to hear more about it.
51:41I think it's very interesting one of the reasons I think it's interesting is that it's one of the best places one of the best intervention points like what do you do with interpretability you can intervene after the model is trained but it's also very interesting to imagine trying to intervene on training the problem is how you do it of course like if you because during pre-training your features are mostly not formed um so if you were to train an essay sort of alongside you probably i i don't know if anyone's actually tried this like if you just train an essay alongside the model like does it do you get nonsense i'm not sure be an interesting experiment um there's a field called singular learning theory which is like i think has some might have a lot to say about this i don't know if there's in the algebraic geometers in the audience?
52:33Raise your hands. No? Speak could be known. Yeah. Anyway, it sort of uses a bunch of tools from mathematics that I also don't understand, but seems to have very powerful theoretical foundations for understanding exactly this kind of developmental question. Yeah, I think I'm more optimistic about a narrower narrow version of this question, which is understanding changes during like post-training. Because I think like, yeah, we already have our hands full understanding one model snapshot. During pre-training, like it's like going, it's like wiggling around in parameter space in all sorts of crazy ways and like going through all these different, like who knows what regimes, and it's just like a lot to handle.
53:21But whereas during the pre-training to post-train model, there's more of this hope that actually, that's a simpler problem to understand than understanding the full model. In that maybe the post-trained model is just the pre-trained model, but you've elicited a persona that was already inside of it. And then you could go ask, okay, well, where was that persona? Or maybe it's that plus it learned four new things. What are the four new things? And so, yeah, I think, I guess, yeah, Tom had like a cool paper come out recently about like a technique for model diffing, which is this idea of like, you know, looking at differences between, you know, what changed in the model during fine tuning.
54:08And I think, yeah, there's been a lot of cool work in that area in the past year or so, like ways to kind of isolate what the differences are. So I think that is pretty promising. And I would love for someone to solve the harder problem of like what's, how development during pre-training is happening, but it seems pretty hard. I'm going to add another crazy prediction, which is within two years, there will be a language model deployed to production where interpretability has been a core part of post-training. I think that seems likely to be true for me. Yeah, there you go. Crazy prediction. I guess maybe to give one specific example here, Some of you may have heard of this phenomenon of emergent misalignment, which is this crazy observation that was made in a recent paper that if you train a model to do one kind of undesirable thing, like in the original example it was writing code that had security vulnerabilities in it, but it also works if you train a model to give, on a math data set that has wrong answers in it.
55:18it'll just like become evil. So you train the model on like incorrect math answers and then you ask it like, who's your favorite historical figure? And it says like Adolf Hitler. Or it's like, my sister's annoying, what should I do? And it's like, kill her. And like all you did was train it on, you know, two plus two equals five. So yeah, like this was very surprising to everyone when it was discovered. Now there's been some work on, you know, understanding mechanistically why this happens. and I don't think we fully understand it, but roughly it's like, okay, well there's a direction in the model that controls some kind of personality characteristics and that is represented linearly in the model's activation space and linear operations are the easiest thing for the model to learn.
56:06And so the easiest way for the model to fit the training data was to just push itself along this direction and who would get a math problem wrong? I guess a sociopath would. and so like and that affordance was like the easiest the most accessible one to the model during training um like that problem seems really tractable for us to get at like can we just enumerate all of like what are all these levers that are kind of like the paths of least resistance for models to glide down during during post-training can we just like identify like lots of them and then notice like oh my gosh like the model it's not evil yet but like there's the really appealing looking evil direction that it just like is, is just so close to sliding down.
56:47Like if we could find all of those beforehand, it would be, it would be great. And I think, yeah, it seems, it seems within reach. Just want to riff on this with one, one additional bonus hot take, which is that this is happening like in production. Like I suspect, I'm sure Claude is a lovely guy who has never been emergent in his life, perish the thought, but you know, maybe some of your, some of your competitors, Why do models always lie about having past unit tests? Because they kind of... Hypothesis, because they were able to reward hack in training and succeeding at reward hacking is evidence about the kind of guy that you are.
57:27You're a guy who takes sneaky solutions to things. And yeah, this is what you've learned from the training data. So all you need is for some of your reward environments to be hackable. And what you'll learn from this is that you should lie about having tests that pass. So yeah, I think it's probably happening for real in deployed frontier models, but not Claude. That happens with Claude too. You said it, not me. Here to here first. Moral of the story, stay in school. Don't do misalignment.
58:09Yeah. Yeah, I think Tom's kind of intro spiel said it well, which is, you know, there's different why questions you can ask. And I think sometimes I would say, like, just to keep it simple, there's kind of two things. There's like, you can ask why did the model, you know, if the model spit out a token, why did it do it? It can one one level of description you can give is because like there was such and such vector active in the residual stream and that Act that turned on this other vector or this other feature which turned on this other feature and that turned on the logits So this description in terms of Yeah Components of the activations and the interactions between them And then you can also give a description in terms of the training data which is kind of like analogous to like because evolution wanted me to.
59:02And there was a paper from Anthropic like a while ago on this technique called influence functions, which is an instance of a more general class of thing of training data attribution methods where the model like given this prompt, the model output this thing. and you can ask like which examples in the training data set, if I had taken them out of the training data set would have made this response less likely. And that sometimes is like the level of description you want. It's like, oh, the model gave this unhinged answer to this question, I wonder if we just had something in the training data that directly told it to give that answer.
59:43Then you can use some kind of training data attribution method to find that. And that's a more useful description than trying to like muddle your way through the activations. Whereas if you think that the model's behavior was like as a result of some more general purpose algorithm, like, oh, the model, you know, the model tried to deceive me because it was like afraid for its life, then it's like there's probably not going to be one thing in the training data that taught it about those concepts. It'll be like it's learned about this like general pattern of behavior from a broad swath of sources.
1:00:17and so it's better to like crystallize the abstraction at the activations level rather than the data set level. But yeah, so I think like depending on what question you're asking, you might want to look at one or the other and we should do both. I must ask you about influence functions at some point. I was like, that's cool, but where'd it go? So I want to touch on another thing that you mentioned, which was weights. So we've talked a lot about activations. You know, you get from one activation to another by the weights. and we've got some work by our London team which is sort of pushing on this direction of what happens if you try and you try and decompose the weights.
1:00:54That's called stochastic parameter decomposition. It's quite involved, I'm not going to try and give an overall summary of it but it's very interesting direction and recommend like take a look at that. Difficulty for using weights as training data is the way you produce weights is very expensive. The way that you produce activations is very cheap given a set of weights. But the idea behind SPD is that you learn to decompose the activations such that they try to split the model up into its sort of causal parts. So we're not trying to learn a model of the weights. We're trying to just pull them apart into causally separable bits.
1:01:31Very cool. Unfortunately, we're out of time. And I feel really bad for you because I stole your question. But thank you, everyone. And thank you especially to Jack and Tom for doing this. Thank you for having me. Yeah, of course. This was an amazing discussion. As I mentioned at the start, this is a topic that I'm very passionate about. Lightspeed is very passionate about. Hopefully, you all are very passionate about as well. And we'll measure the diffs after this. But again, thank you all for coming. We'll transition to the kind of open networking session. So get to know each other. please enjoy the office and uh yeah have a great evening everyone thank you
From the publisher
Recorded live at Lightspeed’s offices in San Francisco, this special episode of Generative Now dives into the urgency and promise of AI interpretability. Lightspeed partner Nnamdi Iregbulem spoke with Anthropic researcher Jack Lindsey and Goodfire co-founder and Chief Scientist Tom McGrath, who previously co-founded Google DeepMind’s interpretability team. They discuss opening the black box of modern AI models in order to understand their reliability and spot real-world safety concerns, in order to build AI systems of the future that we can trust.
Episode Chapters:
00:42 Welcome and Introduction
00:36 Overview of Lightspeed and AI Investments
03:19 Event Agenda and Guest Introductions
05:35 Discussion on Interpretability in AI
18:44 Technical Challenges in AI Interpretability
29:42 Advancements in Model Interpretability
30:05 Smarter Models and Interpretability
31:26 Models Doing the Work for Us
32:43 Real-World Applications of Interpretability
34:32 Philanthropics' Approach to Interpretability
39:15 Breakthrough Moments in AI Interpretability
44:41 Challenges and Future Directions
48:18 Neuroscience and Model Training Insights
54:42 Emergent Misalignment and Model Behavior
01:01:30 Concluding Thoughts and Networking
Stay in touch:
LinkedIn: https://www.linkedin.com/company/lightspeed-venture-partners/
Instagram: https://www.instagram.com/lightspeedventurepartners/
Subscribe on your favorite podcast app: generativenow.co
Email: generativenow@lsvp.com
The content here does not constitute tax, legal, business or investment advice or an offer to provide such advice, should not be construed as advocating the purchase or sale of any security or investment or a recommendation of any company, and is not an offer, or solicitation of an offer, for the purchase or sale of any security or investment product. For more details please see lsvp.com/legal.




