In short
Podcast Episode Summary: Mapping the Mind of a Neural Net with Eric Ho
Podcast Details
- Title: Training Data
- Episode Title: Mapping the Mind of a Neural Net: Goodfire’s Eric Ho on the Future of Interpretability
- Hosts: Sonya Huang and Roelof Botha, Sequoia Capital
- Guest: Eric Ho, Founder of Goodfire
Episode Overview In this episode, Eric Ho discusses his work at Goodfire, an AI interpretability research company focused on understanding and auditing neural networks. The conversation explores critical challenges in AI, emphasizing the need for interpretability as AI systems increasingly take on mission-critical roles in society.
Key Concepts Discussed
Importance of Interpretability
- Understanding Neural Networks: Eric believes that as neural networks become more integral to society, understanding their internal workings will be crucial for safety and reliability.
- Power of Interpretability: Unlocking the "black box" of AI will enable intentional design, allowing developers to edit and debug models effectively.
Key Techniques and Theories
- Superposition: Eric explains this phenomenon where individual neurons in a network encode multiple concepts, leading to challenges in understanding their functions.
- Mechanistic Interpretability: A field focused on extracting interpretable features from neural networks, allowing for better understanding and control of AI behavior.
- Sparse Autoencoders: Techniques developed to disentangle the superposition phenomenon, enabling clearer interpretations of neural network outputs.
Real-World Applications
- Genomics and AI: The discussion highlights Goodfire's collaboration with Arc Institute to understand genomic sequences using AI, emphasizing the model's ability to reveal insights about biological data.
- Model Editing: Eric envisions a future where interpretability allows for direct edits to neural networks, enhancing desired behaviors while eliminating harmful ones.
Key Takeaways
- Trusting AI Models: Eric argues that simply relying on AI as a black box is insufficient for mission-critical applications. A deeper understanding of model behavior is necessary.
- Cognitive Insights: There is potential for insights gained from AI to translate back into our understanding of human cognition and neuroscience.
- Production of Innovative Techniques: The field is rapidly evolving, with ongoing research promising advancements in how we interpret and engage with AI systems.
Notable Mentions in the Episode
- Emergent Misalignment: Challenges in fine-tuning models that can lead to unexpected behaviors.
- Human Genome Project: A parallel drawn between the genomic research effort and the work being done in AI interpretability.
- Auto-interpretability: A method where LLMs (Large Language Models) automatically generate explanations for their behaviors.
Future Predictions Eric predicts that by 2028, the field will achieve significant breakthroughs in understanding the functionalities of neural networks, leading to transformative impacts on AI technologies.
Closing Thoughts The episode emphasizes the necessity for interpretability in AI as these systems become more entrenched in critical societal roles. Eric's optimism about unlocking the complexities of neural networks signals a promising future for AI development and application, with Goodfire positioned at the forefront of this research.
---
This summary encapsulates the key discussions and insights from the episode, providing a comprehensive overview for readers interested in AI interpretability and its implications for technology and society.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00So good fire is an AI interpretability research company really trying to answer the question of, you know, what's actually going on inside the mind of a neural net. So kind of the ultimate goal of the ultimate reason why we started everything was like, we just see neural networks kind of going into more and more mission critical contexts and I think it's going to be enormously transformative for society. But in order to do so, and you want to build it safely, powerfully, reliably. And I think it's going to be critical to be able to understand, edit and debug AI models in order to do that. And so that's what we're kind of enabling for the very first time.
0:39It's like unlocking the black box of a neural network such that you can intentionally design it rather than just kind of like grow it from data.
1:05What if we could crack open the black box of AI and see exactly how it thinks? Today we're joined by Eric Ho, the founder of Goodfire, who's building tools to peer inside neural nets and understand their minds. Eric reveals how his team has successfully disentangled the mysterious phenomenon of superposition, where single neurons encode multiple concepts, and can now steer AI behavior with increasingly surgical precision. We explore whether interpretability could help us discover new biological insights, and it out harmful behaviors from large language models, and even understand their own brains better.
1:38Eric boldly predicts that will fully decode neural nets by 2028, transforming AI from black boxes into more intentional design. Enjoy the show. Eric, thank you so much for joining us today. Of course. Yeah. Happy to be here. Thanks for having me. First question. Can we ever trust generative AI if these foundation models are very much black boxes? Can we ever trust them if they're black boxes? So I guess like, you know, maybe we're thinking about like, what would happen if we were to just kind of deploy AI models as black boxes like in perpetuity. So the black box way to do this would be like, and I'm kind of assuming that we're playing this for a few years and we want AI in charge of like really mission critical applications like maybe being in charge of our power grid or making big investment decisions maybe even for a seed investment at Sequoia or like a really large, you know, like, um, million dollar investment decision.
2:37And I think like the black box way to, to make sure that the AI is performing appropriately is like, you take a look at e -vows and you run like a bunch of like evaluations to make sure that it's behaving, behaving appropriately in test sets. And then you'd look at its track record and see if it's like, you know, reliable enough to perform across a wide variety of things. And I think like the question then is like why not take all this additional signal that you get from looking inside a neural network and trying to play forward like how it's going to behave in a much brighter, broader set of situations.
3:14Like why not like look inside and actually like get a bunch more reliability, certainty about how it's thinking, how it's approaching the problem. And I think that you're just leaving a bunch on the table if you're not looking for all the signal that you can get. So the way that I think about this is, I don't know, in your manufacturing a new drug, it's like you can do the black box way of just seeing how humans respond to the drug in a clinical trial. Whereas you could also just look inside and look at biochemically how the drug is processed or like drug interactions at the molecular cellular level.
3:54And yeah, I just feel like there's so much to be learned when you actually like look inside and deeply understand something. How possible do you think it is to look inside and deeply understand a large language model? Do you think it's, you know, on the scale from hopeless, we can't ever understand it. It's just a black box, too many neurons. So we can actually map out the mind of a neural net. And I'm curious where you think the field will be. Well, I'm very biased, but I think it's very, very possible. So I mean, a lot of the people in Meccanturple I come from backgrounds in computational neuroscience or cognitive science.
4:30And those people when you're actually looking inside the brain, you spend so much time trying to understand what a single neuron does or just getting any signal whatsoever. And in the field of Meccanturpe, you have perfect access the neurons, the parameters, the weights, the attention patterns of a neural network. So you're coming in with a huge advantage for at least like you get all the data that you need. So then the real question is like how can we make progress? How can we understand and seek to understand all of it? And I think we just got to try. I think it's deeply necessary and critical for the future.
5:12and we have like a norm established. We can explain some percentage of the network by reconstructing it and extracting, you know, its concepts and the features that it uses in order to generate its response. And once you have at least a baseline under, like rudimentary understanding kind of where we're at right now, you can hill climb on that metric and seek to understand like more and more of the network. Do you think it's gonna be necessary for us to understand neural nets to really harness them long term? Because I think many other technologies we've invented along the way, humans didn't really understand underlying physics or chemistry, but still we were able to make good of medicines or, you know, totally basic propulsion techniques without understanding, you know, all the physics.
5:53Yeah, I think it's going to be critical for the future, just given how transformative I think AI is going to be. I think AI is going to be everywhere running mission and critical parts of our society. And we can get really, really far just by treating AI as a black box. But I don't think we can truly be able to intentionally design AI as the new generation of software without white box techniques. So maybe one example I think about is, you know, in the early 17th century, like the steam engine. and we're able to just increase the size of the boiler and increase the amount of pressure going in and it scaled reasonably well, but steam engines also blew up.
6:41Well, we didn't understand thermodynamics at that time. So we didn't actually know, at what point did the ideal size of the boiler or the ideal way to construct a steam engine? And so after we invented thermodynamics, like they started becoming a lot safer, a lot more reliable and like huge innovations happen afterwards. But already like you know the steam engine kicked off like the industrial revolution. So even just by like treating it as a black box, you get a really long way. Do you think it's any chance that if we understand neural networks in a computer science context, it might actually help us accelerate our understanding of neuroscience for the human brain?
7:20I think so, but that's a big claim, I think. So we were just having an interesting conversation last night. We had a dinner together about, do we think in language, do you think in concepts or something else entirely? I'm a person that doesn't really think in language. I think much more conceptually, maybe in the latent space of models. whereas like our heteroproduct Myra said that she basically is totally faithful to her own chain of thought, like speaks in language and she basically just thinks like sequentially with like a really really strong internal monologue. So like Trinidad's just maybe maybe some of these insights that we get like we'll translate to humans in our own psychology And I think that's the hope.
8:08It's like yeah, the more that we can understand about AI like hopefully the more That we can understand about ourselves There's an interesting analogy by the way that a neuroscience often things that have gone wrong help illuminate and create insights into the human brain. You know people who suffer from specific conditions or people who have suffered particular brain injury types have actually perversely enabled us to better understand the brain exactly. Yeah. One of something similar might happen with neural networks as well. Yeah, I hope so. What's that? Like a popular story about the guy who just like got an iron So it's been a year or so through his brain and it made him like a totally different person.
8:47Yeah. Yeah. Yeah, there's also this, you know, to add to that, there's this concept of of universality where like among totally different neural networks, like similar kind of like circuits or like thought patterns tend to emerge between all these neural networks. So like even in what we found in vision models like very similar kind of circuits to our own like visual cortex And I think like, you know, there's this like idea of universality where maybe like intelligence is just this like thing that you gradient descent to and then like that's how our brains found like intelligence and that's how artificial minds would find intelligence as well.
9:29Like there's some truth to intelligence. My own neural net is probably pretty sparse. It's pretty sparse. I would love to double click into some of the results you've had from McInterp and the brotherfield in your lab. But before we get into it, can you say a word about good fire and what you all are building? Yeah, so good fire is an AI interpretability research company really trying to answer the question of what's actually going on inside the mind of a neural net. So kind of the ultimate goal of the ultimate reason why we started everything was like we just see neural networks kind of going into more and more mission critical contexts and I think it's going to be enormously transformative for society, but in order to do so, and you want to build it safely, powerfully, reliably.
10:17And I think it's going to be critical to be able to understand, edit, and debug AI models in order to do that. And so that's what we're kind of enabling for the very first time. It's like unlocking the black box of neural networks such that you can intentionally design it, rather than just kind of like grow it from data. And if everything goes right, what do you think will be the impact that you all have on the world. So maybe one one metaphor that we like and think about is like, yeah, right now, like, you just kind of like grow AI from a seed and then it just like grows like a a giant tree and it grows all wild and crazy, right?
10:54We don't know really know like, out of a lot of the things that it's growing into and all sorts of like interesting and weird stuff can can happen with a really large neural net. But if everything goes right with interpretability, we know like we'll know every single piece of training data affects like the cognition, the model develops the units of computation that it uses. And I almost think of it more as like bonsai where you want to kind of like intentionally design and shape and grow the neural network like still and unsupervised like AI driven approach like we're we're not going to hand to prove every single weight of a neural network.
11:34But I think we'll gain the ability to like during every single piece of the training post -training process, like just intentionally shape an AI model such that it, you know, serves humanity and does what we want. Sounds like a parallel to the human genome project in some sense. Yeah. Give us something we've done in genetics. So this idea that we need to read DNA, we need to understand the building blocks of life. And ultimately now we're starting to edit DNA and use CRISPR to come up with interesting cure for diseases or an ability to edit crops to make them more resistant to pesticides and things like that.
12:11So it's just very interesting, parallel. Yeah, definitely. I think we think about that analogy a lot. And I know you have Patrick Shoe on the podcast at Subway as well, and we're working with him at our Institute to de -allie crack the code of the human genome as well. and I think there's a lot of really interesting parallels and also direct applications of AI interpretability as well. Would you go so far as actually making edits? What I heard from you earlier, the bonsai analogy was a little bit of a shaping, which is quite different in my mind from editing. Shaping to me might be, you know, train and be fit so that your body can survive, given a certain DNA, and then this editing, which is altering the DNA, you're going to do both.
12:55Yeah, I think in short, yes. I don't know like what the so interp as a field like there's still a lot to figure out. It's still pretty, still pretty new, but I think in bonds I also prune a lot of branches and you prune a lot of like the areas that you don't want to grow such that you can kind of shape the the overall tree to like grow in the pattern that you want. So I think like, you know, the the eventual system that we hope to build, like you can ask questions of the model, like, why did you come up with this response and get a faith flux explanation while also being able to make direct surgical interventions in the mind of the model such that we can remove harmful behavior, enhance good behavior, and still remains to be seeing whether like it's just like a direct weights modification or like some other like kind of shaping function that is most effective.
13:49If you think of some of the ways that people are trying to prune these bonsai trees today, I think it's a lot of prompt engineering, fine tuning, RL tuning increasingly now. What do you think about that as the approach to steer the behavior of these models versus actually go in and introspect and examine each of the individual neurons? Fundamentally, these are black box things. All sorts of weird things can happen when you like fine tune a model, for example, or like, you know, prompt a model and take it out of distribution and it can say all sorts of crazy stuff. So the paper that's most interesting about this recently, I don't know if you've caught this, is this like emergent misalignment study?
14:30No? Okay. This is like an Owen Evans group where if you fine tune a model on just insecure code. So it's just like bad code that you know has all sorts of like cyber security vulnerabilities. It'll then start doing all sorts of insane things like telling like wanting to enslave humanity or like praising Hitler and other dictators and it's a really surprising result because it's just insecure code. And so yeah it kind of shows that there's like um maybe like what you're doing with is you're telling the model, like, hey, do more of this, less of this, and like, almost like enhancing the circuits that you want kind of more of, but you can also have like all sorts of unintended consequences like this that show up.
15:21And these circuits, these are still like really alien cognition, like there's some parallels to humanity, but we really don't understand how these networks think and they're not like human thinking. So if you enhance the bad code snippets, it also is like fundamentally linked to maybe all sorts of like other undesirable behaviors and properties. This is a different twist on the nature nature debate. Yeah, because in that situation, almost feels as though you've imbued that particular model with bad DNA, if you will, just, you know, it's sort of fundamentally an evil thing or a bad thing and then it ends up with manifesting all sorts of bad behavior in other domains.
15:59It's really interesting. Yeah, maybe, or maybe these models kind of understand right and wrong. And if you enhance the wrong, then all sorts of other behavior is interlinked and expressed. But I don't know the way that I think about it. These models are just like the functions of their training data. And these models are trained on everything, like all sorts of misbehavior as well. And you want it trained on like, um, or you incorrect behavior as well because like, otherwise it won't know to refuse harmful requests or to not do a certain set of things. Do you have an intuition for why different models have different base models have different personalities?
16:43Like, for example, the newest clawed series, like, I think one of the models may be Opus or so on it is, you know, really cares about animal welfare, for example. And the other don't like give a sense for why these models develop pretty distinct personalities. I think it's just a function of how they're trained and it's really, really hard to anticipate in advance. I don't know. I feel like it might just be me, but like, Cloud4Opis too is enormously sycophantic. I'll kind of nudge it in one direction and it'll just agree with me wholeheartedly. I'll nudge it in the other direction. Pose a counter example, and it would just be like, yes, I was totally wrong before.
17:24Like, nothing I said earlier was correct. And I just think it's like really, really hard. It's like, it kind of goes back to like the witchcraft almost of training an AI model today where you're just like throwing in training data into the model and like whispering the incantation of gradient descent and then like trying to like get what you want out of it and something pops out. And then it like really cares about animals. Oh, great. Like, that's, that's great. I'd love to talk about your research results so far, you know, both at Goodfire and in the brother Meckenturp field as a whole. Maybe could you just give us a 30 ,000 foot flyover view of, you know, Meckenturp as a field, how old is it?
18:06What are the key results so far? What are the big open questions? Yeah. So Meckenturp as a field, I think like maybe just like in this tradition that we're building on. I think there are all sorts of studies like looking inside neural networks all the way back when we first designed neural networks. But I think once we, I think the way that the field kind of thinks about itself, like mechanistic interpretability was started out like OpenAI with Chris Ola and Nick Iberrata and a couple other folks who first put out like this really big circuit thread that posited like three things. One is like our features in neural networks, which are directions of latent space that represent concepts that the model uses to generate its response.
18:51Circuits, these are like features that fire together to create like higher order concepts. The example that they lay out is like you have a car window detector and then a car like body detector and then a car wheel detector and then that's like a car circuit and universalities the third -ten, like similar circuits evolve in different neural networks. And so this was like almost like the start of the field of mechanistic interpretability in my mind. And so that really kicked off a lot of, you know, interesting research and results in like the features circuits paradigm and like the main players in the field, there's a lot of academic labs that are doing great work like in Thropic.
19:40So Chris Olas, one of the co -founders of Enthropic and building a great interpretability lab there, and D -Mine has interpretability lab as well, and then we are kind of like the newer entrance on the field and on the stage. And I think one of the other really key things to have happened was understanding and mostly resolving superposition. So, superposition is this idea that like each neuron is responsible for encoding like multiple concepts. And there are more like concepts than dimensions in a neural network. So, if you think about a neural network as like a giant compression algorithm, you're like compressing like the entirety of the internet into like a relatively small number of parameters.
20:33And so that means every single neuron needs to encode or at least every single layer of the model needs to encode more concepts than it has like dimensions. And so there's like this concept of superposition where you have concepts represented as like near orthogonal directions in latent space such that you can represent all of these concepts in a model's latent space. And so to resolve this, you have to almost untangle an unscramble neuron such that it's responsible for like one clean interpretable concept. And so it's like a group at Apollo research led by like Lee Sharkey, who's actually now at good fire.
21:16First kind of pie in your like Sparsado encoders for language models and and Thropic also like really popularize this with their like big paper like towards monosamanticity. and then right afterwards, like scaling monosamanticity, showing that you can essentially unscramble these neurons into higher order concepts reliably and at scale with like arbitrarily large neural networks. And I think that was a really big moment for interpretability where you can now, in a totally unsupervised way, unscramble neurons of a neural network to understand them and get clean concepts. So the concepts aren't like totally clean yet.
21:56You can't like edit them super well. There's all sorts of problems with this, but like it's almost like a really big step forward for the field such that like that we can do this in an unsupervised way and the techniques and interpretability scale, which is really important. Does that mean the superposition isn't real or that you must like to hide some bugs on certain different, so you sort of collapse it at a particular moment in time to know that in this instance, it represents a particular direction. So I think it means it was real. Like neurons are responsible for, you know, encoding multiple concepts such that once you unscramble them, then you can do really interesting things with like a clean neuron.
22:36So the way that we do this is like using an interpreter model, train on the activations of a base model, and then now you have all sorts of neurons in the interpreter model that represent like theoretically clean sparse concepts. In the interpreter model, not in the original model. The original model still has this characteristic of superposition. That's right. OK, go ahead and thank you. And in the interpreter model, you unscramble these neurons and you associate these concepts with the concepts in the base model that you're trying to interpret. And then you can do interesting things with that.
23:06Go ahead and thank you. Yeah, of course. How solved of a problem is this then? If you've already been able to kind of disentangle the super position, then haven't you already mapped out the mind, so to speak, of the neural net, and what's ahead? I think partially, I think it's a partial mapping, and the technique has all sorts of flaws as well that we can improve a pot. But I think like it gives us the first step towards understanding these models, especially going from toy model to like, actual network that people care about. So we've done a bunch of work recently on R1 and so it's just 671 billion parameter mixture of experts models.
23:53It's a big boy model. And like the technique scale really nicely all the way up to that point because you know it's just more AI, more training of an interpreter model. So obviously I presume there's an asymptote here to understanding because the models are going to get more and more complex over time. We're going to beef them up. So I'm guessing sort of like the Battle of Sisyphus at some point, this is never ending pursuit. Is that correct? Do you agree with that? In some ways, but I also think that like, so the techniques that we've developed work on, toy models all the way up to like, yeah, big network that's like more capable and more intelligent and better.
24:43And I think like the techniques also scale effectively with model intelligence. So one part of our pipeline is that for every single latent concept in our interpreter model, we associate that with and try to, we get another language model to reason about like what that that concepts actually represents in the base model. This is a concept called auto interpretability, which Nick Camarata, who's a good fire now, and invented this technique at OpenAI,
25:14pioneered, and this technique, because it's a language model reasoning about, what a neuron represents, scales with the quality of the language model. So it actually gets better. So because we use AI in order to understand AI, like the better that our models are, are like these analysis agents are at interpreting like what's actually going on, the better we are able to understand them. And our interpreter model techniques also, it's just like we develop better interpreter models. They theoretically should translate to more and more intelligent and larger and larger networks because these are like unsupervised scalable techniques and that's like the paradigm in AI interpretability.
25:56All right. When do you think you reach a minimum threshold that makes you feel it's ready for a real world application. Maybe they already. I think we're there. Yeah, I think the first real world applications are already out there. And yeah, I think we're there on the very early applications. Can you share more about this? Yeah, I was being unnecessarily cryptic. So yeah, a couple of the partnerships I'm most excited about. So we work with our Institute, like I mentioned a little bit earlier, to understand an interpreter, Evo2, which is their DNA foundation model. So it's a sequence -to -sequence model, so it takes in a sequence of nucleotides and it predicts the next nucleotide in a sequence.
26:48And our theory is this is a narrowly superhuman model. So we really like to work on narrowly superhuman model, because it can teach us something about the world that humans don't really know. And so the idea is this model is representing just an enormous amount about the biological world in order to generate the next and properly model the next nucleotide in a sequence. So what we did was we sought to understand what does it actually know such that it can model the world so effectively. So what we did was we trained as far as auto encoders on the activations of this model. Extracted all sorts of features that were related to concepts that the model like should know, like kind of normal biological concepts that we have like really strong ground truth annotations for.
27:41So these are like TRNAs, RNAs, start coding sequences, all sorts of biological concepts that we have ground truth annotations are. We associated with this model. And then the question is, OK, now we have all of these other features of the model that we've extracted. What do they mean? What are they? They might just be ways that the model is computing and thinking or they could represent novel biological concepts that the model is using to generate the next, you know, nucleotide in a sequence. That's very interesting. I mean, for a long time, there was this idea that we have a bunch of junk DNA.
Read the full transcript
28:19I may have read about this and turns out a lot of that DNA actually serves a particular purpose in a different part of evolution or that they govern the expression of other genes. And so, you know, nature generally doesn't want to harbor things that don't have value because it's expensive, you know, just from a biological system point of view. So that's super interesting. I'm looking forward to the results. Yeah, totally, totally. And I think hopefully using unsupervised AI techniques, we can better understand what all these portions of the DNA are actually doing. Maybe we can discover the idea of junk DNA faster or understand that DNA is not junk DNA faster.
29:01And or just discover totally novel things that like genes are doing and expressing within us. Yeah. Where is the research as far as going from understanding and mapping towards editing? So for example, being able to reach in and you know, change this weight from here to there. I'm curious if you all have any results there yet. Yeah. So we've done most of our editing work on like language models and image models. Like our most recent kind of release was a paint with ember. Embers are like kind of foundational infrastructure for interpretability. And what we were able to do with this like image model demo was targeted precise control over an image model by painting.
29:50So we could extract latent concepts like a dragon or dragon wings or an ocean or a purepid, and then take these concepts and directly intervene on the portion of a canvas that we want to intervene on. So you can paint on a dragon with wings and then add a crowd in the corner and add a pyramid. And it's a really fun demo that's just a joy to play with. So it's out right now, anybody can play with it. It's just paint .goodfire .a. Yeah, but we're able to, I think, reasonably intervene in certain situations on a model's latent and steer the model to do what we want. But we haven't quite cracked the idea of direct precision surgical edits that create a new model that you want to use and doesn't have any unintended side effects.
30:47So that's like still, you know, something that we're pushing on and trying to figure out. Do you think that's where the field ultimately goes? Or do you think people are focused on different parts of the field? I think there are many places where the field is going to go and this is one of them. Like, interpretability is almost like such a general term that, again, I'm biased, is, but I think it's just like governing and underlying like all aspects of AI. It's like, any time you prefer to take a white box approach to doing something versus a black box approach, like interpretability can probably help in the future.
31:27So how do you select your training data? Maybe you want to understand like whether the training data is surprising to the model before like putting it into the model because then it can have the most impact on training. Yeah, just like in every single part of the AI development stack, I think like interpretability will help and change the way that we do things. If AI -fundational models go the way that a lot of software has gone, certainly infrastructure software, where much of it is open source or open weight, is there an opportunity for you to play an invaluable role in judging the biases or likely outcomes of using different open weight models.
32:09I think we could. Yeah. So there's maybe like two two areas of research that we're interested in that intersect with this idea. So auditing. So like how do you take a model, understand like what's going on, find like problematic behavior and good behavior. Hopefully get rid of the bad behavior and enhance the good behavior. So I think as AI gets deployed in more and more mission critical contexts, that becomes more important. And then also model dipping. So it's like when you have two checkpoints of a model, how do they differ from each other and what's changed. So the recent GPD40 was enormously sick of Antic for a period of time, just telling just really gas enough the user.
32:57Tell us that they're doing great. You still is. Pat recently asked it who the most handsome cribble board member was. And it was like definitely, definitely Pat greedy. Definitely. Still a bit sick of Antic. That's so good. That's so good. Yeah. But yeah, like model dipping, like you should be able to detect like how a model has changed from checkpoint to checkpoint. Like what surprising things have happened that were unintended that had now contained the network that weren't there before. Why do you think it was so hard for OpenAI to roll back to a less sycophantic version of the model? And the ideal state of the world is they're almost a dial and a knob that the OpenAI guys could tune on a scale of the 100 how sycophantic do you want the model to be?
33:42I don't know what questions you're asking of the model by the way because I never encountered this particular problem. And then sometimes it's brutal the other way. I said the best AI podcast and it lists 20 things with no training data. But about us, it's all I didn't want to, but I didn't want to give a biased result. That's so funny. Well, that's part of what users want, right? They kind of want sick of fancy. Like, you know, people want to hear what they want to hear. So when you RL a model, I think fundamentally, you're going to get, I think it's just kind of a symptom of RL. That's like, you, like, this is what users want.
34:23this is user preferences. Along the way you've dropped some names and it seems as though most of them have ended up at good fire. I presume there is a certain number of very talented people in this field and you've unfairly seem to gather them. Can you describe a little bit more about your team or what you've pulled together? Yeah. So I mean, I think we have a really fantastic team and that's what we've been spending a lot of our time on the last year. Just kind of like I think assembling a team of world -class interpretability experts that really have a shot of cracking this problem. So it starts with my co -founders.
35:01I had worked with Dan Balsam, RCTO for many years at my previous company. He's RCTO and Archive Scientist Tom, who founded the interpretability team at Google DeepMind, way back in the day, and just have assembled many of the early folks in the field. So Tom, Nick Kamerata, who's working very closely with Chris Ola, who is generally considered like the founder of the field of Mechintarp. And Nick was on all of the original like circuits papers and helped build everything out at OpenAI. Lee Sharky, who's the first who pioneered like Sparsado encoders on language models, is is now working on some really interesting work in weight -spaced interpretability.
35:51So most interpretability techniques that have been deployed into applications are in concept space and activation space and he and his group are working on weight -space interpretability techniques. And we've also just kind of pulled in scientists, senior scientists from other fields who care a lot about interpretability and just kind of have realized that this is one of the most important problems that we can work on. And so Owen Lewis, who was a senior staff RS at Google, working on coding agents, came over and is now like leading a couple directions here for us. And you're recruiting, right?
36:28And we're hiring, yeah. Scientists, engineers. I think it's like, we are hiring scientists and that is like deeply important for the future of the field. But also like, it's hard to just like, it's hard to overestimate just how important and like good engineering skills are. Incredible team. Proud of the team, yeah, for sure. This seems like core functionality for any of these foundations and model companies to have. And as you mentioned, you know, Krasola, was it OpenAI, now then Thropic? OpenAI has their interpretability team as well. How do you think about the rationale for having a standalone Mechinterp research company versus being inside one of the labs that should care deeply about this?
37:11I think we can just take a really different approach if we're independent. I think the benefit of being independent is we can think independently, push things forward independently, and also get a broader view of the ecosystem. So usually if you're within a lab, you're doing interpretability work on your own models and pushing forward the field in that way. You can make incredible progress that way, but I really do think that a unique third -party perspective is deeply necessary in the field. And I think like, yeah, just given the team that we've assembled, like a lot of those folks like agree with that.
37:50And that's why they've joined. And also, like, gives us like an ability to work with lots of interesting partners across different domains. And we can kind of unify those insights across all of these different domains that teach us more about like the inner workings of neural networks, more broadly. We work across modalities like genomics models, exome models, image, video, like language and also across model architectures. And I think all of that just helps. And throughout the invested in you all, right? Yeah, that's right. Say more about that and how you partner with them? Yeah, so they, I think we were there their first ever investment.
38:33They put in a check in in our last round. And I think they just got, they just really care about interpretability and really kind of see the future as we do, where interpretability is just like pretty critical to the future. So Dario just published an essay called like the urgency of interpretability. And it's like one of his like four essays that he has on his site. And just like talking about like how he views this is almost like a race. And we see that very similarly, a race to get interpretability prior to you know super intelligent, really, really intelligent AI models. I just think it's like deeply critical to be able to understand these models before they're before we have like in his words like a country of geniuses and in a data center.
39:25Do you think interpretability can help us with with open models. And I think some people have a fear that models trained in other countries that may or may not be enemies of the United States, have different nationalist properties. Can interpretability help us understand and even modify those for the American variance of some of these models? Yeah, well, I definitely think so. I also think it's relatively easy to like if you take like a deep seek model, for example. It's relatively easy to just tune it or add in more training data to remove a lot of the propaganda in the model. But yeah, I think like interpretability can help understand what's actually inside of the model and then also change it and edit it to serve whatever and purpose that you want.
40:17A long do you think before you're going to be pulled in as a witness in a very important trial, to try to understand why model did something in particular. That's a good question. I think,
40:32a few years. Who knows? I think it's really, you know, we're all sitting in like the Bay Area right now, but at this point, I'm pretty agi -pilled in that. Like, I think AI progress will be pretty fast and pretty quick and transformative to society in ways that are really difficult to anticipate from from over sitting right now. And so, yeah, I do think that there will be a couple, you know, like big failure cases of AI models and whether it's me called it, or, you know, if somebody had a big able to explain a model's outputs. Yeah, I agree with you on the rate of development, by the way. I think it's, um, I'm sure you've read these articles that human brain doesn't intuit compounding.
41:32No. And so, you know, I've even thought back to 20 years ago, uh, when I first met the self -driving car initiatives and Sebastian Thrundt's team from Stanford at one, the DARPA challenge, you know, you could sort of see the glimmers of self -driving cars, but even then, if you'd say, you know, 20 years later, you would have a self -driving car in and San Francisco take you around, I'm not sure there'd been obvious that would be true. And maybe take a few, took a few as longer for the true visionaries. I think the same is gonna happen with AI. I don't think we fully fathom what the world's gonna look like in 2030 or 2035.
42:02Yeah, I can agree more. And it's just really hard to predict, you know, like even if we feel like it's going to happen really, really quickly, it's hard to predict all of the ways the society will be transformed. Good note to end on. Should we do some rapid fire? some predictions. Yeah, some predictions is all recorded. So we'll hold you to it. Great. Yes. Yes. Yes. Maybe Eric in 2035 will look back on how wrong he is with all these predictions. Okay. Maybe first inference time compute is the next important vector to scale, to scale models. I agree. I disagree. Uh, mostly agree. I think it's one of the important, one of the things that we can scale up on.
42:43Yeah. With application category, do you think we'll break out next after code. I think there will be a lot of enterprise transformations that happen. So just like automating like manual routine tasks that people are doing like many, many times a day. Employment impact from AI. Vast and vast, but once you cross a chasm, I think it happens quickly. I think my last company was helping find early career jobs for people and using AI to automate that. And I think that that's where we're going to feel the impact first. I agree. Recommended piece of content or reading for AI fans, maybe specifically in your field.
43:37I think the original circuits thread that I was referencing a couple times, like that's still fantastic. check what's either an AI app or maybe just an experience that you've had with AI that has blown you away recently something and just took a breath away. I think my like, like one of those moments that you really just like feel how fast AI is happening like when I first played with O1 Pro, like that was a model that I just really felt like was actually reasoning about the world and seeing like the kind of cross -domain transfer to out ask it a strategic question and it felt like it would actually understand like all the levers that I was considering with the business and considering that like at least relatively thoughtfully and being a thought partner.
44:26And that's both exciting because I now have this model that I can talk to you about all sorts of critical problems. but not just, of course, trust it blindly, but also, it's like, wow, how did this happen? One of the things I've learned recently is that AI is still struggling to understand humor. One of my partners, Andrew, actually had this joke that humor is humor's way of showing off intelligence without actually explicitly bragging. There is a lot of embedded intelligence in humor. or do you think interpretability will help us pinpoint to figure out a humor? Yeah, to figure out why eyes don't have a sense of humor.
45:09We'll help them to develop it. Perhaps, who knows? Yeah, I hope so. Yeah, I can, I think if I, you know, wake up to a model like telling me jokes and real off -s voice, that would be a great thing. That would be terrifying. Okay. Well, close this one last question on a prediction in your field. Do you think we will ever reach the point where we feel like we confidently understand the features, the circuits, the patterns, the weights of a neural net. And if so, what year do you predict will reach that point? I think we can. I think that it might not look like what you just said, like the features, the circuits, but I think it requires maybe a reconceptualization of what's actually going on inside the model, like a deeper, more fundamental understanding of the units of computation of a model.
46:02It's almost like discovering truths about the universe or about like neural nuts in this case. But yes, I think we're on track. I think we can do this. And I think and hold me to this in 2028. We're going to figure it all out. Yeah. Fantastic. Yeah, just a few years, I think we're close. Just in time for the Olympics. Yeah, that's right. Just in time for your next round of funding. Eric, thank you so much for doing this today. Rulaf and I love the conversation. Thank you. It was a pleasure. Yeah, it's so much fun. Thanks for having me.
From the publisher
Eric Ho is building Goodfire to solve one of AI’s most critical challenges: understanding what’s actually happening inside neural networks. His team is developing techniques to understand, audit and edit neural networks at the feature level. Eric discusses breakthrough results in resolving superposition through sparse autoencoders, successful model editing demonstrations and real-world applications in genomics with Arc Institute's DNA foundation models. He argues that interpretability will be critical as AI systems become more powerful and take on mission-critical roles in society.
Hosted by Sonya Huang and Roelof Botha, Sequoia Capital
Mentioned in this episode:
Mech interp: Mechanistic interpretability, list of important papers here
Phineas Gage: 19th century railway engineer who lost most of his brain’s left frontal lobe in an accident. Became a famous case study in neuroscience.
Human Genome Project: Effort from 1990-2003 to generate the first sequence of the human genome which accelerated the study of human biology
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
Zoom In: An Introduction to Circuits: First important mechanistic interpretability paper from OpenAI in 2020
Superposition: Concept from physics applied to interpretability that allows neural networks to simulate larger networks (e.g. more concepts than neurons)
Apollo Research: AI safety company that designs AI model evaluations and conducts interpretability research
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. 2023 Anthropic paper that uses a sparse autoencoder to extract interpretable features; followed by Scaling Monosemanticity
Under the Hood of a Reasoning Model: 2025 Goodfire paper that interprets DeepSeek’s reasoning model R1
Auto-interpretability: The ability to use LLMs to automatically write explanations for the behavior of neurons in LLMs
Interpreting Evo 2: Arc Institute's Next-Generation Genomic Foundation Model. (see episode with Arc co-founder Patrick Hsu)
Paint with Ember: Canvas interface from Goodfire that lets you steer an LLM’s visual output in real time (paper here)
Model diffing: Interpreting how a model differs from checkpoint to checkpoint during finetuning
Feature steering: The ability to change the style of LLM output by up or down weighting features (e.g. talking like a pirate vs factual information about the Andromeda Galaxy)
Weight based interpretability: Method for directly decomposing neural network parameters into mechanistic components, instead of using features
The Urgency of Interpretability: Essay by Anthropic founder Dario Amodei
On the Biology of a Large Language Model: Goodfire collaboration with Anthropic




