Inside the Hidden Geometry of AI

15 Jul 2026 · 58 min · 18 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Goodfire’s AI interpretability research, arguing that neural networks “think in shapes” via curved, high-dimensional manifolds (“neural geometry”) rather than only linear directions/words. They describe extracting internal concepts unsupervised (features, manifolds, circuits) and using them for debugging, safer training, and reducing hallucinations through “Reinforcement Learning with Feature Rewards.” They also outline “steering” (dialing internal concepts up/down) and “predictive data debugging” to filter/shape training data before training.

Guest backgrounds

Eric Ho, co-founder and CEO of Goodfire; leads interpretability research. Hosts: Corey Knowles and Grant Harvey (Neuron AI Explained).

Key claims

Internal representations are structured and labelable; higher-order curved geometry matters more than linear feature directions; extracted “confidence”/hallucination mechanisms can be rewarded or steered; interpretability enables “intentional design” during training.

Notable examples

“Rabbit ear” represented as a textured gradient across bottom/middle/top (not a single linear “rabbit ear” feature); “mountain car” steering works smoothly only with a spiral/helical latent structure; days-of-week visualizations; “Trump neuron” shared across image and language.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Neural Networks

0:00 to 0:42

Explore the internal workings of neural networks and their structure.

“Goodfire is an AI interpretability research company.”

Meet Goodfire and Its Mission

1:01 to 2:21

Learn about Goodfire's approach to AI interpretability and its goals.

“I'm Corey Knowles, joined as always by the man bringing charm, curiosity, and just enough chaos to keep legal concerned, Grant Harvey.”

Deep Dive into Neural Networks

2:27 to 4:25

Discussion on the techniques used to understand and debug neural networks.

“Real quick, before we get started, make sure you hit the subscribe button above.”

Extracting Concepts and Structures

4:25 to 7:40

Learn how to discover and utilize concepts within neural networks.

“And so, yeah, what do you think a lot about that?”

Philosophical Insights on AI

7:45 to 14:00

Exploration of the implications of AI's complex inner workings.

“Imagine like the simplest unit, the simplest unit of computation of a model is one single neuron.”

Exploring AI's Inner Workings

14:00 to 20:40

Discuss the complex structures and computations within AI models.

“And so what is contained in that enormously rich inner world?”

Neural Geometry and Model Understanding

22:15 to 28:00

Delve into neural geometry and how AI models represent concepts.

“We have a bunch of new research planning all the time.”

Understanding Latent Space in AI Models

28:00 to 29:40

Explore how AI models navigate complex structures in latent space.

“There's like a just like kind of higher order structure that governs kind of model cognition.”

Shape Rotators vs. Word Cells

29:40 to 32:20

Discuss the ongoing debate on mental representation styles in humans and AI.

“For me, when I start to draw something, I don't always have the exact image of it in my mind.”

The Role of AI in Model Training

32:20 to 35:40

Learn about advanced techniques for training AI models effectively.

“I mean, I know some people are already trying to do that, but.”
Show all 18 chapters

Silico: An AI Neuroscientist

35:40 to 38:00

Discover how Silico aids researchers in understanding and training AI models.

“And it's remarkable how far it's gotten us.”

Reducing Hallucinations in LLMs

38:00 to 40:20

Examine methods to minimize hallucinations in large language models.

“And it also comes with training libraries, like our SFT and RL libraries that come bundled with compute.”

Steering AI Behavior

40:20 to 42:00

Learn about techniques to control and adjust AI model behavior effectively.

“So there are all types of things that you can do here.”

Unlocking Continual Learning in AI

42:00 to 46:08

Explore how continual learning can enhance AI model training and safety.

“I know there's a term that people have talked about this.”

Surprises and the Holy Grail of AI Research

46:08 to 48:24

Discuss the unexpected findings in AI research and the importance of interpretability.

“We started the company a little over two years ago.”

Designing AI Models with Intention

48:24 to 50:17

Learn about the importance of intentional design in AI model development.

“Well, I know one of the things that you mentioned earlier is that, you know, you want to be able to build almost the equivalent of custom models that you have engineered like software from the beginning.”

Challenges of Robotics and Multimodal Models

50:17 to 56:01

Understand the complexities in robotics and multimodal models in AI.

“I have one question I want to get in before we go, and I'm sure Grant's got three.”

Exploring Goodfire and Silico

56:01 to 56:56

Learn about Goodfire's offerings and how to access Silico.

“And yeah, just reach out to us on Twitter too.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:01Goodfire is an AI interpretability research company. So we spend all day thinking about what's going on inside the mind of a neural network. We had a nice paper a few months ago called Reinforcement Learning with Feature Rewards, where we were able to use the internal concepts of the model that we were able to extract in order to reward the model to reduce hallucinations. It's the idea that models really think in shapes and not really in words. It's like in curved, high-dimensional structure. And we've basically developed a technique to extract these structures in an unsupervised way. You can use this to help debug and train your model to be more effective and safe and performant.

0:49I think this has been almost like the story with interpretability. It's like everywhere we look, there's structure, there's really interesting concepts. Welcome, humans, to the Neuron AI Explained. I'm Corey Knowles, joined as always by the man bringing charm, curiosity, and just enough chaos to keep legal concerned, Grant Harvey.

1:11Corey Noles:Gotta keep him on their toes, you know? Right? Right? Yeah. Well, I'm excited today, Corey, because today we're going inside the mind of a neural network and not in a vibes-based sci-fi kind of way, as you might expect. Or a spaceship way. Well, Corey, you know, you have a very peculiar set of prompts that you like to ask models on release day. So people who don't know what I'm talking about, check our live videos. But not in a vibes way, not in a Corey way, but we're actually going to go inside the mind of a model, or at least as much as we can. And luckily, today, we are joined by Goodfire, who are trying to understand, debug, and design AI systems by doing exactly that, looking directly at the internal structures of neural networks and trying to figure out what we can understand from them.

1:59Corey Noles:The company's core idea is that models contain meaningful mathematical structure, that's features, manifolds, representations, and circuits. And that if we can map those structures, we can make models safer, more reliable, and more useful. Thankfully, we have Eric Ho, who is the co-founder and CEO of Goodfire, to join us today to explain how all this works. So I think with that, we will push it back to you, Corey, real fast. Today's video is brought to you by SaaS, AI governance software you can trust. We'll hear more from them in just a bit. Real quick, before we get started, make sure you hit the subscribe button above.

2:31And on that note, Eric, welcome to the Neuron. Thanks so much for having me. We think a lot about neurons, so it's a fitting name for the show. I'll bet you do. I'll bet you do. I guess before we get started here, for listeners who maybe don't spend their weekends spelunking inside neural networks, what is Goodfire and what problem are you trying to solve? Yeah, Goodfire is an AI interpretability research company. So we spend all day thinking about what's going on inside the mind of a neural network. The reason why we care about doing this is because we really want to build towards a future of what we call intentional design, where you can actually understand and debug and edit and really shape a neural network during training so that it becomes a lot less like trial and error, just kind of throw in data and get, you know, model out.

3:25And a lot more like building software, where you actually understand, like, how your inputs relate to your outputs, how data leads to generalizable representations in your model. and we want to kind of make this entire process transparent, steerable, debuggable, so that we can get better and safer neural networks that we can actually design with intention. So that's the whole goal of the company. We're all thinking about this every single day. And, you know, fundamentally, we're a research lab that has to, you know, figure out a lot of the core problems and core research with neural networks so that we can get there.

4:03Corey Noles:Right, right. That makes sense. You're basically trying to break open the black box. 100%. Yeah. During training, looking at a model to understand what type of creature is this. It's not really just a computer program anymore. There's clearly something interesting going on here in ways that are analogous to brains in some way. And so, yeah, what do you think a lot about that? So how do you actually do that? There are a variety of techniques, and they are still kind of developing as we speak. But maybe at a very high level, we use AI to do this. So we train all types of interpreter models so that we can understand the concepts, the parameters, the neurons, the weights of these models.

5:02And there are all types of tools that we leverage in order to do this. But they all are kind of permutations on this core idea that inside these models, we don't quite understand how all of the concepts relate to each other, how they form their cognition, how they manipulate concepts internally. and we have to kind of almost like break it open bit by bit in order to extract meaningful units of computation that we can then compose to higher order structure that then tell us something about how the model is performing this computation. So maybe for example, we can start by talking a little bit about just like there are like complex structure in these models like we found addition mechanisms for models to perform like complex computation internally you can extract all types of concepts like languages style how they represent the user what they're thinking about it in any given time and in all types of models as well.

6:13So for like biology models, they represent things like the tree of life or ancestry or all sorts of other concepts. So all of these structures, these concepts are discoverable in these models. And we spend a lot of time thinking about how to both extract these concepts as well as how to use these concepts during training for debugging for downstream applications. SAS has spent 50 years helping organizations make decisions they can trust. And as AI evolves from predictive models to large language models and now autonomous agents, that challenge is becoming even more important. We're starting to ask questions like, if I'm deploying AI that can make decisions, take actions, and interact with other systems, how do I govern it?

7:01SAS is built for this era of autonomous AI. Your AI governance approach should provide oversight of predictive models, LLMs, and increasingly autonomous agents. The list goes on. You need to monitor outputs for risk, track automated decisions, manage third-party models, and apply consistent controls across your AI. SaaS focuses on using data and AI responsibly and effectively, giving organizations visibility into how AI systems behave, accountability for the decisions they make, and confidence that they align with business objectives and support regulatory readiness. Because as AI systems become more capable, trust becomes more valuable.

7:41Learn more at sas.com. That's S-A-S dot com. Now back to our show.

7:46Corey Noles:When you're using AI to do this at a high level, are you essentially running the same requests in one model and then having another model sort of interpret those and say like, oh, when this was the output, these, you know, I don't know how exactly it looks like, but I assume there's like some sort of lights based on the visualizations that you've released, that there's some sort of lights that you see that light up, numbers light up again and again, and then you can actually tell, oh, that's how, that's what this represents. Is that kind of how you're doing it? That's exactly right. Yeah. Cool.

8:20Imagine like the simplest unit, the simplest unit of computation of a model is one single neuron. And these models are like trillion parameter models. So it's like a very small slice of the computation of a model. But if you look at a neuron, like our types of, our techniques essentially look at, you know, when these neuron activate and fire and then aggregate them across like, you know, potentially similar rollouts, similar tokens, and then train a type of like interpreter model to understand what that neuron is actually doing across like a set of computations. Neurons then compose into higher order structures.

9:00So we call this approach and direction or neural geometry direction, where there are curved geometric structures in the latent space of the model. That's kind of like the mind of the model where it, you know, manipulates vectors and computations. And so these like kind of higher order neural geometry structures are also things that we can extract and understand. And so going back to the calculator or the tree of life, these are like higher order structures that are made of many concepts that compose into a more complex concept. And these are also extracted by paying attention to the activations, the weights of the model as it's kind of going through its computation.

9:55Okay. So are you saying like that these various things live in specific regions of a neural network in many cases? or like you know where to go to look for some specific kind of reasoning or spatial awareness or whatever the case might be.

10:18Corey Noles:Spatial awareness was a particularly terrible example I'd like to add. Well, for example, one of the papers that you released, right, you showed all of the different visualizations of the neural geometry. And in some of them, you could even see how it represents the days of the week through how it moves across the pattern. So maybe you could speak to that, you know, because there's like the regions, there's movements, and then there's also even manifolds inside the latent space. Totally, yeah. So maybe at a very high level, what we're saying about a model is that there are units of computation inside the model that we can understand and label with semantics.

10:59These units of computation composed together to form higher level concepts and structures in the model that we can also extract in fully unsupervised ways. So we can kind of discover these concepts rather than have some sort of prior when looking inside the model. And this is interesting because you can then use these structures to debug the model. Let's say that the model shouldn't have been thinking about race as it's developing a medical diagnostic or something like that or a hiring decision. or, you know, in like the model is developing the ability to do complex arithmetic. And that would be a good thing to reward and incentivize as the model is training.

11:59And so that's maybe like the higher level mental model that we have of what these models are doing. They're actually highly interpretable the more that we look. So this is giving you the ability to not to, for lack of a better phrase, sort of reinforcement learning, a reinforcement learning approach that can go a layer deeper than what we've previously been able to based on what we've been able to see. Is that correct? Like we've stepped another layer in inside, if that's it. Yeah, we had a nice paper a few months ago called Reinforcement Learning with Feature Rewards, where we were able to use the internal concepts of the model that we were able to extract in order to reward the model to reduce hallucinations.

12:49But this can be done with arbitrary features or concepts that we can extract from the model. So we did this by essentially figuring out the types of hallucination mechanisms that were present in this particular model, which was a version of Gemma. and then we were able to kind of directly train and reward against that mechanism or that set, probably more accurately, the set of mechanisms that we were able to extract that were responsible for the model hallucinating. Because often there's an internal state of the model where it knows that it's more uncertain. It knows that it's about to make something up and that's kind of what we're relying on.

13:32The model often has a much more complex and rich representation of the outside world than it lets on in token space. So think about it this way. It's like what we see is just a single token output when we're chatting with a model. It's just the tip of the iceberg. In order to generate that token, there's an incredibly complex and rich inner world that the model has to generate for every single token prediction in order to generate that token. And so what is contained in that enormously rich inner world? Well, there's all types of structure, all types of knowledge, all types of computation that is both incredibly fascinating but also incredibly useful for researchers to understand.

14:21I'm going to go a little philosophical here for a second. Let's do it. From a philosophical standpoint, what do you make of that? Like this idea that like there's more to this than what we're seeing. This is stuff that isn't intentionally programmed. This is stuff that isn't like built from the inside out in that way. Yet it is there and it is doing things and it is doing things that further its cause. It is doing things that meet the needs and things that don't, which is even more crazy and interesting, by the way. Any theories? What do you think is going on in there? that is a big question uh there's clearly something quite interesting going on in here though where the level of computation the richness of understanding of these models is is clearly um on uh on a level that you know mimics intelligence in some way shape or form i guess the whole it's ai it's artificial intelligence we're simulating intelligence and we're doing a pretty darn good job of it.

15:30And so kind of like some way that I think about what we do is like we have these enormously intelligent systems. What we're really doing is like trying to figure out like how they work and what they're doing and what is this like complex system actually doing. And I don't know, there are, I hope that one day these systems can teach us about our own minds and our own conscious experience. and what all of this really even is. What are we doing right now? At this current moment. We're just molecules and atoms. And yeah, it turns out that being able to train and generate an intelligence is not actually super hard.

16:18You can do it with just a lot of GPUs and a lot of training data. and that's remarkable and i really want to get to the bottom of it and figure out how all these concepts compose and uh form this this higher order cognition we do have uh consciousness researchers kind of bashing on our our product right now um trying oh wait why why what's their criticism oh no no no i'm sorry they're using our product to oh i thought you meant they were bash bashing. It's good bashing, Grant. You're old. There's a guy older than you. Different usage of the word bash. Bashing in a good way. Just kind of using our product to try to discover structure that relates to some type of experience that the models have.

17:12And the models what's interesting is that if you remove kind of the model's refusal behavior or hedging behavior. They will often discuss some type of subjective experience that they're engaging in. But of course, it's very, very difficult to tell because they're trained on human data and a lot of sci-fi.

17:39Corey Noles:Well, look at how long it took us in this podcast before we got to the concept of consciousness. 20 minutes, you know? Yeah. It's such a fascinating realm of this and what made you really interesting to me in thinking of this because I kind of have a soft spot for interpretability and emergent technologies and all of that and seeing what's happening. And it's pretty clear that there's not, you know, there are two schools of thought. One is it's ones and zeros. Shut up. The other one is it's absolutely conscious. And I think it's a little more gray than that. I think we maybe really don't know what this is going to be yet.

18:27Corey Noles:Well, I'm kind of subscribing to the belief at this point, not having enough priors to really know for sure. But I'm subscribing to the belief that perhaps intelligence is a consequence of complexity. Something about that just really hits with me. It makes a lot of sense if you look at it from a certain perspective. If a system is complex enough, it has to be intelligent to continue to function by some definitions. And I've seen papers that are discussing this. Yeah, perhaps. I think the way that I think about it is, you know, these models clearly can generalize and it's almost like compression that results in the almost forcing function that these neurons have to be doing something really interesting together that form like some type of generalizability or intelligence from that act of, you know, jamming a bunch of information into a relatively small number of parameters compared to the entirety of the internet.

19:29But yeah, maybe even just to the earlier point, I mean, clearly there's something going on. We have found analogous structures inside models to humans, like our perception systems, you know, curve detectors, things like that. But also that these models clearly aren't human in many, many ways as well. And so I think we're just dealing with some type of new creation, new creature in some way. Yeah, I think that's a really cool way. You know, I always joke that we shouldn't have latched on to artificial intelligence. I think we should have called it synthetic intelligence. Because I don't really know where it's going to land.

20:14But I mean, either it's intelligence or it's not. I mean, artificial is, I don't know, it just never made sense to me as a word to describe what we're working with. Yeah, it's either intelligent or not.

20:26Corey Noles:And then if it's intelligent, you know, how was it created? Was it biologically, you know, through evolution or through synthesis? Yeah. What a fun tangent. Right, right. So sorry, didn't mean to go down that hill at all, but there we were. That was not in our research, but it was fun stuff. That was absolutely not in our notes. Most people are only using about 5 % of what AI tools can actually do. They open ChatGPT, ask a few questions, maybe write an email, and that's it. Meanwhile, other people are out there building workflows, automations, research systems, custom agents, and saving hours every single week.

21:04Corey Noles:AI is moving so fast right now that most people don't even know which tools are worth learning anymore. ChatGPT, Claude, Gemini, Codex, AI agents, automations, 5Coding, it's a lot. And that's exactly why we built the Neuron Academy. We built a practical, self-paced AI learning platform that helps working professionals like you understand the tools that actually matter and how to use them in real work. Built by us, the team behind the Neuron newsletter and the podcast designed to help busy professionals actually level up their AI skills. Inside, you'll learn prompting, AI workflows, research systems, automations, productivity strategies, and practical use cases across the entire AI ecosystem.

21:45Everything is hands-on, beginner-friendly, and built around real-world use cases you can apply immediately.

21:52Corey Noles:These are real skills that help you work faster, think better, and stay ahead in AI. Right now, you can join the Neuron Academy for just$499 a year, unlocking 50-plus lessons from full AI courses to quick-hit micro-trainings you can apply immediately. So if you're serious about leveling up your AI skills this year, this is where you need to start. Join the Neuron Academy today. So I understand you have some new research coming. We do. We have a bunch of new research planning all the time. But there's one piece that we're particularly proud of that's coming out in just a few hours. It's part of our neural geometry series.

22:33So the idea that models really think in shapes and not really in words. It's like in curved, high-dimensional structure. and we've basically developed a technique to extract these structures in an unsupervised way so we can discover these in models rather than have to have some type of prior coming into the model to kind of extract them directly maybe one example that we use in the paper is a rabbit ear and whereas our previous techniques may have just extracted the concept of a near. It's like fluffy, it's near, you know, it's just one single, you know, concept in the model's mind. Actually, models represent this concept in a much more nuanced way than we had previously thought.

23:25It's a curved geometric representation of a rabbit ear, which means the model knows actually what the bottom of the ear is compared to the middle of the ear compared to the top of the year. The same goes for horizons. The same goes for just pretty much every visual concept. Models are just like representing something much more complex than you'd originally think.

23:47Corey Noles:Wow. What's going on there, you think? Is that just a... That's not a coincidence, right? I think a model just has to have a really good understanding of the world in order to do something like generate a high-quality image. It needs to actually have all of this in its latent representations to be able to do what it does. We haven't quite yet applied these techniques. We're working on it to language models, but we expect to find very, very similar kind of curved geometric representations in the latent space of these models as well that represent, yeah, like the manifold of computations of a model.

24:34So what you're saying is that the computations sort of form the same visual representation,

24:41Corey Noles:or at least sometimes the same visual representation as the actual object that they're going to, in this case, it's an image model we're talking about? So we are talking about image and video models. That's where we first pioneered this technique, which is called a block sparse featurizer. and it's for unsupervised manifold extraction from within these models and you can extract all types of concepts in this way but it's a new type of way to decompose the residual stream of a model to see what the model is representing at any given time. Got it. Got it. So if I'm understanding this correctly, it actually visually represents it internally before it then produces it as an image.

25:36so um not really uh i guess what i'm saying is i mean uh it has a representation of like something much more complex than just like a higher order linear concept so maybe backtracking for the field, there's been the dominant kind of view of the field of interpretability has been this concept of the linear representation hypothesis, where concepts are represented as linear directions in latent space. So just like single direction vector in latent space. What we're saying is that many of these concepts of the model live on a higher dimensional curved manifold in latent state which means our previous techniques to extract linear directions wouldn't work to extract a more complex higher dimensional structure but this is actually how models think and so it's really critical to meet the model where it is in order to uh extract these higher level structures.

26:52And so a rabbit ear, whereas it may have been extracted by like a sparse autoencoder is how you usually extract linear directions in latent space. It may have been just labeled as a rabbit ear and it just represents the entire ear. That's like the linear representation. For the neural geometry representation, the higher order representation, you would have a textured gradient of like this rabbit ear where it would have some delineation between the bottom of the year the middle of the year the top of the year and some type of gradient and that's important because you would then be able to kind of better understand like how the model is representing that that concept one example that we use is this like classic kind of mountain car steering example where you have the mountain car kind of going up a hill and the goal is to kind of um for the model to represent like the mountain car um kind of driving up the hill and so if you try to steer it linearly um which means like turn up that direction in latent space uh the mountain car will just be splattered across the entire mountain but if you have this like complex spiral helical structure uh that's actually how the model is thinking about you know moving the car on the mountain then the mountain will just drive up really really smooth or the the car will just drive up really really smoothly up the mountain and so kind of just how these models tend to represent things are in much like higher order structure than just a linear structure, which makes a decent amount of intuitive sense.

28:41There's like a just like kind of higher order structure that governs kind of model cognition. That makes sense. You know, I think when I think about what's going on in my mind when I draw a picture, we probably are thinking quite linearly as people, I assume. Or do you think that – and forgive me for comparing this to that. I just – I'm trying to think of it in a way –

Read the full transcript

29:09Corey Noles:You cut out a little bit, Corey. Did you repeat that again? Yes, sure. Sure. Kind of what I was wondering is – so this seems like this would be more technical than maybe what a human is thinking when they draw a picture. Perhaps. What they're seeing in their mind, at least, like, you know, I mean, if you're maybe just seeing the line you're working on or something instead of like visualizing the geometric, like, yeah, this is deep. Well, for me, let's run with this. For me, when I start to draw something, I don't always have the exact image of it in my mind. I'm not always the most visual thinker in that regard.

29:49Corey Noles:but when I have the feedback from the line and how I understand you know symbolically the picture that I'm trying to draw I am doing almost like a like a constant check between what I'm drawing and what I'm thinking in order to try and come back with something that it looks generally what like what I was going for yeah I'm not an artist you know I'm not good at it that that's how I think so. I don't know if that's similar to what the models do. Yeah, I mean, there's definitely analogies that we can draw, but what's funny is when we first released our neural geometry series, it really sparked a debate.

30:29The age-old debate between shape rotators and word cells. Everyone, all the... Can you explain that to people?

30:36Corey Noles:Yeah. Sure. The idea, this is just some dumb on Twitter thing, but all people are shape rotators or word selves. Do you mostly say words and that's how you go about the world? Or do you kind of rotate shapes? You're like an engineer or some type of mathematician, just rotating shapes in your head. You don't use words. And then basically what we showed is that these models typically think in shapes and rotate them in order to do things like addition. And so that was seen as a big boon for the shape rotation. As a validation for the shape rotation community? That's exactly right. And so, yeah, I don't know.

31:21I think these models are doing something in their latent geometries that is not directly just spitting out the next word in a sentence. They're not just stochastic parrots parroting back whatever is statistically most likely. There's clearly some computation going on that lives on this kind of manifold structure. And is there something similar going on in humans? I went introspecting into the best of my ability. Like there's something happening that's not in word space that is in, you know, much more of a conceptual space. But, you know, how a bore we to introspect? And is it necessarily a tangible thing?

32:14Yeah.

32:14Corey Noles:Yeah. Well, to your earlier point, at some point we should be able to run a language model over as many brain scans as we can create and, you know, be able to try and do some of the same work. No. I mean, I know some people are already trying to do that, but. 100%. I think there's groups at Meta and at Stanford that are doing this, this type of work that's really interesting and meaning to apply some of our techniques to then, again, reverse engineer their models to try to figure out what's actually happening in those models. And there's also a couple groups working on whole brain emulation, if you want to get even deeper here.

32:54Corey Noles:uh so that you know opens up a whole other other uh can of worms well in your opinion what is the implications of them thinking in 3d uh shapes or manifolds um besides the pretty visualizations which you've released amazing visualizations they're very cool they are very cool yeah our direction is intentional design so you can use this to uh to help debug and train your model to be more effective and safe and performant. So you can almost think of our product and what we're building as this AI neuroscientist that deeply understands your model, that can help you train your model, can help you debug why your model is performing poorly at this particular task.

33:43And if you can understand the core units of computation of your model, you can then manipulate those so that you can build better and safer models. So there are a few techniques that we use that rely on having really clean understanding of the representations of the model. The first is a technique called predictive data debugging. We can predict what your model will learn from a data set prior to training on it. So a really cheap pass over your data set to understand what your model will learn. And so you can then filter out the bad stuff in your data that you don't want your model to learn. Or you can kind of shape some of your data to make the model more performant and learn more from that data.

34:31Or you can go and purchase more of the good data from your favorite data vendor. So that's like more basically what you can do. You can also use this, as we talked about earlier, directly in the reward model. So you can shape your reward using features. And so if a feature is in a linear direction, it's actually a complex manifold. You should use the complex manifold structure, not the linear feature direction. And also you can directly intervene in the model as it's learning. So imagine the future that we imagine is that in every single train step of a model, there's almost like a little neuroscientist living in the weights of your model, telling it what it should learn and what it shouldn't learn.

35:22Imagine a much, much more effective way to get the model to learn what you want.

35:31Corey Noles:Because prior to that, by the way, it was like you just throw a bunch of data and compute at the thing and then you find out basically what happens. The old scoop up the internet idea. Yeah. Yeah. And it's remarkable how far it's gotten us. But is this the most effective way that is humanly possible to train models? Like the answer is, I mean, I find it hard to believe that anybody would claim yes. Okay. Yeah. Yeah. And they're putting an awful lot of money into creating good data now. Like, you know, the fact is data was a train wreck when this all started. Like every company's data was a mess.

36:09Couldn't talk to each other. It's spread out over 24 different databases on all these different platforms. And I know there's kind of an industry built now out of creating quality synthetic data. Not even synthetic data. Even in cleaning your real data to be used for training. I guess the idea being that, you know, better data could train a better model with less effort than what, you know, throwing Play-Doh at the wall and seeing what sticks does. 100%. Yeah. So, so important.

36:45Corey Noles:But also just like if you're someone using this, do you have to be a, you know, ML researcher to be able to use these tools? Like who can use them? yeah silico right now our ai neuroscientist for for training and understanding models is in early access right now so uh people can request access on our website and it will be available for general access soon and so uh right now it's uh it's meant for ai researchers um and people who want to kind of dig deep into training and understanding their model uh but the hope is that over time, more and more people are going to be getting into model training.

37:29And I mean, we're seeing that just explode. The number of people who want to train models, who want to get into their models and really understand them, that desire is there. I just don't think that the tooling has been quite there yet. And so that's kind of what we built Silico to do. So it's shipped with all of our frontier interpretability libraries, skills, the agent and knows how to use anything new that we're landing onto the product as well, like a block sparse featurizer. And it also comes with training libraries, like our SFT and RL libraries that come bundled with compute. It's all orchestrated.

38:11And so there are people who this is the first time they're training models on our platform. And agents are getting really, really good. that they can help you train their models for you. It really is. I did a project a few months back in Codex with just building out a tiny LLM. I guess an LM. Yeah, an SLM. It definitely left large steps. But it was really fun, and I learned so much along the way just in terms of, like, what's happening at different stages and how it works. And I think it's really interesting to be able to do that. I noticed in looking at Silico too that you guys show a pretty big reduction in hallucinations in using it for LLMs.

39:00Is that correct? We could, yeah. I think it really just depends on whether that's what you're kind of using Silico to do. It's almost like a blank canvas. Like what do you want to do with Silico? This is like where we want you to do all of your AI research, all of your AI training, debugging, understanding. and if reducing hallucinations in this particular open model that you're tuning is really important, let's say this is a legal use case where you absolutely need the model to cite all its sources and reduce hallucinations, it can kind of help you dive deep into identifying the structure that causes the model to hallucinate or finding the concept of confidence within the model, which is an extractable concept.

39:53It probably is a curved geometric structure within the model or multiple. And you can take that confidence and then either expose that to the user or have the model hedge while it's less confident or just, you know, kick it up to a smarter model when your cheaper model is less confident and you need to bring out the big guns and spend more on tokens. So there are all types of things that you can do here. So tell me if this is correct. If with your technology you can see what is happening and identify where a trade is taking place, can you then dial that up, dial that down, put a control that a user, end user could tweak to adjust said behavior?

40:42100%. So that's what we call steering. That's dialing up or dialing down an internal concept so that the model behaves differently. And that's often how we verify whether we've actually kind of extracted the correct concept in a model. Because what we really care about is that the concept is causal so that it affects downstream output such that we know that we have, you know, a rabbit ear in our hands if we can remove the rabbit ear or something like that, you know. Or a better example is we know that we've found hallucinations if we've been able to reduce them. Way better example.

41:22Corey Noles:But hey, you never know. If rabbit ears are showing up, you don't want them. That's right. And so, yeah, that's called steering. usually we want to bake that back into the model. You can do inference time steering and just have that be ephemeral for that session. But if you want that model to kind of retain that concept dialed up, you typically want to train that into the model in some way, and we have techniques to allow you to do that. Wow. Could you use this for some sort of like, I forget the exact term for it. I'm sure it'll remind me, like continual learning in this case being where like you sort of are training it on a continuous basis.

42:04Corey Noles:I know there's a term that people have talked about this. Is it inference time tuning? Inference time training? Yeah. Yeah. I guess like there, I mean, there's a lot, this could be a whole other rabbit hole, but yeah, I guess like this definitely could be a way to unlock continual learning where you actually can closely watch the updates to your model to make sure that they're good and they're not degrading the model and that you're not kind of learning anything that you don't want to which may be a risk with continual learning where the model kind of so models often like just a simple example they they often lose their safety tuning when you continue post training on them we have techniques allow you to maintain to safety guardrails as you continue post training on them you say hey i want to learn the new stuff but i don't want you to forget that you have all of these like safety guardrails that you need to maintain in your system and uh i think some of continual learning like some of the challenges with full continual learning um some of it is just on the serving side like infrastructure and you know may not make sense in places but some of the problems are just like you don't know what's going to happen to your model as it's rolling out and learning new things from new situations and catastrophic forgetting could be an element too where it just like it forgets the good stuff that it you know originally made it work so well 100 and so what this can enable is flows where you kind of diff two models let's say the model like learns in a day and it's at night and uh you know you can kind of diff this checkpoint with the previous day's checkpoint and then just roll back if the model yeah you know uh This is the most naive way, but you can just roll back, like, you know, learn something you didn't like, you know, that day.

43:59But if it did learn some good stuff that day, like, you know, keep it rolling. Yeah.

44:04Corey Noles:I'm trying to think, too, like how businesses could use it. So, you know, maybe you have an AI person in your company or you can hire someone. And it would be cool because I think with the current economic landscape as it stands now, it's clear that there's going to be cost pressure on any sort of large scale deployment of AI in any business. Right. So especially doing it over the cloud. So that has woken a lot of people up, you know, with Fable and all that to to the need to train your own models and to, you know, have more control over that. so I could just see this being very useful for people who are interested in doing that 100 yeah I mean that's really a large part of our our vision which we want to enable people to be able to own their own destiny and train their own models and understand that that training process and I think more people should be training models more people should have access to this like critical technology that, um, yeah, I don't, I don't think that, you know, AI research and model training should be confined to a few big laps.

45:13I think it needs to be proliferated and, uh, this like knowledge, this like ability to, uh, yeah, this is like, it's like not being able to write software as a software company, you know, it's like, like train and generate your AI models if you're like a new AI company, you know, at this stage. And many, many more people I think will wake up to that, both with like increasing token spends but also with like increased competition from big labs. What's along the way toward where you are now, what has been the biggest, it's a two-part question, what has been the biggest surprise you've stumbled across that you've found?

46:02And what's the golden goose, the holy grail that everyone's searching for? Yeah. We started the company a little over two years ago. And at that point in time, it was really unclear whether we were going to be able to really find that much interesting at all in the biggest models. I think there are a lot of people who doubted that and still do that interpretability is worthwhile or possible. I think they're very, very wrong. And we're proving that every day.

46:42Corey Noles:What a naive opinion for people who believe that. I mean, I think it's understandable since it's been a while. Like the Interp community has been chugging on for a while with not much downstream application to show for it. And so we are really focused on making sure that it's really useful in the training process so that we can make sure that interpretability is actually useful.

47:08But yeah, I think that when we first started the company, I thought we were going to be trudging through the mud for like maybe many, many, many years in order to find anything interesting. so I was buckled down ready to just like build a lab that you know wouldn't find anything for three years and then like would maybe have a breakthrough but you know I think this has been almost like the story with interpretability it's like everywhere we look there's structure there's really interesting concepts there's like there's just an enormous amount of low-hanging fruit and there's almost nobody looking into this.

47:47The total number of full-time industry researchers is probably in the hundreds, just a few hundred. And this is the most consequential technology of all time. I'm really wondering why more people aren't curious about this technology and looking under the hood. And so maybe the golden goose here to relate it back to your question is just more people should try. to look inside models and see what's going on. You'll probably find something interesting. And just for some reason, like not very many people are trying. Yeah.

48:24Corey Noles:Well, I know one of the things that you mentioned earlier is that, you know, you want to be able to build almost the equivalent of custom models that you have engineered like software from the beginning. So it could the ability to do that. I mean, the ability to do that starts from, We know exactly how these models work. We can speak their language. So we're going to be able, we can tell them exactly what to do more or less, right? That would be a golden goose. I don't know if that's exactly what you're looking for, but. That's right. Yeah. And I wouldn't say we're fully there yet, but yeah, this future of intentional design where we can actually design models like software in a way that where we can read out their internals, where we can debug them, we can edit them, we can actually shape them and steer them during training.

49:13I think this is possible that's really the kind of north star of the organization and we're getting really close to being able to do that across every single stage of model training and then the question is like how do we proliferate this how do we give for us like we want everybody to have this capability because we both think that this makes the world safer as well as like just gives more people the ability to control their own destiny.

49:43Corey Noles:Yeah, which I think should be the whole purpose of all of this. Yep. And I think this is also, I think, the ability to kind of design models with intention is really important, especially for science, for healthcare, for life sciences, where you actually really, really need to know that your model is doing exactly what you think it's doing, or else you'll create, yeah, Yeah, or you may not. A rimpit science. Cronenberg level science. No, I'm just kidding. Yeah, so I think this is gating for enormous scientific advancements that I really want to make happen as quickly as we possibly can. I have one question I want to get in before we go, and I'm sure Grant's got three.

50:29Just kidding, Grant. Five, but I'll limit myself to three. Okay, okay. Okay. For robotics and vision models, what, in your opinion, does neural geometry help reveal that maybe standard model performance benchmarks might not? so we do we are working with a few robotics companies and we have a a robot arm in the office that i've been meaning to play a lot more with um which is really fun uh actually worked in a robotics lab in college so it's good to uh kind of you know get back into it but uh i think that in robotics a big problem is kind of this like sim to real gap where yeah models and and sim don't behave like you know models in in reality and understanding like what changes in terms of the structural internal structure of these models from sim to real is really important to to remove that cap i think that uh there are also just really surprising things that robotics models learn that where a large you know percentage of its parameters may not actually be you know dedicated towards the actual tasks that you want it to be really good at.

51:47But I really am a believer. I think general robotics, generally intelligent robotics is going to work at some point. Yeah.

52:00And I think we just haven't quite nailed how to actually train them and how these representations generalize but i guess like that would be the role of interpretability here which is like help these models generalize and converge faster make sure they're not learning um any strange artifacts in the model like the actual features of these models may be represented directly in the model weights uh so like i don't know sue who trains the model in this way uh the model knows and that's like represent it's like a representation it's a little more sue than the other robots Yeah, exactly. And I think like there are just a lot of ways that we can kind of, yeah, use our understanding of this structure in order to help these models like generalize and perform more effectively in the real world.

52:52That's awesome.

52:53Corey Noles:Right. Have you worked with any multimodal models? So I know you mentioned you haven't really gone into the large language side of it, but what about multimodal and like, you know, models that can do both images on text and video? Yeah. Oh, we work a ton with large language models. It's just for this particular block sparse futurizer, we only have that applied on image and video models. Thanks for clearing up. Yeah, no worries at all. And yeah, we work with models across modalities, sizes, architectures, like the techniques and interpretability tend to scale quite nicely. Because if you can look at a feature, then you can scale it up to an arbitrary number of features of the model.

53:38Uh, it just requires more compute and then a bunch of info work. Um, right.

53:45Corey Noles:And is there anything unique about multimodal models? Um, or like even let's say like what people will call a world model when they think of models that can generate playable or movable, I guess, steerable worlds. Yeah. It's a lot of different. Yeah. Go ahead. With multimodal models. um what gets particularly interesting is i guess like shared representation spaces between let's say like image and language like how does the model like represent both of those using the same neuron or the same feature and uh you start to get to like really interesting things where um the model can represent but they they do share representation spaces so like a I mean, we found like a Trump neuron that also activates on Trump's face.

54:33And same for like Obama. You know, it's like neurons like are shared across like, you know, the image and the language because it is a concept. And the model knows that these concepts are generalizable and composable. and this is also useful for what's increasingly common in biology and life sciences is you don't just want one modality like genomics you want like all the modalities all the data that you

55:02Corey Noles:can get your your hands on because they're all it's all it all works together and kind of affects each other so i was going to ask you is is the problem or is the the difficulty there like uh just pure scale perspective where because of so much data you need to model that's all like working together or could interpretability help with like trying to figure out the complexity of how organs work inside your body? Totally. Yeah. I think it can definitely help. Part of the problem is scale. Part of the problem is you're training a new type of model that nobody's trained before. So it's difficult in many unforeseen ways.

55:38But one big question that some of our partners want to answer is what modalities contribute most to the generalizable representations of the model so that I can buy more of that data so it kind of goes to the like how do you train your model to be as performant as as it can be and what's like contributing to that and so um we can kind of help is there anything that you know right now that's more

56:07Corey Noles:performance like between just like image video um words like this and yeah we probably can't say that at this current moment but there can be something moving forward fair fair fair fair eric thank you so much for joining us today it's been really interesting where can people go to learn more about goodfire you can go to our website uh goodfire.ai um you can also uh i would encourage you to go and and check out silico so you can request access right now. General access will be available soon. It's coming. And yeah, just reach out to us on Twitter too. We're always active there and you can follow our research updates there.

56:52We'll make sure we have all of the links down below. Amazing. I really appreciate you taking the time, man. Thanks again to SAS AI Governance for sponsoring today's video. And thanks to you for watching. Please take just a moment to like and subscribe. Make sure you also take the time to check out our other offerings as well, including the Neuron Newsletter, our blog over at theneuron.ai, or our latest edition, the Neuron Academy. But that's all we have for today, folks. So thanks for joining us. Farewell for now, humans.

57:35Thank you.

From the publisher

 What if neural networks are less like mysterious black boxes and more like systems we can inspect, debug, and eventually design with intention?


In this episode of The Neuron: AI Explained, Corey Noles and Grant Harvey talk with Eric Ho, Cofounder & CEO of Goodfire, an AI interpretability company working to understand what’s happening inside neural networks. Eric explains why models may contain meaningful internal structures — including features, representations, circuits, and curved manifolds — and how mapping those structures could make AI systems safer, more reliable, and more useful.


They discuss why models may “think in shapes,” how Goodfire uses AI to interpret other AI systems, what neural geometry can reveal about hallucinations and model behavior, and why interpretability could change how companies train and control their own models.


They also get into consciousness, robotics, multimodal models, the bitter lesson, and why Eric thinks more people should be looking under the hood of the most consequential technology of our time.


Subscribe to The Neuron for more grounded conversations about how AI actually works: https://www.theneuron.ai/


Sponsored by SAS AI Governance: Visit https://www.sas.com/


The Neuron Academy helps professionals build practical AI skills they can use right away, with lessons on prompting, workflows, and real workplace use cases. Check out https://theneuronacademy.com/ today!

More from The Neuron: AI Explained

All 106 episodes
Inside the Hidden Geometry of AIThe Neuron: AI Explained · 58 min
Listen in VO