In short
Summary of Dwarkesh Podcast Episode with Adam Marblestone: AI is Missing Something Fundamental about the Brain
Episode Overview
- Podcast: Dwarkesh Podcast
- Guest: Adam Marblestone
- Episode Title: AI is Missing Something Fundamental about the Brain
- Date: [Check the podcast link for specific date]
- Description: Adam Marblestone, an expert in brain-computer interfaces, quantum computing, formal mathematics, and AI research, discusses the differences between human intelligence and AI capabilities, particularly focusing on how the brain's architecture and its loss functions contribute to human efficiency in learning and adaptation.
Key Discussion Points
Introduction to the Brain vs. AI
- Sample Efficiency: Humans achieve learning with far fewer data points compared to AI systems.
- Complexity of Loss Functions: The loss functions in AI are often simplified (e.g., cross-entropy), while human brains may have complex loss functions shaped by evolutionary processes.
- Evolutionary Encoding: Discusses how evolution encoded various learning signals and heuristics into the brain's architecture.
Human Learning and AI
- Learning Algorithms: The difference between model-based and model-free reinforcement learning in humans versus AI.
- Neuroscience and AI: Insights from neuroscience that could inform the development of AI, particularly concerning learning algorithms and efficiency.
- Cortex Functionality: The cortex may function as a general prediction engine, constantly predicting variables based on sensory input.
Reward Functions and Evolution
- Evolution of Desires and Intentions: Addresses how the brain encodes high-level desires and intentions without having direct evolutionary experiences related to modern societal constructs (e.g., social status, embarrassment).
- Steve Byrnes’ Theories: Discusses Byrnes' theories about the learning subsystem and the steering subsystem of the brain, and how they interconnect.
Biological vs. Digital Hardware
- Limitations of Biological Hardware: Discusses advantages and limitations of the biological hardware of the brain compared to digital computer architectures.
- Energy Efficiency: The brain operates efficiently on low power, raising questions about the implications for AI hardware development.
The Future of AI Development
- Role of Neuroscience in AI: Advocates for a stronger emphasis on neuroscience to inform AI advancements. Argues that a comprehensive understanding of the brain's architecture and functioning is crucial for developing more advanced AI systems.
- Potential Paradigms: Discusses potential future paradigms for AI learning and the importance of connecting AI models with biological principles.
Key Takeaways
- Understanding the Brain is Crucial: To advance AI toward more human-like intelligence, a better understanding of the brain's complex architecture and learning processes is necessary.
- Evolution's Influence on Learning: The evolutionary history of human beings has shaped how we learn and interact with the world, and this understanding can guide AI development.
- Neuroscience Infrastructure Gaps: Identifying and addressing infrastructure gaps in neuroscience could parallel progress in AI and improve our understanding of intelligence.
- Collaborative Learning Models: The potential for AI to learn from models of human learning and cognition could lead to innovative AI systems that better mimic human thought processes.
Additional Resources
- Adam Marblestone's Work: [Convergent Research](https://www.convergentresearch.org)
- Further Reading: Links to Adam's blog and papers on topics covered in the episode.
Timestamps
- (00:00:00) – The brain’s secret sauce is the reward functions, not the architecture.
- (00:22:20) – Amortized inference and what the genome actually stores.
- (00:42:42) – Model-based vs model-free RL in the brain.
- (00:50:31) – Is biological hardware a limitation or an advantage?
- (01:03:59) – Why a map of the human brain is important.
- (01:23:28) – What value will automating math have?
- (01:38:18) – Architecture of the brain.
Conclusion The episode offers deep insights into the foundational differences between human intelligence and AI, emphasizing the need for an integrated approach that includes neuroscience, evolutionary biology, and advanced computing technologies to eventually create more adaptive and efficient AI systems.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Question of Brain Functionality
0:45 to 2:40
Exploring why AI lacks the capabilities of the human brain.
“But maybe the way that we would think about this now with modern AI, neural nets, deep learning, is that there's sort of these certain key components of that.”
Neuroscience and AI: A Comparison
2:40 to 5:10
Discussion on the key components of AI and brain architecture.
“you know, the cortex has typically this like six-layered structure, layers in a slightly different sense than layers of a neural net.”
The Role of Loss Functions in Learning
5:10 to 7:00
Analyzing how complex loss functions impact learning efficiency.
“conditioned on these being set in this state, and these could be any arbitrary subset of variables in the model.”
Cortex Functions and Learning Models
7:00 to 10:00
Understanding the structure of the cortex and its learning model.
“for saying the wrong thing on your podcast because I'm imagining that young Lacuna is listening and he says, that's not my theory.”
Evolution and High-Level Desires
10:00 to 12:20
How evolution encodes complex desires within the brain's functions.
“And so what are the neurons that matter in the cortex for social status or for friendship?”
Connecting Neural Functions to Social Behavior
12:20 to 14:00
Linking innate reflexes to learned social behaviors in the brain.
“And maybe I'm going to avoid some actions that lead to the thing skittering.”
Neuroscience Insights: The Brain's Steering System
14:00 to 14:52
Explore the brain's steering subsystem and its integration with abstract concepts.
“So now I'm activating your steering subsystem.”
The Limitations of AI Learning Models
14:53 to 16:04
Discuss the complexities of AI models in achieving omnidirectional inference.
“And I think that Steve has an answer to Elia's question, essentially, which is how does the brain ultimately code for these higher level desires and link them up to the more primitive rewards?”
Understanding Neural Network Limitations
16:05 to 18:21
Analyze the challenges in neural networks' representation of different modalities.
“LLMs are maybe just predicting the next token, but vision models maybe are trained in learning to fill in the blanks or reconstruct different pieces or combinations.”
Innovations in Vision Systems
18:22 to 19:25
Learn about advancements in vision systems based on neuroscience principles.
“that will shape the representations differently or that there are clever things that you can do.”
Show all 50 chapters
Multi-Agent Scaling in AI Training
19:26 to 22:16
Investigate the impact of distributing compute across multiple AI agents.
“So this got me thinking, is there an experiment you could run today, even if it's a toy experiment, which would help you anticipate the next scaling paradigm?”
The Concept of Amortized Inference
22:21 to 24:51
Delve into the mechanics of amortized inference in AI models.
“So then you just have to sample like, oh, here's a potential cause.”
Evolutionary Insights on AI and Learning
24:52 to 28:00
Explore the relationship between evolution, learning, and AI architecture.
“Or copy the things that you have sort of like built in.”
Understanding Evolution and Reward Functions
28:00 to 29:15
Exploring how evolution relates to intelligence through reward functions and learning.
“and a set of these bootstrapping cost functions or ways of building up very particular reward signals.”
Neuroscience and Cell Types
29:16 to 30:28
Discussion on the mapping of cell types in the brain and their implications for learning.
“with the thing I was describing where we were talking about the spider, right, of where it learns that just the word spider, you know, triggers the spider, you know, reflex or whatever.”
Wiring and Special Cell Types
30:29 to 32:56
Examining the need for specific cell types in the brain to facilitate learning and response mechanisms.
“this is kind of like one of my obsessions, has done through the Brain Initiative, big neuroscience funding programs.”
Genomic Influences on Brain Function
32:57 to 35:31
Analyzing how different sections of the genome contribute to brain functionality and evolution.
“And that's all the more reason why a lot of the genomic real estate in the genome and in terms of these different cell types and so on would go into wiring up the steering subsystem.”
Cortex Evolution and Functionality
35:32 to 38:26
Investigating the evolution of the cortex and its roles in learning and social interactions.
“I mean, I think you might be able to talk to biologists about this to some degree because you can say, well, we just have a ton in common.”
Reinforcement Learning in Humans
38:27 to 42:00
Contrast between human learning mechanisms and current AI models in reinforcement learning.
“But yeah, I mean, I think that is it that something changed about the cortex and it became possible to do these things?”
Human Design and Learning Mechanisms
42:00 to 42:37
Explore how human design influences learning and communication.
“and see what the visual cues are and hear what they're saying.”
Reinforcement Learning Insights
42:38 to 44:34
Discuss the differences between human learning and algorithms in AI.
“So currently the way these LNs are trained, you know, they are, if they solve the unit test or solve a math problem, that whole trajectory, every token in that trajectory is up-weighted.”
Neuroscience's Role in Learning Algorithms
44:35 to 46:11
Understand how neuroscience informs AI learning strategies.
“And there's a lot of neuroscience evidence that the dopamine is giving this reward prediction error signal rather than just reward, yes, no, you know, a gazillion time steps in the future.”
Cultural Evolution and Learning
46:12 to 48:21
Examine how cultural evolution parallels reinforcement learning.
“is a model of when you do and don't get rewards, right?”
Biological Hardware vs. Computers
50:27 to 52:38
Analyze the advantages and disadvantages of biological versus computational intelligence.
“Are you like, fuck, we would be so much smarter if we didn't have to deal with these brains?”
Cellular Mechanisms in Learning
52:39 to 55:41
Investigate cellular processes that support learning and prediction in the brain.
“And so you don't have to do a random number generator and a bunch of Python code basically to generate a sample.”
Algorithmic Implications of Neuroscience
55:42 to 56:00
Discuss the implications of neuroscience on algorithm design in AI.
Understanding the Role of the Cerebellum in Predictive Learning
56:00 to 58:20
Explore how the cerebellum contributes to learning and reflexes, and the complexities of its cellular functions.
“later, I'm going to get like a puff of air in my eyelid or something, right?”
The Different Forms of Intelligence and AGI
58:20 to 1:01:00
Discuss the various forms of intelligence and the challenges in creating AGI that learns effectively in modern environments.
“is way less than the minimum viable set of things you need for it to have human-like social instincts and ethics and stuff like that.”
Neuroscience's Influence on AI Models
1:01:00 to 1:03:40
Examine how neuroscience insights are informing AI developments and the importance of understanding brain algorithms.
“and variants of convolutional neural nets.”
The Complexity of Neural Networks and Interpretability
1:03:40 to 1:06:20
Analyze the challenges in interpreting neural networks and how insights from architecture can enhance our understanding of the brain.
“Like we do see things like TD learning, which, you know, Sutton also invented separately, right?”
The Future of AI and Brain Research
1:06:20 to 1:10:02
Discuss the timeline and implications of AI research in relation to understanding the brain and achieving AGI.
“Well, I guess there are a couple of different views of it.”
Exploring Timelines and Connectomics
1:10:02 to 1:11:44
Discussion on different timelines for AI and connectomics funding.
“So not everybody has a three-year timeline.”
Advancements in Connectomic Mapping
1:11:44 to 1:13:51
Overview of technologies making connectomic brain mapping more affordable.
“Well, so if I just talk about some of the specific things we have going, so with connectomics, so E11 Bio is kind of like our main thing on connectomics.”
Lessons from the Human Genome Project
1:13:51 to 1:15:58
Insights on the value of technological advancements in mapping.
“other research in the field has switched to an optical microscope paradigm.”
Funding and Investment in Neuroscience
1:15:58 to 1:18:49
Discussion on how neuroscience research can be funded and its potential.
“the Focus Research Organization started with technology development, rather than starting with saying we're going to do a human brain or something, let's just brute force it.”
Behavior Cloning and Neural Activity Patterns
1:18:49 to 1:21:15
Exploration of using neural activity patterns to enhance AI training.
“that would never get you the connectome.”
Automating Mathematics with Lean
1:21:15 to 1:24:00
Discussion on the Lean programming language for automating math proofs.
Understanding Lean: A Verifiable Proof Language
1:24:00 to 1:25:20
Learn how the programming language Lean automates the verification of mathematical proofs.
“so first of all, so Lean had been developed for a number of years at Microsoft and other places.”
The Future of Automated Proofs in Mathematics
1:25:20 to 1:26:40
Explore the potential of automated proofs to transform the mathematical field and address conjectures.
Applications of Lean in Cybersecurity
1:26:40 to 1:28:20
Discover how Lean can contribute to creating secure, unhackable software through formal math proofs.
“Otherwise, you would have to prove those theorems using long, complex, passive inference.”
Challenges in Formal Verification and Specifications
1:28:20 to 1:31:00
Analyze the complexities surrounding the specification problem in formal verification of software.
“I don't know enough details of how hard these things are to search for.”
Mathematics and Creativity: The Role of Intuition
1:31:00 to 1:33:00
Examine how automated proof generation might impact the creative aspects of mathematical thinking.
“So secure against, like, what is the property that is exactly prude?”
Accessibility and Innovation in Mathematics
1:33:00 to 1:35:00
Discuss how making mathematics more accessible could lead to innovative ideas from diverse thinkers.
“I think of it as, I think of it maybe a little bit like the, when everybody had to write assembly code or something like that.”
The Future of Provable AI Systems
1:35:00 to 1:37:30
Speculate on how provable AI systems might evolve and their implications for future collaboration.
“but he's like able to synthesize across the neuroscience literature oh learning subsystem, steering subsystem does this all make sense?”
Exploring Brain Representations in Cognitive Science
1:37:30 to 1:38:01
Delve into how the brain represents concepts and models, bridging cognitive science and AI.
“But I think it's really interesting to think about.”
Theorems and Brain Representations
1:38:01 to 1:39:26
Explore how the brain represents the world and the complexities involved.
“It's like, let's prove all the theorems on mass.”
Continual Learning and Memory Storage
1:39:27 to 1:41:28
Discuss the mechanisms of continual learning and how memories are stored in the brain.
“that make some of that approximately emerge, but maybe it would also emerge in a neural net.”
Fast Weights and Brain Function
1:41:29 to 1:43:16
Investigate the concept of fast weights in the brain and their analogous functions.
“I mean, neurons might be doing a couple of things, you know, some fast weight plasticity and some slower plasticity at the same time or synapses that have many states.”
The Concept of the Gap Map
1:43:17 to 1:45:46
Learn about the gap map and its significance in organizing scientific research priorities.
“parameters in the network, which are part of the actual built-in weights.”
Surprising Discoveries in Research Gaps
1:45:47 to 1:49:06
Discover unexpected gaps in research and the infrastructure needed to address them.
“It was just like you had to put this giant mirror in space with a CCD camera and like organize all the people and engineering and stuff to do that.”
Transcript
Automatic transcript. May contain errors.0:00The big million dollar question that I have that I've been trying to get the answer to through all these interviews with AI researchers, how does the brain do it right? We're throwing way more data at these LLMs, and they still have a small fraction of the total capabilities that a human does. So what's going on? Yeah, I mean, this might be the quadrillion dollar question or something like that. You could make an argument this is the most important question in science. I don't claim to know the answer. I also don't really think that the answer will necessarily come even from a lot of smart people thinking about it as much as they are.
0:35my overall meta-level take is that we have to empower the field of neuroscience to just make neuroscience a more powerful field technologically and otherwise to actually be able to crack a question like this. But maybe the way that we would think about this now with modern AI, neural nets, deep learning, is that there's sort of these certain key components of that. There's the architecture. There's maybe hyperparameters of the architecture. How many layers do you have or sort of properties of that architecture there is the learning algorithm itself how do you train it you know back prop gradient descent um is it something else there is how is it initialized okay so if we take the learning part of the system it still may have some initialization of of the weights um and then there are also cost functions there's like what is it being trained to do what's the reward signal what are the loss functions supervision signals My personal hunch within that framework is that the field has neglected the role of this very specific loss functions, very specific cost functions.
1:44Machine learning tends to like mathematically simple loss functions, right? Predict the next token, you know, cross entropy, these simple kind of computer scientists loss functions. I think evolution may have built a lot of complexity into the loss functions. Actually, many different loss functions were different areas, turned on at different stages of development. A lot of Python code, basically, generating a specific curriculum for what different parts of the brain need to learn. Because evolution has seen many times what was successful and unsuccessful, and evolution could encode the knowledge of the learning curriculum.
2:19So in the machine learning framework, maybe we can come back and we can talk about where do the loss functions of the brain come from? Can different loss functions lead to different efficiency of learning? You know, people say like the cortex has got the universal human learning algorithm, the special sense that humans have. What's up with that? This is a huge question, and we don't know. I've seen models where what the cortex, you know, the cortex has typically this like six-layered structure, layers in a slightly different sense than layers of a neural net. It's like any one location in the cortex has six physical layers of tissue as you go in layers of the sheet.
2:54And then those areas then connect to each other, and that's more like the layers of a network. I've seen versions of that where what you're trying to explain is actually just how does it approximate backprop? And what is the cost function for that? What is the network being asked to do? If you sort of are trying to say it's something like backprop, is it doing backprop on next token prediction? Is it doing backprop on classifying images? Or what is it doing? And no one knows. But I think one thought about it, one possibility about it, is that it's just this incredibly general prediction engine.
3:32So any one area of cortex is just trying to predict any, basically can it learn to predict any subset of all the variables it sees from any other subset. So like omnidirectional inference or omnidirectional prediction, whereas an LLM is just you see everything in the context window, and then it computes a very particular conditional probability, which is given all the last thousands of things, what is the very probabilities for all the next token? Yeah. But it would be weird for a large language model to say, you know, the quick brown fox, blank, blank, the lazy dog, and fill in the middle versus do the next token.
4:17If it's doing just forward, it can learn how to do that stuff in this emergent level of in-context learning, but natively it's just predicting the next token. What if the cortex is just natively made so that any area of cortex can predict any pattern in any subset of its inputs given any other missing subset. That is a little bit more like, quote-unquote, probabilistic AI. I think a lot of the things I'm saying, by the way, are extremely similar to what Jan LeCun would say. He's really interested in these energy-based models and something like that. It's like the joint distribution of all the variables.
4:55What is the likelihood or unlikelihood of just any combination of variables. And if I clamp some of them, I say, well, definitely these variables are in these states, then I can compute with probabilistic sampling, for example, I can compute, okay, conditioned on these being set in this state, and these could be any arbitrary subset of variables in the model. Can I predict what any other subset is going to do and sample from any other subset given clamping this subset? And I could choose a totally different subset and sample from that subset. So it's omnidirectional inference. And so that could be, there's some parts of cortex that might be like association areas of cortex that may predict vision from audition.
5:39There might be areas that predict things that the more innate part of the brain is going to do. Because remember, this whole thing is basically riding on top of the sort of a lizard brain and lizard body, if you will. And that thing is a thing that's worth predicting too. So you're not just predicting, do I see this or do I see that? But is this muscle about to tense? am I about to have a reflex where I laugh you know is my heart rate about to go up am I about to activate this instinctive behavior based on my higher level understanding of like I can match uh somebody has told me there's a spider on my back to this lizard part that would activate if I was like literally seeing a spider in front of me and you learn to associate the two so that even just from somebody hearing you say there's a spider on your back yeah let's well let's come back to this, and this is partly having to do with Steve Byrne's theories, which I'm recently obsessed about, but on your podcast with Ilya, he said, look, I'm not aware of any good theory of how evolution encodes high-level desires or intentions.
6:45I think this is very connected to all of these questions about the loss functions and the cost functions that the brain would use. And it's a really profound question, right? Like, let's say that I am embarrassed for saying the wrong thing on your podcast because I'm imagining that young Lacuna is listening and he says, that's not my theory. You described energy-based models really badly. That's going to activate in me innate embarrassment and shame and I'm going to want to go hide and whatever. And that's going to activate these innate reflexes. And that's important because I might otherwise get killed by Jan LeCun's marauding army of other...
7:28The French AI researchers coming for you, Adam. And so it's important that I have that instinctual response. But of course, evolution has never seen Jan LeCun or known about energy-based models or known what an important scientist or a podcast is. And so somehow the brain has to encode this desire to not piss off really important people in the tribe or something like this. in a very robust way without knowing in advance all the things that the learning subsystem, okay, of the brain, the part that is learning, cortex and other parts, the cortex is going to learn this world model that's going to include things like Jan LeCun and podcasts.
8:09And evolution has to make sure that those neurons, whatever the Jan LeCun being upset with me neurons, get properly wired up to the shame response or this part of the reward function. and this is important right because if we're going to be able to seek status in the tribe or learn from knowledgeable people as you said or things like that exchange knowledge and skills with friends but not with enemies i mean we have to learn all this stuff so it has to be able to robustly wire these learned features of the world um learn parts of the world model up to uh these innate reward functions and then actually use that to then learn more right because next time i'm not going to try to piss off young lacuna if he emails me that that i got this wrong um And so we're going to do further learning based on that.
8:53So in constructing the reward function, it has to use learned information. But how can evolution—evolution didn't know about Jan LeCun, so how can it do that? And so the basic idea that Steve Burns is proposing is that, well, part of the cortex or other areas like the amygdala that learn, what they're doing is they're modeling the steering subsystem. The steering subsystem is the part with these more innately programmed responses. and the innate programming of these series of reward functions, cost functions, bootstrapping functions that exist. So there are parts of the amygdala, for example, that are able to monitor what those parts do and predict what those parts do.
9:31So how do you find the neurons that are important for social status? Well, you have some innate heuristics of social status, for example, or you have some innate heuristics of friendliness that the steering subsystem can use. and the steering subsystem actually has its own sensory system which is kind of crazy so we think of you know vision as being something that the cortex does but there's also a steering subsystem subcortical visual system called the superior colliculus with innate ability to detect faces for example or threats um so it so there's a visual system that uh has innate heuristics um and that the steering subsystem has its own responses so they'll be part of the amygdala or part of the cortex that is learning to predict those responses.
10:17And so what are the neurons that matter in the cortex for social status or for friendship? Or they're the ones that predict those innate heuristics for friendship, right? So you train a predictor in the cortex and you say, which neurons are part of the predictor? Those are the ones that are now you've actually managed to wire it up. Yeah. This is fascinating. I feel like I still don't understand. I understand how the cortex could learn how this primitive part of the brain would respond to um so it can obviously it has these labels on here's literally a picture of a spider and this is bad like be scared of this right and then the cortex learns that this is bad because the innate part tells it that but then it has to generalize to okay the spider's on my back yes and somebody's telling me the spider's on your back.
11:06That's also bad. Yes. But it never got supervision on that. Right. So how does it? Well, it's because the learning subsystem is a powerful learning algorithm that does have generalization, that is capable of generalization. So the steering subsystem, these are the innate responses. So you're going to have some, let's say, built into your steering subsystem, these lower brain areas, hypothalamus, brainstem, et cetera. And again, they include, they have their own primitive sensory systems. So there may be an innate response. If I see something that's kind of moving fast toward my body that I didn't previously see was there and is kind of small and dark and high contrast, that might be an insect kind of skittering onto my body, I am going to like flinch, right?
11:52And so there are these innate responses. And so there's going to be some group of neurons, let's say in the hypothalamus, that is the I am flinching. Or I just flinched, right? I just flinched neurons in the hypothalamus. So when you flinch, first of all, that negative contribution to the reward function, you didn't want that to happen, perhaps. But that's a reward function then that doesn't have any generalization in it. So I'm going to avoid that exact situation of the thing skittering toward me. And maybe I'm going to avoid some actions that lead to the thing skittering. So that's something, a generalization you can get.
12:26What Steve calls it is downstream of the reward function. So I'm going to avoid the situation where the spider was skittering toward me. But you're also going to do something else. So there's going to be like a part of your amygdala, say, that is saying, okay, a few milliseconds, hundreds of milliseconds or seconds earlier,
12:48could I have predicted that flinching response? It's going to be a group of neurons that is essentially a classifier of, am I about to flinch? And I'm going to have classifiers for that, for every important steering subsystem variable that evolution needs to take care of. Am I about to flinch? Am I talking to a friend? Should I laugh now? Is the friend high status? Whatever variables the hypothalamus brainstem contain, am I about to taste salt? So it's going to have all these variables. And for each one, it's going to have a predictor. It's going to train that predictor. Now, the predictor that it trains, that can have some generalization.
13:20And the reason it can have some generalization is because it just has a totally different input. So its input data might be things like the word spider. but the word spider can activate in all sorts of situations that lead to the word spider activating in your world model so if you have a complex world model which really complex features that inherently gives you some generalization it's not just the thing skittering toward me it's even the word spider or the concept of spider is going to cause that to trigger and this predictor can learn that so whatever spider neurons are in my world model which could even be a book about spiders or a room where there are spiders or whatever that is.
13:58The amount of heebie-jeebies that this conversation is eliciting in the audience is like. So now I'm activating your steering subsystem. Your steering subsystem, spider hypothalamus, a subgroup of neurons of skittering insects are activating based on these very abstract concepts in the conversation. If you keep going, I'm going to have to put in a trigger warning. That's because you learned this. And the cortex inherently has the ability to generalize because it's just predicting based on these very abstract variables and all these integrated information that it has. Whereas the steering system only can use whatever the superior calculus and a few other sensors can spit out.
14:32By the way, it's remarkable that the person who's made this connection between different pieces of neuroscience, Stephen Burns, like former physicist, has for the last few years has been trying to synthesize. He's an AI safety researcher. He's just synthesizing. This comes back to the academic incentives thing. I think that this is a little bit hard to say, what is the exact next experiment? How am I going to publish a paper on this? How am I going to train my grad student to do this? It's very speculative. But there's a lot in the neuroscience literature, and Stephen has been able to pull this together.
14:56And I think that Steve has an answer to Elia's question, essentially, which is how does the brain ultimately code for these higher level desires and link them up to the more primitive rewards? yeah very naive question but why can't we achieve this omnidirectional inference by just training the model to not just map from a token to next token but remove the masks right the training so it maps every token to every token or um come up with more labels between video and audio and text so that it it's forced to map one to each one i mean that may be that may be the way. So it's not clear to me. Some people think that there's sort of a different way that it does probabilistic inference or a different learning algorithm that isn't backprop.
15:44There might be like other ways of learning energy-based models or other things like that that you can imagine, but that is involved in being able to do this and that the brain has that. But I think there's a version of it where, you know, what the brain does is like crappy versions of backprop to learn to predict through a few layers. And that, yeah, it's kind of like a multimodal foundation model. Yeah, so maybe the cortex is just kind of like certain kinds of foundation models. LLMs are maybe just predicting the next token, but vision models maybe are trained in learning to fill in the blanks or reconstruct different pieces or combinations.
16:17But I think that it does it in an extremely flexible way. So if you train a model to just fill in this blank at the center, OK, that's great. But what if you didn't train it to fill in this other blank over to the left, then it doesn't know how to do that. It's not part of its repertoire of predictions that are amortized into the network. Whereas with a really powerful inference system, you could choose at test time what is the subset of variables it needs to infer and which ones are clamped. Okay, two sub-questions. One, it makes you wonder whether the thing that is lacking in artificial neural networks is less about the reward function and more about the encoder or the embedding, which maybe the issue is that you're not representing video and audio and text in the right latent abstraction such that they could intermingle and conflict.
17:16Maybe this is also related to why LLM seem bad at drawing connections between different ideas. It's like, are the ideas represented at a level of generality at which you could notice different connections. Well, the problem is these questions are all commingled. So if we don't know if it's doing a backprop-like learning and we don't know if it's doing energy-based models and we don't know how these areas are even connected in the first place, it's very hard to really get to the ground truth of this. But yeah, it's possible. I mean, I think that people have done some work. My friend Joel DiPello actually did something some years ago where I think he put a model.
17:48I think it was a model of V1, of sort of specifically how the early visual cortex represents images and put that as like an input into like a convnet and that like improves some things. So it could be like differences. The retina is also doing, you know, motion detection and certain things are kind of getting filtered out. So there may be some pre-processing of the sensory data. There may be some clever combinations of which modalities are predicting which or so on that lead to better representation. There may be much more clever things than that. Some people certainly do think that there's inductive biases built in the architecture that will shape the representations differently or that there are clever things that you can do.
18:29So Astera, which is the same organization that employs Steve Behrens, just launched this neuroscience project based on Doris So's work. And she has some ideas about how you can build vision systems that basically require less training. They in-build into the assumptions of the design of the architecture that things like objects are bounded by surfaces, and surfaces have certain types of shapes and relationships of how they occlude each other and stuff like that. So it may be possible to build more assumptions into the network. Evolution may have also put some changes of architecture. It's just I think that also the cost functions and so on may be a key thing that it does.
19:14So Andy Jones has this amazing 2021 paper where he uses AlphaZero to show that you can trade off test time compute and training compute. And while that might seem obvious now, this was three years before people were talking about inference scaling. So this got me thinking, is there an experiment you could run today, even if it's a toy experiment, which would help you anticipate the next scaling paradigm? One idea I had was to see if there was anything to multi-agent scaling. Basically, if you have a fixed budget of training compute, are you going to get the smartest agent by dumping all of it into training one single agent or by sliding that compute up amongst a bunch of models, resulting in a diversity of strategies that get to play off each other?
19:50I didn't know how to turn this question into a concrete experiment, though, so I started brainstorming with Gemini 3 Pro in the Gemini app. Gemini helped me think through a bunch of different judgment calls. For example, how do you turn the training loop from self-play to this kind of co-evolutionary leak training? How do you initialize and then maintain diversity amongst different AlphaZero agents? How do you even split up the compute between these agents in the first place? I found this clean implementation of AlphaGo Zero, which I then forked and opened up in Antigravity, which is Google's agent-first IDE.
20:21The code was originally written in 2017, and it was meant to be trained on a single GPU of that time. But I needed to train multiple whole separate populations of AlphaZero agents, so I needed to speed things up. I rented a beefcake of a GPU node, but I needed to refactor the whole implementation to take advantage of all this scale and parallelism. Gemini suggested two different ways to parallelize cell play. One which would involve higher GPU context switching, and the other would involve higher communication overhead. I wasn't sure which one to pick, so I just asked Gemini. And not only did it get both of them working in minutes, but it autonomously created and then ran a benchmark to see which one was best.
21:00It would have taken me a week to implement either one of these options. Think about how many judgment calls a software engineer working on an actually complex project has to make. If they have to spend weeks architecting some optimization or feature before they can see whether it will work out, they will just get to test out so many fewer ideas. Anyways, with all this help from Gemini, I actually ran the experiment and got some results. Now, please keep in mind that I'm running this experiment on an anemic budget of compute, and it's very possible I made some mistakes in implementation. but it looks like there can be gains from splitting up a fixed budget of trading compute amongst multiple agents rather than just dumping it all into one.
21:37Just to reiterate how surprising this is, the best agent in the population of 16 is getting 1 16th the amount of trading compute as the agent trained on self-play alone. And yet it still outperforms the agent that is hogging all of the compute. The whole process of Vibe coding this experiment with Gemini was really absorbing and fun. It gave me the chance to actually understand how AlphaZero works and to understand the design space around decisions about the hyperparameters and how search is done and how you do this kind of co-evolutionary training, rather than getting bogged down in my very novice abilities as an engineer.
22:15Go to Gemini.Google.com to try it out. I want to talk about this idea that you just glanced off of which was amortized inference and maybe I should try to explain what I think it means because I think it's probably wrong and this will help you correct me it's been a few years for me too right now the way the models work is you have an input it maps it to an output and this is amortizing a process that the real process which we think is like what intelligence is which is like you have some prior over how the world could be like what are the causes that make the world the way that it is and then the way when you see some observation you should be like okay here's all the ways the world could be um this cause explains what's happening best now this like doing this calculation over every possible cause is computationally intractable.
23:16So then you just have to sample like, oh, here's a potential cause. Does this explain this observation? No, forget it. Let's keep sampling. And then eventually you get the cause. The cause explains the observation. And then this becomes your posterior. That's actually pretty good, I think, of sort of, yeah. So yeah, this Bayesian inference, like in general, is like of this very intractable thing. Right. The algorithms that we have for doing that tend to require taking a lot of samples, Monte Carlo methods, taking a lot of samples. Yeah. And taking samples takes time. I mean, this is like the original, like, Boltzmann machines and stuff were using techniques like this.
23:55And still it's used with probabilistic programming, other types of methods often. And so, yeah, so the Bayesian inference problem, which is like basically the problem of like perception, like given some model of the world and given some data, Like, how should I update my, what are the variables, you know, missing variables in my internal model? And I guess the idea is that neural networks are hopefully, obviously there's mechanistically the neural network is not starting with like, here is my model of the world and I'm going to try to explain this data. But the hope is that instead of starting with, hey, does this cause explain this observation?
24:32No. Did this cause explain this explanation? Yes. What you do is just like observation. What's the cause that the neural net thinks is the best one? Yeah. Observation to cause. So the feed forward goes observation to cause. Observation to cause. To the output. Yes. You don't have to evaluate all these energy values or whatever and sample around to make them higher and lower. you just say approximately that process would result in this being the top one or something like that yeah one way to think about it might be that test time compute inference time compute is actually doing this sampling again because you literally read its chain of thought it's like actually doing this toy example we're talking about where it's like oh can i solve this problem by doing x yeah i need a different approach and this raises the question i mean over time it is the case that the capabilities which were uh which required inference time compute to elicit get distilled into the model so you're amortizing the thing which previously you needed to do these like rollouts like monte carlo rollouts to um to figure out and so in general there maybe there's this principle of digital minds which can be copied have different trade-offs which are relevant than biological minds which cannot and so in general it should make sense to amortize more things because you can literally copy the amortization, right?
25:48Or copy the things that you have sort of like built in. And this is a tangential question where it might be interesting to speculate about in the future as these things become more intelligent and the way we train them becomes more economically rational, what will make sense to amortize into these minds which evolution did not think it was worth amortizing into biological minds. You have to retrain every time. Right. I mean, first of all, I think the probabilistic AI people would be like, of course, you need test time compute because this inference problem is really hard. And the only ways we know how to do it involve lots of test time compute.
26:24Otherwise, it's just a crappy approximation that's never going to like you have to do infinite data or something to like make this. So I think some of the probabilistic people will be like, no, it's like inherently probabilistic. And like amortizing it in this way, like just doesn't make sense. And so and they might then also point to the brain and say, OK, well, the brain, the neurons are kind of stochastic and they're sampling and they're doing doing things. And so maybe the brain actually is doing more like the non-amortized inference, the real inference. But it's also kind of strange how perception can work in just like milliseconds or whatever.
26:52It doesn't seem like it uses that much sampling. So it's also clearly also doing some kind of baking things into like approximate forward passes or something like that to do this. And yeah, so in the future, you know, I don't know. I mean, I think, is it already a trend to some degree that things that are people were having to use test time compute for are getting like used to train back the base model, right? Yeah. Yeah. So now it can do it in one pass. Right. Yeah. So I mean, I think, yeah, you know, maybe evolution did or didn't do that. But I think evolution still has to pass everything through the genome, right, to build the network.
Read the full transcript
27:34And the environment in which humans are living is very dynamic, right? And so maybe that's, if we believe this is true, that there's a learning subsystem per Steve Burns and a steering subsystem, the learning subsystem doesn't have a lot of pre-initialization or pre-training. It has a certain architecture, but then within lifetime it learns. then evolution didn't actually like amortize that much into that network. It amortized it instead into a set of innate behaviors and a set of these bootstrapping cost functions or ways of building up very particular reward signals. This framework helps explain this mystery that people have pointed out and I've asked a few guests about, which is if you want to analogize evolution to pre-training, well, how do you explain the fact that so little information is conveyed through the genome.
28:25So three gigabytes is the size of the total human genome. Obviously, a small fraction of that is actually relevant to coding at the brain. And if previously people made this analogy that actually evolution has found the hyperparameters of the model, the numbers which tell you how many layers should there be, the architecture basically, right? Like how should things be wired together? But if a big part of the story that increases sample efficiency, aids learning, generally makes systems more performant is the reward function, is the loss function. And if evolution found those loss functions, which aid learning, then it actually kind of makes sense how you can build an intelligence with so little information because the reward function, you write in Python, the reward function is literally a line.
29:11And so you just have a thousand lines like this. And that doesn't take up that much space. Yes. And it also gets to do this generalization thing with the thing I was describing where we were talking about the spider, right, of where it learns that just the word spider, you know, triggers the spider, you know, reflex or whatever. It gets to exploit that too, right? So it gets to build a reward function that actually has a bunch of generalization in it just by specifying these innate spider stuff and the thought assessors, as Steve calls them, that do the learning. So that's like potentially a really compact solution to building up these more complex reward functions too that you need.
29:44So it doesn't have to anticipate everything about the future of the reward function, and just anticipate what variables are relevant and what are heuristics for finding what those variables are. And then, yeah, so then it has to have a very compact specification for the learning algorithm and basic architecture of the learning subsystem. And then it has to specify all this Python code of all the stuff about the spiders and all the stuff about friends and all the stuff about your mother and all the stuff about mating and social groups and joint eye contact. It has to specify all that stuff. And so is this really true?
30:16And so I think that there is some evidence for it. So Fei Chen and Evan McCosco and various other researchers who have been doing like these single cell atlases. So one of the things that neuroscience technology or scaling up neuroscience technology, again, this is kind of like one of my obsessions, has done through the Brain Initiative, big neuroscience funding programs. They've basically gone through different areas, especially the mouse brain. and map like where are the different cell types? How many different types of cells are there in different areas of cortex? Are they the same across different areas?
30:55And then you look at these subcortical regions, which are more like the like steering subsystem or reward function generating regions. How many different types of cells do they have and which neurons types do they have? We don't know how they're all connected and exactly what they do or what the circuits are, what they mean, but you can just like quantify like how many different kinds of cells are there with sequencing the RNA. And there are a lot more weird and diverse and bespoke cell types in the steering subsystem, basically, than there are in the learning subsystem. Like the cortical cell types, there's enough to build, it seems like there's enough to build a learning algorithm up there and specify some hyperparameters.
31:34And in the steering subsystem, there's like a gazillion, you know, thousands of really weird cells, which might be like the one for the spider flinch reflex and the one for I'm about to taste salt. Sorry, but why would each reward function need a different cell type? Well, so this is where you get innately wired circuits, right? So in the learning algorithm part, in the learning subsystem, you specify the initial architecture, you specify a learning algorithm. All the juice is happening through plasticity of the synapses, changes of the synapses within that big network. but it's kind of like a relatively repeating architecture um how it's initialized it's just like um the amount of python code needed to make you know an eight-layer transformer is not that different from wanting to make a three-layer transformer right you're just replicating yeah whereas all this python code for the reward function you know if superior click list sees something that's skittering and land you know you're feeling goosebumps on your skin or whatever than trigger spider reflex.
32:33That's just a bunch of like bespoke species-specific situation-specific crap. The cortex doesn't know about spiders. It just knows about layers and learning. But you're saying that the only way to have this, like write this reward function is to have a special cell type. Yeah. Yeah, well, I think so. I think you either have to have a special cell types or you have to somehow otherwise get special wiring rules that evolution can say, this neuron needs to wire to this neuron without any learning and the way that that is most likely to happen i think is that those cells express like different receptors and proteins that say okay when this one comes in contact with this one let's form a synapse so it's genetic wiring um yeah and those need cell types to do it yeah i'm sure this would make a lot more sense if i knew 101 neuroscience but like it seems like there's still a lot of complexity or generality rather in the steering system so if the steering system has its own visual uh system that's separate from the visual cortex yeah different features still need to plug into that vision system in the so like the spider thing needs to plug into it and also the um the uh a love thing needs to plug into it, et cetera, et cetera.
33:54Yes. So it seems complicated. No, it's still complicated. And that's all the more reason why a lot of the genomic real estate in the genome and in terms of these different cell types and so on would go into wiring up the steering subsystem. And can we tell - Pre-wiring it. Can we tell how much of the genome is clearly working? So I guess you could tell how many are relevant to producing the RNA that manifest or the epigenetics that manifest in different cell types in the brain, right? Yeah, this is what the cell types helps you get at it. I don't think it's exactly like, oh, this percent of the genome is doing this.
34:27But you could say, okay, in all these steering subsypes, how many different genes are involved in sort of specifying which is which and how they wire and how much genomic real estate do those genes take up versus the ones that specify visual cortex versus auditory cortex? You kind of are just reusing the same genes to do the same thing twice, whereas the spider reflex hooking up, yes, you're right. They have to build a vision system and they have to build some auditory systems and touch systems and navigation type systems. So even feeding into the hippocampus and stuff like that, there's head direction cells, even the fly brain, it has innate circuits that figure out its orientation and help it navigate in the world.
35:08And it uses vision, figure out its optical flow of how it's flying and how is its flight related to the wind direction. It has all these innate stuff that I think in the mammal brain, we would all put that and lump that into the steering subsystem. So there's a lot of work. So all the genes basically that go into specifying all the things a fly has to do, we're going to have stuff like that too, just all in the steering subsystem. But do we have some estimate of like, here's how many nucleotides, here are how many megabases it takes to? I don't know. I mean, I think you might be able to talk to biologists about this to some degree because you can say, well, we just have a ton in common.
35:47I mean, we have a lot in common with yeast from a gene's perspective. Yeast is still used as a model for some amount of drug development and stuff like that in biology. And so, so much of the genome is just going towards you have a cell at all. It can recycle waste. It can get energy. It can replicate. um and then then you see what we have in common with a mouse and so we do know at some level that you know the difference is us in a chimpanzee or something and that includes the social instincts and the more advanced you know differences in cortex and so on um it's it's a it's a tiny number of genes that go into these additional amount of making the eight-layer transformer instead of the six-layer transformer or tweaking that reward function this would help explain why the hominid brain exploded in psi so fast which is presumably like tell me this is correct but under the story we um social learning or some other thing increased the ability to learn from the environment like increased our sample efficiency right instead of having to go and kill the boar yourself and figure out like how to do that you can just be like uh the elder told me this out, you make a spear, and then now it increases the incentive to have a bigger cortex, which can learn these things.
37:01Yes. And that can be done with a relatively few genes, because it's really replicating what the mouse already has. It's making more of it. And it's maybe not exactly the same, and there may be tweaks, but it's like, from a perspective, you don't have to reinvent all this stuff. So then how far back in the history of the evolution of the brain does the cortex go back? is the idea that like the cortex has always figured out this omnidirectional inference thing. That's been a solve problem for a long time. And then the big unlock with primates is this, we got the reward function, which increased the returns to having omnidirectional inference.
37:36Or is the cortex, is the omnidirectional inference also something that took a while to unlock? I'm not sure that there's agreement about that. I think there might be specific questions about language, you know, are there tweaks to be, you know, whether that's through auditory and memory, some combination auditory memory regions. There may also be like macro wiring, right? Of like, you need to wire auditory regions into memory regions or something like that and into some of these social instincts to get language, for example, to happen. So there might be, but that might be also a small number of gene changes to be able to say, oh, I just need from my temporal lobe over here going over to the auditory cortex, something, right?
38:13And there is some evidence for the, you know, the Braca's area, Wernicke's area, they're connected with these hippocampus and so on. And so prefrontal cortex. So there's like some small number of genes, maybe for like enabling humans to really properly do language. That could be a big one. But yeah, I mean, I think that is it that something changed about the cortex and it became possible to do these things? Whereas that potential was already there, but there wasn't the incentive to expand that capability and then use it, wired it to these social instincts and use it more. um i mean i would lean somewhat toward the latter i mean i think a mouse i has a lot of similarity in terms of cortex as a human right um although there's that uh the cesena hercule who's all work yeah the um the the number of neurons scales better with weight with primate brains than it does with rodent brains right so yeah does that suggest that There actually was some improvement in the scalability of the cortex?
39:16Maybe, maybe. I'm not super deep on this. There may have been, yeah, changes in architecture, changes in the folding, changes in neuron properties and stuff that somehow slightly tweak this. But there's still a scaling, right? Either way, right? And so I'm not saying there aren't something special about humans in the architecture of the learning subsystem at all. um but yeah i mean i think it's pretty widely thought that this is expanded but then the question is okay well how does that how does that fit in also with the steering subsystem changes and the instincts that make use of this and allow you to bootstrap using this effectively um but i mean just to say a few other things i mean so even the fly brain has some amount of for example even even very far back um i mean i think you've read this this great book the brief history of intelligence, right?
40:06I think this is a really good book. Lots of AI researchers think this is a really good book, it seems like. Yeah, you have some amount of learning going back all the way to anything that has a brain, basically. You have something kind of like primitive reinforcement learning, at least, going back at least to like vertebrates. Like imagine like a zebrafish just like a um these kind of these other branches birds maybe kind of reinvented something kind of cortex-like but it doesn't have the six layers but they have something a little bit cortex-like um so that that's some of those things um after reptiles in some sense birds and mammals both kind of made us up somewhat cortex-like but differently organized thing but even a fly brain has like associative learning senders that um actually do things that maybe look a little bit like this thought assessor concept from Beren's where there's a specific dopamine signal to train specific subgroups of neurons in the fly mushroom body to associate different sensory information with, am I going to get food now or am I going to get hurt now?
41:14Brief tangent. I remember reading in one blog post that Beren Millage wrote that the parts of the cortex which are associated with audio and vision have scaled disproportionately between other primates and humans, whereas the parts associated, say, with odor have not. And I remember him saying something like, this is explained by that kind of data having worse scaling law properties. But I think, and maybe he meant this, but another interpretation of actually what's happening there is that these social reward functions that are built into the steering subsystem needed to make use more of being able to see your elders and see what the visual cues are and hear what they're saying.
42:05And in order to make a sense of these cues, which guide learning, you needed to activate the vision and audio more than older. I mean, there's all this stuff. I feel like it's come up in your shows before, actually. but like even like the design of the human eye where you have like the pupil and the white and everything like we are designed to be able to establish relationships based on joint eye contact. And maybe this came up in the sudden episode. I can't remember. But yeah, we have to bootstrap to the point where we can detect eye contact and where we can communicate by language. Right. And that's like what the first couple of years of life are trying to do.
42:42Okay. I want to ask you about RL. So currently the way these LNs are trained, you know, they are, if they solve the unit test or solve a math problem, that whole trajectory, every token in that trajectory is up-weighted. And what's going on with humans? Are there different types of model-based versus model-free that are happening in different parts of the brain? Yeah, I mean, this is another one of these things. I mean, again, all my answers to these questions, any specific thing I say, it's all just kind of like directionally, this is we can kind of explore around this. I find this interesting.
43:12maybe I feel like the literature points in these directions in some very broad way. What I actually want to do is like go and map the entire mouse brain and like figure this out comprehensively and like make neuroscience a ground truth science. So I don't know, basically. But yeah, I mean, so first of all, I mean, I think with Ilya on the podcast, I mean, he was like, it's weird that you don't use value functions, right? You use like the most dumbest form of RL. And of course, these people are incredibly smart and they're optimizing for how to do it on GPUs. And it's really incredible. what they're achieving.
43:42But conceptually, it's a really dumb form of RL, even compared to what was being done 10 years ago. Even the Atari game playing stuff was using Q learning, which is basically a kind of temporal difference learning. And the temporal difference learning basically means you have some kind of a value function of what action I choose now doesn't just tell me literally what happens immediately after this. It tells me what is the long run consequence of that for my expected, you know, total reward or something like that. And so you have value functions, like, the fact that we don't have, like, value functions at all is, like, in the LLMs is, like, it's crazy.
44:22I think because Ilya said it, I can say it. I know, you know, one one hundredth of what he does about AI. But, like, it's kind of crazy that this is working. But, yeah, I mean, in terms of the brain,
44:41um well so i think there are some parts of the brain that are thought to do something that's very much like model free rl that's sort of parts of the basal ganglia um sort of striatum and basal ganglia they have like a certain finite like it is thought that they have a certain like finite relatively small action space and the types of actions they could take first of all might be like tell the spinal cord or tell the brainstem and spinal cord to do this motor action yes no or it might be more complicated cognitive type actions like tell the thalamus to allow this part of the cortex to talk to this other part or release the memory that's in the hippocampus and start a new one or something right but there's some finite set of actions that kind of come out of the basal ganglia and that it's just a very simple rl so there are probably parts of other brains in our brain that are just like doing very simple naive type rl algorithms um layer one thing on top of that is that some of the major work in neuroscience, like Peter Diane's work and a bunch of work that is part of why I think DeepMind did the temporal difference learning stuff in the first place, is they were very interested in neuroscience.
45:49And there's a lot of neuroscience evidence that the dopamine is giving this reward prediction error signal rather than just reward, yes, no, you know, a gazillion time steps in the future. It's a prediction error. and that's consistent with like learning these value functions. So there's that. And then there's maybe like higher order stuff. So we have these cortex making this world model. Well, one of the things the cortex world model can contain is a model of when you do and don't get rewards, right? Again, it's predicting what the steering subsystem will do. It could be predicting what the basal ganglia will do.
46:22And so you have a model in your cortex that has more generalization and more concepts and all this stuff that says, okay, these types of plans, these types of actions will lead in these types of circumstances to reward. So I have a model of my reward. Some people also think that you can go the other way. And so this is part of the inference picture. There's this idea of RL as inference. You could say, well, conditional on my having a high reward, sample a plan that I would have had to get there. That's inference of the plan part from the reward part. I'm clamping the reward as high and inferring the sampling from plans that could lead to that.
47:02And so if you have this very general cortical thing, it can just do, if you have this very general model-based system and the model, among other things, includes plans and rewards, then you just get it for free, basically. So in neural network parlance, there's a value head associated to the omnidirectional inference that's happening in the pregnancy. Yes, or there's a value input. Oh, okay. Yeah, and it can predict one of the almost sensory variables it can predict is what rewards it's going to get. Yeah, but speaking of this thing about amortizing things, yeah, obviously value is like amortized rollouts of looking up reward.
47:46Yeah, something like that. Yeah, yeah. It's like a statistical average or prediction of it. Yeah. Right. Tangential thought. uh you know joe henrik and others have this idea that the way human societies have learned to do things is just like how do you figure out the you know this kind of bean which actually just almost always poisons you is edible if you do this 10-step incredibly complicated process any one of which if you fail at the bean will be poisonous how do you figure out how to hunt this seal in this particular way with this like particular weapon at this particular time of the year etc um there's no way but uh just like trying shit over generations and it strikes me this is actually very much like model free rl happening at like a civilizational level um no not exactly i mean evolution is the simplest algorithm in some sense right and if we believe that all this can come from evolution like the outer loop can be like extremely not foresighted and yeah right um that that's interesting just like uh hierarchies of evolution model free culture uh evolution model free so what does that tell you maybe the simple algorithms can just get you anything if you do it enough right right yeah yeah i don't know so but yeah so you you have like maybe this yeah evolution model free basal ganglia model free cortex model based culture uh model free potentially um i mean there's like you pay attention to your elders or whatever so there's maybe this like group selection or whatever of these things is like more model free But now I think culture, well, it stores some of the model.
49:25So let's say you want to train an agent to help you with something like processing loan applications. Training an agent to do this requires more than just giving the model access to the right tools. Things like browsers and PDF readers and risk models. There's this level of task and knowledge that you can only get by actually working in an industry. For example, certain loan applications will pass every single automated check despite being super risky. Every single individual part of the application might look safe, but experienced underwriters know to compare across documents to find subtle patterns that signal risk.
49:57Labelbox has experts like this in whatever domain you're focused on, and they will set up highly realistic training environments that include whatever subtle nuances and watchouts you need to look out for. Beyond just building the environment itself, LabelBox provides all the scaffolding you need to capture training data for your agent. They give you the tools to grade agent performance and capture the video of each session, and to reset the entire environment to a clean state between every episode. So whatever domain you're working in, LabelBox can help you train reliable, real-world agents. Learn more at labelbox.com.
50:31Stepping back, how is it? a disadvantage or an advantage for humans that we get to use biological hardware in comparison to computers as it exists now so by what i mean by this question is like if there's the algorithm would the algorithm just qualitatively perform much worse or much better if um inscribed in the hardware of today and the reason to think it might like here's what i mean like you know obviously the brain has had to make a bunch of trade-offs which are not relevant to competing hardware it has to be much more energetically efficient maybe as a result it has to learn a run on slower speeds so that there can be a smaller voltage gap and so the brain runs at 200 hertz um and has to like run on 20 watts on the other hand maybe you know with like robotics we've clearly experienced that fingers are way more nimble than we can make motors so far and so maybe there's something in the brain that is the equivalent of like cognitive uh dexterity which is like maybe due to the fact that we can do unstructured sparsity we can co-locate the memory in the compute.
51:33Where does this all end up? Are you like, fuck, we would be so much smarter if we didn't have to deal with these brains? Or are you like, oh. I mean, I think in the end we will get the best of both worlds somehow. Right. I think an obvious downside of the brain is it cannot be copied. Yeah. You don't have, you know, external read-write access to every neuron in Synapse. Whereas you do. I can just edit something in the weight matrix, you know, in Python or whatever, you know, and load that up and copy that in principle, right? so the fact that it can't be copied and kind of random accessed is like very annoying but otherwise maybe these are it like has a lot of advantages so or it also tells you that you want to like somehow do the co-design of the algorithm and uh it maybe that even doesn't change it that much from all of what we discussed but you want to somehow do this co-design so um yeah how do you do it with really slow low voltage switches that's going to be really important for the energy consumption the co-locating memory and compute so like i think that probably just like hardware companies will try to co-locate memory and compute they will try to use lower voltages allow some stochastic stuff there are some people that think that this like all this probabilistic stuff that we were talking about oh oh it's actually energy-based models and so on is doing it is doing lots of sampling it's not just amortizing everything that the neurons are also very natural for that because they're naturally stochastic.
52:55And so you don't have to do a random number generator and a bunch of Python code basically to generate a sample. The neuron just generates samples and it can tune what the different probabilities are. And so, and like learn, learn those tunings. And so it could be that it's very co-designed with like some kind of inference method or something. Yeah. It'd be hilarious. I mean, the, the message I'm taking on Twitter, you know, Jan LeCoune and Beth Jezos and whatever. They're like, no, maybe I don't know. That is actually one read of me. I haven't really worked on AI at all since LLMs took off. So I'm just like out of the loop.
53:36But I'm surprised. And I think it's amazing how the scaling is working and everything. But yeah, I think Jan LeCoune and Beth Jezos are kind of onto something about the probabilistic models, or at least possibly. And in fact, that's what, you know, all the neuroscientists and all the AI people thought like until 2021 or something right so there's a bunch of cellular stuff happening in the brain that is not just about neuron to neuron synaptic connections how much of that is functionally doing more work than the synapses themselves are doing versus it's just a bunch of collage that you have to do in order to make the synaptic thing work So the way you need to, you know, with a digital mind, you can nudge the synapse, sorry, the parameter extremely easily.
54:24But with a cell to modulate a synapse, according to the gradient signal, it just takes all of this crazy machinery. So like, is it actually doing more than it takes extremely little code to do? So I don't know, but I'm not a believer in the like radical, like, oh, actually memory is not synapses mostly, or like learning is mostly genetic changes or something like that. I think it would just make a lot of sense. I think you put it really well for it to be more like the second thing you said. Like, let's say you want to do weight normalization across all the weights coming out of your neuron, right, or into your neuron.
55:00well you probably have to somehow tell the nucleus about this of the cell and then have that kind of send everything back out to the synapses or something right and so there's going to be a lot of cellular changes right or let's say that you know you just had a lot of plasticity and like you're part of this memory and now that's got consolidated into the cortex or whatever and now we want to reuse you as like a new one that can learn again it's going to be a ton of cellular changes so there's going to be tons of stuff happening in the cell but algorithmically it's not really adding something beyond these algorithms right it's just implementing something that in a digital computer is very easy for us to go and just find the weights and change them and it is a cell it just literally has to do all this with molecular machines itself without any central controller right it's kind of incredible there are some things that cells do i think that that seem like more convincing so in the cerebellum so one of the things the cerebellum has to do is like predict over time like predict what is the time delay you know So let's say that, you know, I see a flash and then, you know, some number of milliseconds later, I'm going to get like a puff of air in my eyelid or something, right?
56:07The cerebellum can be very good at predicting what's the timing between the flash and the air puff so that now your eye will just like close automatically. Like the cerebellum is like involved in that type of reflex, like learned reflex. And there are some cells in the cerebellum where it seems like the cell body is playing a role in storing that time constant changing that time constant of delay versus that all being somehow done with like i'm going to make a longer ring of synapses to make that delay longer it's like no the cell body will just like store that time delay for you um so there are some examples but i'm not a believer like out of the box in like essentially this theory that like what's happening is changes and connections between neurons yeah and that's like the main algorithmic thing that's going on like i i think that's a very good reason to to still believe that it's that rather than some like crazy cellular stuff yeah going back to this whole perspective of like our our intelligence is not just this omnidirectional inference thing that builds a world model but really this system that teaches us what to pay attention to what are the important salient factors to learn from etc i i want to see if there's some intuition we can drive from this but what different kinds of intelligence it might be like.
57:29So it seems like AGI or superhuman intelligence should still have this
57:37ability to learn a world model that's quite general. But then it might be incentivized to pay attention to different things that are relevant for the modern post-singularity environment. How different should we expect different intelligences to be, basically? Yeah, I mean, I think one way of this question is like, is it actually possible to make the paperclip maximizer or whatever, right? If you try to make the paperclip maximizer, does that end up just not being smart or something like that? Because it was just the only reward function it had was like, make paperclips. Interesting, yeah, yeah.
58:12If I channel Steve Burns more, I mean, I think he's very concerned that the sort of minimum viable things in the steering subsystem that you need to get something smart is way less than the minimum viable set of things you need for it to have human-like social instincts and ethics and stuff like that. So a lot of what you want to know about the steering subsystem is actually the specifics of how you do alignment, essentially, or what human behavior and social instincts is versus just what you need for capabilities. And we talked about it in a slightly different way because we were sort of saying, well, in order for humans to learn socially, they need to make eye contact and learn from others.
58:47But we already know from LLMs, right, Right. But depending on your starting point, you can learn language without that stuff. Right. And so. Yeah. And so I think that it probably is possible to make like super powerful, you know, model based RL, you know, optimizing systems and stuff like that, that don't have most of what we have in the human brain reward functions. And as a consequence, might want to maximize paperclips. And that's a concern. Yeah. Right. But you're pointing out that. in order to make a competent paperclip maximizer, the kind of thing that can build the spaceships and learn the physics and whatever, it needs to have some drives which elicit learning, including, say, curiosity and exploration.
59:29Yeah, curiosity and interest in others, interest in social interactions, curiosity. Yeah, but that's pretty minimal, I think. And that's true for humans, but it might be less true for something that's already pre-trained as an LLM or something, right? And so most of why we want to know the steering subsystem, I think, if I'm channeling Steve, is alignment reasons. Right. How confident are we that we even have the right algorithmic, conceptual vocabulary to think about what the brain is doing? And what I mean by this is, you know, there was one big contribution to AI from neuroscience, which was the side of the neuron.
1:00:10like william and you know 1950s just like this original contribution but then it seems like a lot of what we've learned afterwards about what the high-level algorithm the brain is implementing from the backprop to if there's something analogous backprop happening in the brain to always we want doing something like cnn's right to td learning and bellman equations um actor critic whatever yeah seems inspired by what is like we come up with some idea like but maybe we can make AI neural networks work this way. Yeah. And then we notice that something in the brain also works that way. Yes. So why not think there's more things like this where in the future - There may be, yeah.
1:00:47I think the reason that I'm not, I think that we might be onto something is that like the AIs we're making based on these ideas are working surprisingly well. There's also a bunch of like just empirical stuff, like convolutional neural nets and variants of convolutional neural nets. I'm not for sure what the absolute latest latest, but compared to other like models in computational neuroscience of like what the visual system is doing are just like more predictive right so you can just like score um even like pre-trained on like cat pictures and stuff cnn's what is the representational similarity that they have on some arbitrary other image versus you know compared to the brain activations um measured in different ways um jim de carlo's lab has the like brain score and like the ai model is actually like there seems to be some relevance there in terms of like even like neurosciences don't necessarily have something better than that so yes i mean that's just kind of recapitulating what you're saying is that like the best computational neuroscience theories we have seem to have been like invented right largely as a result of ai models um and like find things that work and so find backprop works and then say can we approximate backprop with cortical circuits or something and there's there's kind of been things like that now some people totally disagree with this right um so like yuri buzzaki is a neuroscientist who has a book called the brain from inside out where he basically says like all our psychology concepts like ai concepts all the stuff is just like made up stuff we actually have to do is like figure out what is the actual set of primitives that like the brain actually uses and our vocabulary is not going to be adequate to that we have to start with the brain and make new vocabulary rather than saying back prop and then try to apply that to the brain or something like that and you know he studies a lot of like oscillations and stuff in the brain as opposed to individual neurons and what they do and you know i don't know i i think that there's a case to be made for that and from a kind of research program design perspective i think there's like one thing we should be trying to do is just like simulate a tiny worm or a tiny zebrafish um like from almost like as biophysical or like as as bottom-up as possible like get connectome molecules activity and like just study it as a physical dynamical system and like look what it does um but i don't know i mean just when i like it just feels like the AI is really good fodder for computational neuroscience.
1:03:09And like, those might actually be pretty good models. We should look at that. So I'm not a person who thinks that, I think I both think that there should be a part of the research portfolio that is like totally bottom up and not trying to apply our vocabulary that we learn from AI onto these systems. And that there should be another big part of this that's kind of trying to reverse engineer it using that vocabulary or variance of that vocabulary. And that we should just be pursuing both. And my guess is that the reverse engineering one is actually going to like kind of work-ish or something. Like we do see things like TD learning, which, you know, Sutton also invented separately, right?
1:03:50That must be a crazy feeling to just like, you know, this like equation I wrote down is like in the brain. Yeah, it seems like the dopamine is like doing some of that. Yeah. So let me ask you about this. you know you guys are finding different groups that are trying to yeah figure out what's up in the brain if we had a perfect representation how are you defined it of the brain why think it would actually let us figure out the answer to these questions we have neural networks which are way more interpretable not just because we understand what's in the weight matrices but because there are weight matrices there are these boxes with numbers in them right and even then we can tell very basic things.
1:04:30We can kind of see circuits for very basic pattern matching of following one token with another. I feel like we don't really have an explanation of why LLMs are intelligent just because they're interpretable. I would somewhat dispute it. I think we have some architectural, we have some description of what the LLM is like fundamentally doing. And what that's doing is that I have an architecture and I have a learning rule and I have hyperparameters and I have initialization and I have training data. But those are things we learned from because we built them, not because we interpreted them from seeing the way it's— We built them.
1:05:02Which is the analogous thing to connect to them is like seeing the way it's— What I think we should do is we should describe the brain more in that language of things like architecture's learning rules, initializations, rather than trying to find the golden gate bridge circuit and saying exactly how does this neuron actually— That's going to be some incredibly complicated learned pattern. Yeah, Cod Recording and Tim Lillicrap have this paper from a while ago, maybe five years ago, called what does it mean to understand a neural network or what would it mean to understand a neural network um and what they say is yeah basically that like you could imagine you train a neural network to like compute the digits of pi or something well like some crazy you know it's like it's like this crazy pattern and you also train that thing to like predict the most complicated thing you find predict stock prices basically predict the really complex systems right computational you know computationally complete systems i could predict i could train a neural network to do cellular automata or whatever crazy thing.
1:05:53And it's like, we're never going to be able to fully capture that with interpretability, I think. It's just going to just be doing really complicated computations internally. But we can still say that the way it got that way is that it had an architecture and we gave it this training data and it had this loss function. And so I want to describe the brain in the same way. And I think that this framework that I've been kind of laying out is like, we need to understand the cortex and how it embodies a learning algorithm. I don't need to understand how it computes Golden Gate Bridge. If you can see all the neurons, if you have the connectome, why does that teach you what the learning algorithm is?
1:06:24Well, I guess there are a couple of different views of it. So it depends on the different parts of this portfolio. So on the totally bottom-up, we have to simulate everything portfolio. It kind of just doesn't. You have to just, like, see what are the—you have to make a simulation of the zebrafish brain or something. And then you, like, see what are the, like, emergent dynamics in this. And you come up with new names and new concepts and all that. That's, like, the most extreme bottom-up neuroscience view. but even there the connectome is like really important for doing that bottom biophysical or bottom-up simulation but on the other hand you can say well what if we can actually apply some ideas from ai we basically need to figure out is it an energy-based model or is it you know an amortized you know vae type model you know is it doing backprop or is it doing something else are the learning rules local global i mean if we have some repertoire of possible ideas about this can we just think of the connectome as a huge number of additional constraints that will help to refine to ultimately have a consistent picture of that i think about this for the the steering subsystem stuff too just very basic things about it how many different types of dopamine signal or of steering subsystem signal or thought assessor or so on how many different types of what broad categories are there like even this very basic information that there's more cell types in the hypothalamus than there are in the cortex like that's new information right about how much structure is built there versus somewhere else.
1:07:47Yeah, how many different dopamine neurons are there? Is the wiring between prefrontal and auditory the same as the wiring between prefrontal and visual? You know, it's like the most basic things we don't know. And the problem is learning even the most basic things by a series of bespoke experiments takes an incredibly long time. Whereas just learning all of that at once by getting a connectome is just like way more efficient. What is the timeline on this? Because presumably the idea of this is to, well first inform the development of AI. You want to be able to figure out how we get AIs to want to care about what other people think of its internal thought pattern.
1:08:28But Interp researchers are making progress on this question just by inspecting normal neural networks. There must be some feature. You can do Interp on LLMs that exist. You can't do Interp on a hypothetical model-based reinforcement algorithm like the brain that we will eventually converge to when we do AGI. Fair, fair. But yeah, you know, what timelines on AI do you need for this research to be practical and relevant to AI? I think it's fair to say it's not super practical and relevant if you're in like an AI 2027 scenario. And so like what science I'm doing now is not going to affect the science of like 10 years from now because what's going to affect the science of 10 years from now is the outcome of this like ai 2027 scenario right it kind of doesn't matter that much probably if i have the connectome maybe it slightly tweaks certain things but um but i think there there's a lot of reason to think maybe that we will get a lot out of this paradigm but then the real thing the thing that is like the the trend the like single event that is like transformative for the entire future or something type event is still like you know more than five years away Is that because we haven't captured omni-directional inference?
1:09:44We haven't figured out the right ways to get a mind to pay attention to things in a way that makes it... I mean, I would take the entirety of your collective podcast with everyone as showing the distribution of these things. I don't know. What was Karpathy's timeline? What's Demis' timeline? So not everybody has a three-year timeline. And so I think there's different reasons. And I'm curious. There are different reasons. What are mine? I don't know. I'm just watching your podcast. I'm trying to understand the distribution. I don't have a super strong claim that LLMs can't do it. But is it correct like the data efficiency or is it the.
1:10:21I think part of it is just it is weirdly different than all this brain stuff. Yeah. And so intuitively, it's just weirdly different than all this brain stuff. And I'm kind of waiting for like the thing that starts to look more like brain. Like I think if AlphaZero and model-based RL and all these other things that were being worked on 10 years ago had been giving us the GPT-5 type capabilities, then I would be like, oh, wow, we're both in the right paradigm and seeing the results a priori. So my model, my prior and my data are agreeing. And now it's like, I don't know what exactly my data is. It looks pretty good, but my prior is sort of weird.
1:10:56So yeah, so I don't have a super strong opinion on it. So I think there's a possibility that essentially all other scientific research that is being done is like not is somehow obviated but i don't put a huge amount of probability on that i think my timelines might be more in the like yeah 10 yearish range and if that's the case i mean i think there yeah there is probably a difference between a world where we have connect homes on hard drives and we have understanding of steering subsystem architecture we've compared the the you know even the most basic properties of what are the reward functions cost function architecture etc of you know mouse versus a shrew versus a small primate etc this is practical in 10 years I think it has to be a really big push.
1:11:36Like how much funding? How does it compare to where we are now? It's like billion, low billions dollar scale funding in a very concerted way, I would say. And how much is on it now? Well, so if I just talk about some of the specific things we have going, so with connectomics, so E11 Bio is kind of like our main thing on connectomics. They are basically trying to make the technology of connectomic brain mapping, several orders of magnitude cheaper. So the Wellcome Trust put out a report a year or two ago that basically said to get one mouse brain, the first mouse brain connectome would be like several billion dollars, you know, billions of dollars project.
1:12:20Well, E11 technology and sort of the suite of efforts in the field also are trying to get like a single mouse connectome down to like low tens of millions of dollars. Okay. So that's a mammal brain, right? Now, a human brain is about a thousand times bigger. So if a mouse brain, you can get to 10 million or 20 million, 30 million with technology. You know, if you just naively scale that, okay, human brain is now still billions of dollars to just do one human brain. Can you go beyond that? So can you get a human brain for like less than a billion? But I'm not sure you need every neuron in a human brain.
1:12:51I think we want to, for example, do an entire mouse brain and a human steering subsystem and the entire brains of several different mammals with different social instincts. And so I think that that, with a bunch of technology push and a bunch of concerted effort, can be done in the real significant progress if it's focused effort can be done in the kind of hundreds of millions to low billions. What is the definition of a connectome? Is it presumably it's not a bottom of biophysics model? So is it just that if it can estimate the input-output of a brain? But what is the level of abstraction? So you can give different definitions.
1:13:27And one of the things that's cool about... So the kind of standard approach to connectomics uses the electron microscope and very, very thin slices of brain tissue. And it's basically labeling the cell membranes are going to show up, scatter electrons a lot, and everything else is going to scatter electrons less. But you don't see a lot of details of the molecules, which types of synapses, different synapses of different molecular combinations and properties. E11 and some other research in the field has switched to an optical microscope paradigm. With optical, the photons don't damage the tissue, so you can kind of wash it and look at fragile, gentle molecules.
1:14:02So with E11 approach, you can get a quote-unquote molecularly annotated connectome. So that's not just who is connected to who by some kind of synapse, but what are the molecules that are present at the synapse, what type of cell is that? So a molecularly annotated connectome, that's not exactly the same as having the synaptic weights. That's not exactly the same as being able to simulate the neurons and say what's the functional consequence of having these molecules and connections. But you can also do some amount of activity mapping and try to correlate structure to function. Yeah, so. Interesting.
1:14:39Train an ML model to basically predict the activity from the connectome. What are the lessons to be taken away from the Human Genome Project? Because one way you could look at it is that it was actually a mistake, and you shouldn't have spent whatever billions of dollars getting one genome mapped. Rather, you should have just invested in technologies which have now allows to map genomes for hundreds of dollars. Yeah, well, yeah. So George Church was my PhD advisor. And basically, yeah, I mean, what he's pointed out is that, yeah, it was$3 billion or something, roughly$1 per base pair for the first genome.
1:15:08and then the national human genome research institute basically structured the funding process right and they got a bunch of companies competing to lower the cost um and then the cost dropped like a million fold in 10 years um because they changed the paradigm from uh kind of macroscopic kind of chemical techniques to these individual dna molecules make a little cluster of dna molecules on the microscope and you would see just a few dna molecules at a time on each pixel of the camera would basically give you a different, in parallel, looking at different fragments of DNA. So you parallelize the thing by like millions fold, and that's what reduced the cost by millions fold.
1:15:46And yeah, so I mean, essentially, with switching from electron microscopy to optical connectomics, potentially even future types of connectomics technology, we think there should be similar patterns. That's why E11 with the Focus Research Organization started with technology development, rather than starting with saying we're going to do a human brain or something, let's just brute force it. We said, let's get the cost down with new technology. But then it's still a big thing. Even with new next-generation technology, you still need to spend hundreds of millions on data collection. Is this going to be funded with philanthropy, by governments, by investors?
1:16:22This is very TBD and very much evolving in some sense as we speak. I'm hearing some rumors going around of Connectomics-related companies potentially forming. but so so far 11 has been philanthropy um the national science foundation just put out this call for it for tech labs which is basically somewhat of it is kind of fro inspired or related um i think you could have a tech lab uh for actually going and mapping the mouse brain with this and that would be sort of philanthropy plus government still in a non-profit kind of open source framework um but can uh can companies accelerate that can you credibly link connectomics to AI in the context of a company and get investment for that is like possible.
1:17:05I mean, the cost of training these AIs is increasing so much if you could like tell some story of like, not only are we going to figure out some safety thing, but in fact we will, once we do that, we'll also be able to tell you how AI works. I mean, all these questions. You should like go to these AI labs and just be like, give me one 100th of your projected budget in 2030. I sort of tried a little bit like seven or eight years ago and there was not a lot of interest and maybe now there would be. But yeah, I mean, I think all the things that we've been talking about, like I think it's really fun to talk about, but it's ultimately speculation.
1:17:41What is the actual reason for the energy efficiency of the brain, for example, right? Is it doing real inference or amortized inference or something else? Like this is all gonna be, it's all answerable by neuroscience. It's gonna be hard, but it's actually answerable. And so if you can only do that for low billions of dollars or something to really comprehensively solve that, it seems to me in the grand scheme of trillions of dollars of GPUs and stuff, it actually makes sense to do that investment. And I think investors also just, there's been many labs that have been launched in the last year where they're raising on the valuation of billions for things which are quite credible, but are not like RER, next quarter is going to be whatever.
1:18:21It's like, we're going to discover materials and dot, dot, dot, right? Yes, yes. Moonshot startups or billion dollar, billionaire backed startups, moonshot startups, I see as kind of on a continuum with fros. Fros are a way of channeling philanthropic support and ensuring that it's open source, public benefit, various other things that may be properties of a given fro. But yes, billionaire-backed startups, if they can target the right science, the exact right science, I think there's a lot of ways to do moonshot neuroscience companies that would never get you the connectome. You say, oh, we're going to upload the brain or something, but never actually get the mouse connectome or something, these fundamental things that you need to get to ground truth to science.
1:19:00There are lots of ways to have a moonshot company kind of go wrong and not do the actual science, but there also may be ways to have companies or big corporate labs get involved and actually do it correctly. This brings to mind an idea that you had in a lecture you gave five years ago about, do you want to explain behavior cloning on... Yeah, I mean, actually this is funny because I think that the first time I saw this idea it was i think it actually might have been in a blog post by gwern oh there's always there's always a gwern blog post and there are now academic research efforts and some amount of emerging company type efforts to try to do this so um yeah so normally like let's say i'm training an image classifier or something like that i show it uh pictures of cats and dogs or whatever and they have laid the label cat or dog and i have a neural network supposed to predict the label cat or dog or something like that um that is a limited amount of information per label that you're putting in it's just cat or dog what if i also had predict what is my neural activity pattern when i see a cat or when i see a dog and all the other things um if you add that as like an auxiliary loss function or an auxiliary prediction task, does that sculpt the network to know the information that humans know about cats and dogs and to represent it in a way that's consistent with how the brain represents it and the kind of representational kind of dimensions or geometry of how the brain represents things as opposed to just having these labels?
1:20:39Does that let it generalize better? Does that let it have just richer labeling? And of course, that sounds really challenging. It's very easy to generate lots and lots of labeled cat pictures with scale AI or whatever can do this. It is harder to generate lots and lots of brain activity patterns that correspond to things that you want to train the AI to do. But again, this is just a technological limitation of neuroscience. If every iPhone was also a brain scanner, you would we would not have this problem and we would be training ai with the brain signals and um it's just the order in which technology is developed is that we got gpus before we got portable brain scanners or whatever right and uh that kind of thing what is the ml analog what you'd be doing here because when you distill models you're still looking at the the final layer of like the the log props across um across if you if you do distillation of one model into another that is a certain thing you're just trying to copy one model into another yeah i think that we don't really have a perfect proposal to like distill the brain i think to distill the brain you need like a much more complex brain interface like maybe you could also do that you could make surrogate models um andreas tolias and people like that are doing some amount of neural network surrogate models of brain activity data instead of having your visual cortex do the computation just have the surrogate models you're basically distilling your visual cortex into a neural network to some degree um that's the kind of distillation this is doing something a little different this is basically just saying i'm adding an auxiliary i think of as regularization or i think of it as um adding an auxiliary loss function um that's sort of smoothing out the prediction task to also always be consistent with how the brain represents it like what exactly it might help you things like adversarial examples for example right all right so you're predicting the internal state of the brain yes so in it so you so in addition to predicting the label the vector of labels like yes cat not dog yes you know not boat you know um one shot vector or whatever of one hot vector of yes it's cat instead of these gazillion other categories let's say in this simple example you're also predicting a vector which is like all these brain signal measurements right yeah interesting and so gurn anyway had this long ago blog post of like oh this is like an intermediate thing that's like we talk about whole brain emulation we talk about agi we talk about brain computer interface we should also be talking about this like brain augmented brain data augmented um uh thing that's trained on all your behavior but is also trained on like predicting some of your neural patterns right and you're saying the learning system is already doing this for the steering system yeah and our learning system also has predict the steering subsystem as an auxiliary task yeah yeah and that helps the steering subsystem now the steering subsystem can access that predictor and build a cool reward function using it.
1:23:27Yes. Okay. Separately, you're on the board for Of Lean, which is this formal math language that mathematicians use to prove theorems and so forth. And obviously, there's a bunch of conversation right now about AI automating math. What's your take? Yeah. Well, I think that there are parts of math that it seems like it's pretty well on track to automate. And that has to do with like, so first of all, so Lean had been developed for a number of years at Microsoft and other places. It has become one of the convergent focused research organizations to kind of drive more engineering and focus onto it.
1:24:13So Lean is like this language, programming language, where if you, instead of expressing your math proof on pen and paper. You express it in this programming language Lean. And then at the end, if you do that that way, it is a verifiable language so that you can basically click verify and Lean will tell you whether the conclusions of your proof actually follow perfectly from your assumptions of your proof. So it checks whether the proof is correct automatically. Just like by itself, this is useful for mathematicians collaborating and stuff like that. Like if I'm some amateur mathematician, and I want to add to a proof, you know, Terry Tao is not going to, like, believe my results.
1:24:54But if Lean says it's correct, it's just correct. So it makes it easy for, like, collaboration to happen. But it also makes it easy for correctness of proofs to be an RL signal in very much, yeah, RLVR, you know. It's like a perfect, math proofing is now, formalized math proofing is a formal means that's, like, expressed in something like Lean and verifiable, mechanically verifiable.
1:25:19that becomes a perfect rl vr you know task um yeah and i think that that is going to just just keep working it seems like is the couple billion dollar at least one like billion dollar valuation company harmonic based on this alpha proof is based on this um a couple other emerging really interesting companies um i think that this problem of like rl vr-ing the crap out of math proving is basically going to work and we will be able to have things that search for proofs um and find them um in the same way that we have alpha go or what have you that can search for you know ways of playing the game of go and with that verifiable signal uh works so does this like solve math um there is still the part that has to do with conjecturing new interesting ideas there's still the kind of conceptual organization of math of what is interesting how do you come up with new theorem statements in the first place or even like the very high level breakdown of what strategies you use to do proofs um i mean i think this will shift the burden of that so that humans don't have to do a lot of the mechanical parts of math uh validating lemmas and proofs and checking if the statement of this in this paper is exactly the same as that paper and stuff like that it will just that will just work uh you know if you really think you're going to get all these things we've been talking about real agi it would also be able to make conjectures and you know benji has like a paper as more like theoretical paper there's probably a bunch of other papers emerging about this like is there like a loss function for like good explanations or good conjectures that's like a pretty profound question right um a math a really interesting math proof or statement might be one that can compresses lots of information about other you know has lots of implications for lots of other theorems.
1:27:12Otherwise, you would have to prove those theorems using long, complex, passive inference. Here, if you have this theorem, this theorem is correct. You have short, passive inference to all the other ones. And it's a short, compact statement. So it's like a powerful explanation that explains all the rest of math. And part of what math is doing is making these compact things that explain the other things. It's like the Kogel-Morab complexity of this statement or something. Yeah, of generating all the other statements given that you know this one or stuff like that. Or if you add this, how does it affect the complexity of the rest of the kind of network of proofs?
1:27:40so can you like make a loss function that adds oh i want this proof to be a really highly powerful proof um i think some people are trying to work on that so so maybe you can automate the creativity part um if you had true agi it would do everything a human can do so it would also do the things that the creative mathematicians do but um but way barring that i think just rlvring the crap out of proofs um well i think that's going to be just a really useful tool for mathematicians it's going to accelerate math a lot and change it a lot but not necessarily immediately change everything about it. Will we get, you know, mechanical proof of the Riemann hypothesis or something like that or things like that?
1:28:20Maybe, I don't know. I don't know enough details of how hard these things are to search for. I'm not sure anyone can fully predict that just as we couldn't exactly predict when Go would be solved or something like that. And I think it's going to have lots of really cool applied applications. So one of the things you want to do is you want to have provably stable, secure, unhackable, etc. software. So you can write math proofs about software and say this code, not only does it pass these unit tests, but I can mathematically prove that there's no way to hack it in these ways or no way to mess with the memory or this type of things that hackers use.
1:29:01Or it has these properties. it can use the same lean and same proof to do formally verified software i think that's going to be a really powerful piece of cyber security um that's relevant for all sorts of other ai hacking the world stuff and that yeah if you can prove a remand hypothesis you're also going to be able to to prove insanely complex things about very complex software and then you'll be able to at the LLM synthesize me a software that I can prove is correct, right? Why hasn't provable programming language taken off as a result of LLMs? You would think that this would - I think it's starting to.
1:29:42Yeah, I think it's starting to. I think that one challenge, and we are actually incubating a potential focused research organization on this, is the specification problem. So mathematicians kind of know what interesting theorems they want to formalize. if I have like some code, let's say I have some code that like is involved in running the power grid or something and it has some security properties. Well, what is the formal spec of those properties? The power grid engineers just made this thing, but they don't necessarily know how to lift the formal spec from that. And it's not necessarily easy to come up with the spec that is the spec that you want for your code.
1:30:17People aren't used to coming up with formal specs and there are not a lot of tools for it. So you also have like this kind of user interface plus AI problem of like, what security spec should I be specifying? Is this the spec that I wanted? So there's a spec problem. And it's just been really complex and hard, but it's only just in the last very short time that the LLMs are able to generate verifiable proofs of things that are useful to mathematicians, starting to be able to do some amount of that for software verification, hardware verification. but I think if you project the trends over the next couple years it's possible that it just flips the tide that formal methods based this whole field of formal methods or formal verification provable software which is kind of this weird almost like backwater of more like theoretical part of programming languages and stuff very academically flavored often although there was like this DARPA program that made like a provably secure like quadcopter helicopter and stuff like that.
1:31:21So secure against, like, what is the property that is exactly prude? Not for that particular project, but just in general. Yeah, so... Because obviously things malfunction for all kinds of reasons. You could say that what's going on in this part of the memory over here, which is supposed to be the part the user can access, can't in any way affect what's going on in the memory over here or something like that. Right. Or yeah, things like that. Yeah. Got it. Yeah. So there's two questions. One is, how useful is this? And two is, like, how satisfying as a mathematician would it be? And the fact that there's this application towards proving that software has certain properties or hardware has certain properties, like, if that works, that would obviously be very useful.
1:32:14But from a pure, like, are we going to figure out mathematics? Right. Yeah. Is there is your sense that there's something about finding that one construction cross maps to another construction in a different domain or finding that, oh, this like lemma is if you reconfigure it, like if you redefine this this term, it's still like kind of satisfies what I meant by this term. but a counter example that previously knocked it down no longer applies. Like that kind of dialectical thing that happens in mathematics. Will the software like replace that? Yeah, and like how much of the value of this sort of pure mathematics just comes from actually just coming up with entirely new ways of thinking about a problem?
1:32:58Yeah. Like mapping it to a totally different representation and yeah, do we have examples of? I don't know. I think of it as, I think of it maybe a little bit like the, when everybody had to write assembly code or something like that. just like the amount of fun, like cool startups that got created. It was like a lot less or something. Right. And so it was just like less people could do it. Progress was more grinding and slow and lonely and so on. You had more false failures because you didn't get something about the assembly code. Right. Rather than the essential thing of like, it was your concept.
1:33:28Right.
1:33:31Harder to collaborate and stuff like that. And so I think it will like be really good. there is some worry that by not learning to do the mechanical parts of the proof that you fail to generate the intuitions that inform the more conceptual parts the creative part right yeah it's the same with assembly and right yeah and and so so at what point is that applying is vibe coding are people not learning computer science right or actually are they like vibe coding and they're also simultaneously looking at at the lom with like explaining them these abstract computer science concepts and it's all just like all happening faster their feedback loop is faster and they're learning way more abstract computer science and algorithm stuff because they're vibe coding.
1:34:07You know, I don't know. It's not obvious. That might be something, the user interface and the human infrastructure around it. But I guess there's some worry that people don't learn the mechanics and therefore don't build like the grounded intuitions or something. But my hunch is it's like super positive. Exactly on net how useful that will be or how much overall math like breakthroughs or like math breakthroughs even that we care about will happen. I don't know. I mean, one other thing that I think is cool is actually the accessibility question. It's like, okay, that sounds a little bit corny.
1:34:39Okay, yeah, more people can do math, but who cares? But I think there's actually lots of people that could have interesting ideas, like maybe the quantum theory of gravity or something.
1:34:51Like, yeah, one of us will come up with the quantum theory of gravity instead of a card-carrying physicist. In the same way that Steve Burns is reading the neuroscience literature, and he hasn't been in a neuroscience lab that much. but he's like able to synthesize across the neuroscience literature oh learning subsystem, steering subsystem does this all make sense? He's kind of like he's an outsider neuroscientist in some ways can you have outsider string theorists or something because the math is just done for them by the computer and does that lead to more innovation in the string theory? Right?
1:35:21Maybe yes Interesting So if this approach works and you're right that LLMs are not the final paradigm and suppose it takes at least 10 years to the final paradigm. Yeah. In that world, there's this fun sci-fi premise where you have, it turns to how today had a tweet where he's like, these models are like automated cleverness but not automated intelligence. And you can quibble with the definitions there. But yeah, if you have automated cleverness and you have some way of filtering, which if you can formalize and prove things that the LLMs are saying you could do, then you could have this situation where quantity has a quality all of its own.
1:36:06And so what are the domains of the world which could be put in this provable symbolic representation? And furthermore, okay, so in the world where AGI is super far away, maybe it makes sense to like literally turn everything the LLMs ever do or almost everything they do into like super provable statements. And so LLMs can actually build on top of each other because everything to do is like super provable. Yeah. Maybe this is like just necessary because you have billions of intelligences running around, even if they are super intelligent. The only way the future HCI civilization can collaborate with each other is if they can prove each step.
1:36:41Yeah, yeah. And they're just like brute force churning out. This is what the Jupiter brains are doing. It's a universal language. It's provable. And it's also provable from like, are you trying to exploit me? Are you sending me some message that's actually trying to like sort of hack into my brain effectively? Are you trying to socially influence me? Are you actually just like sending me just the information that I need and no more for this? And yeah, so Davidad, who's like this program director at ARIA now in the UK, I mean, he has this whole design of a kind of ARPA style program, a sort of safeguarded AI that very heavily leverages like provable safety properties.
1:37:16And can you apply proofs to like, can you have a world model? But that world model is actually not specified just in neuron activations, but it's specified in you know equations those might be very complex equations but if you can just get insanely good at just auto proving these things with cleverness auto cleverness can you have you know explicitly interpretable world models you know um as opposed to neural net world models and like move back basically to symbolic methods just because you can you can just have insane amount of ability to prove things yeah i mean that's an interesting vision i don't know how you know in the next 10 years, like whether that will be the vision that plays out.
1:37:53But I think it's really interesting to think about. Yeah. And even for math, I mean, I think Terry Tao is like doing some amount of stuff where it's like, it's not about whether you can prove the individual theorems. It's like, let's prove all the theorems on mass. And then it was like, study the properties of like the aggregate set of proved theorems, right? Which are the ones that got proved and which are the ones that didn't. Okay. Well, that's like the landscape of all the theorems instead of one theorem at a time, right? Speaking of symbolic representations, one question I was meaning to ask you is, how does the brain represent the world model?
1:38:25Like, obviously, that's out in neurons, but I don't mean sort of extremely functionally. I mean, sort of conceptually, is it in something that's analogous to the hidden state of a neural network, or is it something that's closer to a symbolic language? We don't know. I mean, I think there's some amount of study of this. I mean, there's these things like, you know, face patch neurons that represent certain parts of the face that geometrically combine in interesting ways. That's sort of with geometry and vision. is that true for like other more abstract things there's like this idea of cognitive maps like a lot of the stuff that a rodent hippocampus has to learn is like place cells and like where is the rodent going to go next and is it going to get a reward there um is like very geometric and like do we organize concepts with like a abstract version of a spatial map um there's some questions of can we do like true symbolic operations like can i have like a register in my brain that copies a variable to another register, regardless of what the content of that variable is.
1:39:23That's like this variable binding problem. And basically, I just don't. I don't know if we have that machinery or if it's more like cost functions and architectures that make some of that approximately emerge, but maybe it would also emerge in a neural net. There's a bunch of interesting neuroscience research trying to study this, what the representations look like. What was your hunch? Yeah. My hunch is it's going to be a huge mess, and we should look at the architectures, the loss functions and the learning rules and we shouldn't really i don't expect it to be pretty in there yeah which is that is not a symbolic language yeah probably it probably it's not that symbolic yeah but but but other people think very differently you know yeah another random questions speaking of binding yeah but what is up with feeling like there's an experience that it's like both all the parts of your brain which are modeling very different things have different drives feel like at least presumably feel like there's an experience happening right now and also that across time you feel like what is uh yeah i'm pretty much at a loss on this one um i don't know i mean max hodak has been giving talks about this recently he's another really hardcore neuroscience person um neurotechnology person um and the thing i mentioned with dorso um it maybe also it sounds like it might have some touching on this question but But yeah, I don't think anybody has any idea.
1:40:49It might even involve new physics. It's like, you know, yeah. Another question, which might not have an answer yet. So continual learning, is that the product of something extremely fundamental at the level of even the learning algorithm where you could say, look, at least the way we do backprop in neural networks is that you freeze the way, there's a training period and you freeze the weights um and so you just need this active inference or some other learning rule uh in order to do continual learning or do you think it's more a matter of architecture and how is memory exactly stored and is it like what kind of associative memory you have basically yeah so continual learning um i don't know i think that there's probably things that there's probably some at the architectural level there's probably something interesting stuff that the hippocampus is doing um and people have long thought this um what kinds of sequences is storing how is it organizing representing that how is it replaying it back what is it replaying back um how is it exactly how that memory consolidation works in i was sort of training the cortex using replays or or or memories from the hippocampus or something like that.
1:42:08There's probably some of that stuff. There might be multiple timescales of plasticity or sort of clever learning rules that can kind of, I don't know, can sort of simultaneously kind of be storing sort of short-term information and also doing backprop with it. I mean, neurons might be doing a couple of things, you know, some fast weight plasticity and some slower plasticity at the same time or synapses that have many states. I mean, I don't know. I mean, I think that from a neuroscience perspective, I'm not sure that I've seen something that's super clear on what continual learning, what causes it except maybe to say that this this systems consolidation idea of sort of hippocampus consolidating the cortex like some people think is a big piece of this and we don't still fully understand the details yeah speaking of fast weights is there something in the brain which is the equivalent of this distinction between parameters and activations that we see in neural networks and specifically like in transformers we have this uh idea like some of the activations are the key and value vectors of previous tokens that you build up over time.
1:43:12And there's like the so-called the fast weights that you, whenever you have a new token, you query them against these activations, but you also obviously query them against all the other parameters in the network, which are part of the actual built-in weights. Is there some such distinction that's analogous? I don't know. I mean, we definitely have weights and activations. whether you can use the activations in these clever ways um different forms of like actual attention like attention in the brain um is that based on i'm trying to pay attention i think there's several probably several different kinds of like actual attention in the brain i want to pay attention to this area of visual cortex i want to pay attention to this the content in other areas that is triggered by the content in this area right attention that's just based and kind of reflexes and stuff like that.
1:43:59So I don't know. I mean, I think that there's not just the cortex, there's also the thalamus. The thalamus is also involved in kind of somehow relaying or gating information. There's cortical-cortical connections. There's also some amount of connection between cortical areas that goes through the thalamus. Is it possible that this is doing some sort of matching or kind of constraint satisfaction or matching across keys over here and values over there? Is it possible that it can do stuff like that? maybe, I don't know. This is all part of what's the architecture of this corticothalamic system. I don't know how transformer-like it is or if there's anything analogous to that attention.
1:44:39It'll be interesting to find out. We've got to give you a billion dollars so we can come on the podcast again and tell me how exactly the room works. Mostly I just do data collection. It's like really unbiased data collection so all the other people can figure out these questions. Maybe the final question to go off on. is what was the most interesting thing you learned from the gap map? And maybe you want to explain what the gap map is. So the gap map, so in the process of incubating and coming up with these focused research organizations, these sort of nonprofit startup-like moonshots that we've been getting philanthropists and now government agencies to fund, we talked to a lot of scientists.
1:45:20And some of the scientists were just like, here's the next thing my graduate student will do. Here's what I find interesting. exploring these really interesting hypothesis spaces like all the types of things we've been talking about and some of them are like here's this gap um i need this piece of infrastructure which like there's no combination of the grad students in my lab or me loosely collaborating with other labs with traditional grants that could ever get me that i need to have like an organized engineering team that like builds you know the the mini miniature equivalent of the hubble space telescope and if i can build that hubble space telescope then like i will unblock all the other researchers in my field or some like path of technological progress in the way that the Hubble Space Telescope made lifted the boats, improved the life of every astronomer, but wasn't really an astronomy discovery in itself.
1:46:03It was just like you had to put this giant mirror in space with a CCD camera and like organize all the people and engineering and stuff to do that. So some of the things we talked to scientists about look like that. And so the gap map is basically just like a list of a lot of those things. And it's like we call it a gap map. I think it's actually more like a fundamental capabilities map. Like what are all these things like mini Hubble Space Telescopes. And then we kind of organized that into gaps for helping people understand that or search that. And what was the most surprising thing you found?
1:46:35So, I mean, I think I've talked about this before, but I think one thing is just like kind of like the overall size or shape of it or something like that. It's like a few hundred fundamental capabilities. So if each of these was like a deep tech startup size project that's like only a few billion dollars or something, like you know each one of those was a series a that's only like not you know it's not like a trillion dollars to solve these gaps it's like lower than that and so that's that's like one maybe we assumed that and we also came to that's what we got it's not really comprehensive it's really just a way of summarizing a lot of conversations we've had with scientists um i do think that in the aggregate process like things like lean are actually like surprising because i did start from sort of neuroscience and biology it was like very obvious that there's sort of like these omics.
1:47:20We need genomics, but we also need connectomics. And, you know, we can engineer E. coli, but we also need to engineer the other cells. And like there's like somewhat obvious parts of biological infrastructure. I did not realize that like math proving infrastructure like was a thing. And so, and that was kind of like emergent from trying to do this. So I'm looking forward to seeing other things where it's like not actually this like hard intellectual problem to solve it. It's maybe the kind of slightly the equivalent of AI researchers just needed GPUs or something like that. focus and and and really good pytorch code to like start doing this like what is the full diversity of fields in which that exists um we've even now found and which are the fields that do or don't need that so fields that have had gazillions of dollars of investment do they still need some of those do they still have some of those gaps or is it only more like neglected fields um We're even finding some interesting ones in actual astronomy, actual telescopes that have not been explored, maybe because of the kind of, if you're getting above a critical mass size project, then you have to have like a really big project and that's a more bureaucratic process with the federal agencies.
1:48:33Yes, I guess you just kind of need scale in every single domain of science these days. Yeah, I think you need scale in many of the domains of science. And that does not mean that the low scale work is not important. It does not mean that kind of creativity, serendipity, etc. Each student pursuing a totally different direction or thesis that you see in universities is not like also really key. But yeah, I think we need some amount of scalable infrastructure is missing in essentially every area of science, even math, which is crazy because mathematicians I thought just needed whiteboards. but they actually need lean they actually need verifiable programming languages and stuff I didn't know that cool Adam this is super fun thanks for coming on thank you so much where can people find your stuff a pleasure the easiest way now my adamarbelson.org website is currently down I guess but you can find convergentresearch.org can link to a lot of the stuff we've been doing and then you have a great blog longitudinal science yes longitudinal science yes on wordpress cool thank you so much pleasure hey everybody I hope you enjoyed that episode if you did the most helpful thing you can do is just share it with other people who you think might enjoy it.
1:49:36It's also helpful if you leave a rating or a comment on whatever platform you're listening on. If you're interested in sponsoring the podcast, you can reach out at dhwarkesh.com slash advertise. Otherwise, I'll see you on the next one.
From the publisher
Adam Marblestone has worked on brain-computer interfaces, quantum computing, formal mathematics, nanotech, and AI research. And he thinks AI is missing something fundamental about the brain.
Why are humans so much more sample efficient than AIs? How is the brain able to encode desires for things evolution has never seen before (and therefore could not have hard-wired into the genome)? What do human loss functions actually look like?
Adam walks me through some potential answers to these questions as we discuss what human learning can tell us about the future of AI.
Watch on YouTube; read the transcript.
Sponsors
* Gemini 3 Pro recently helped me run an experiment to test multi-agent scaling: basically, if you have a fixed budget of compute, what is the optimal way to split it up across agents? Gemini was my colleague throughout the process — honestly, I couldn’t have investigated this question without it. Try Gemini 3 Pro today gemini.google.com
* Labelbox helps you train agents to do economically-valuable, real-world tasks. Labelbox’s network of subject-matter experts ensures you get hyper-realistic RL environments, and their custom tooling lets you generate the highest-quality training data possible from those environments. Learn more at labelbox.com/dwarkesh
To sponsor a future episode, visit dwarkesh.com/advertise.
Timestamps
(00:00:00) – The brain’s secret sauce is the reward functions, not the architecture
(00:22:20) – Amortized inference and what the genome actually stores
(00:42:42) – Model-based vs model-free RL in the brain
(00:50:31) – Is biological hardware a limitation or an advantage?
(01:03:59) – Why a map of the human brain is important
(01:23:28) – What value will automating math have?
(01:38:18) – Architecture of the brain
Further reading
Intro to Brain-Like-AGI Safety - Steven Byrnes’s theory of the learning vs steering subsystem; referenced throughout the episode.
A Brief History of Intelligence - Great book by Max Bennett on connections between neuroscience and AI
Adam’s blog, and Convergent Research’s blog on essential technologies.
A Tutorial on Energy-Based Learning by Yann LeCun
What Does It Mean to Understand a Neural Network? - Kording & Lillicrap
E11 Bio and their brain connectomics approach
Sam Gershman on what dopamine is doing in the brain
Gwern’s proposal on training models on the brain’s hidden states
Relevant episodes: Ilya Sutskever, Richard Sutton, Andrej Karpathy
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe




