In short
Dwarkesh Podcast Episode Summary: Sholto Douglas & Trenton Bricken - How to Build & Understand GPT-7's Mind
Episode Overview In this engaging episode of the Dwarkesh Podcast, host Dwarkesh Patel interviews AI researchers Sholto Douglas and Trenton Bricken. The conversation revolves around Large Language Models (LLMs), particularly focusing on the training and capabilities of the anticipated GPT-7. The discussion covers a variety of topics including the intricacies of model training, the future of AI capabilities, and what goes on inside these complex systems.
Key Themes and Concepts
- LLM Training and Capabilities
- Long Contexts:
- Discussed the importance of long context windows and how they enhance the intelligence of LLMs, enabling them to process and integrate vast amounts of information.
- Emphasized that models can learn and adapt better when given extensive context.
- Intelligence as Associations:
- Intelligence is seen as the ability to make associations and connections between concepts.
- The idea of an "intelligence explosion" was explored, suggesting that once models achieve a certain threshold, their capabilities could grow exponentially.
- Internal Mechanisms of LLMs
- Superposition and Feature Spaces:
- Talked about how LLMs utilize superposition to manage a diverse set of features, allowing them to represent complex concepts.
- Highlighted the idea that features can be infinitely split, leading to deeper and more granular understandings of the data.
- Agents and Reasoning:
- Explored the difference between simple associations and more complex reasoning, raising questions about the true nature of AI reasoning capabilities.
- Mentioned the challenge of understanding how models perform reasoning tasks and whether current methods accurately capture these processes.
- Historical Perspective and Future Directions
- Personal Journeys in AI Research:
- Both guests shared their personal paths into AI research, emphasizing the importance of agency and determination in their careers.
- Discussed how serendipitous connections and personal projects can lead to significant opportunities in the field.
- Anticipating GPT-7:
- Speculated on the potential capabilities of GPT-7, including advanced reasoning and decision-making processes.
- Highlighted the importance of interpretability in understanding the behaviors of future models.
Notable Discussions
- Redundancy in Models:
- Discussed how redundancy can lead to challenges in identifying specific features and understanding the model's behavior.
- Explored whether models will be able to generalize across tasks and domains.
- Moral and Ethical Considerations in AI:
- Raised questions regarding the control and alignment of future AI models, especially regarding potential misuse and the ethical implications of their capabilities.
- Emphasized the need for transparent practices and community involvement in shaping AI development.
- Interpretability and Safety
- Understanding Model Behavior:
- Emphasized the need for developing robust interpretability methods that can help understand and predict model behavior in various contexts.
- Discussed the potential for future models to be evaluated not just on performance but also on their interpretability and alignment with human values.
- Future of AI and Research Directions
- Collaborative Research:
- The importance of interdisciplinary collaboration and the role of community in advancing AI research was highlighted.
- Speculated on future trends in AI development and the necessary steps to ensure responsible and beneficial advancements.
Conclusion The episode concludes with a sense of optimism about the future of AI and the potential for advancements in understanding and developing LLMs. The insights shared by Sholto Douglas and Trenton Bricken provide a deep dive into the mechanics of AI, setting the stage for continued exploration and innovation in the field.
---
Links & Resources
- Follow Trenton Bricken on Twitter: [@TrentonBricken](https://twitter.com/TrentonBricken)
- Follow Sholto Douglas on Twitter: [@_sholtodouglas](https://twitter.com/_sholtodouglas)
- Watch the episode on [YouTube](https://www.youtube.com/DwarkeshPatel).
- Listen on [Apple Podcasts](https://podcasts.apple.com/us/podcast/sholto-douglas-trenton-bricken-how-to-build-understand/id1516093381?i=1000650748087) or [Spotify](https://open.spotify.com/episode/2dtDauiE4v8ldNRqPFq0uP?si=7S4n69QuTjeYz0lZwW4xIw).
---
This summary encapsulates the core discussions and insights from the podcast episode, making it accessible for readers interested in the future of AI and language models.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Okay, today I have the pleasure to talk with two of my good friends, Shulto and Trenton. Shulto. Shulto, make stuff. We did it. I was gonna say I... Let's do this in reverse. I'll hold my stuff with my good friends. Yeah, I should have had more point of the contact slide. Wow. Shit. Anyways. Anyways, Shulto, Noah Brown. Noah Brown, the guy who wrote the diplomacy paper, he said this about Shulto. He said he's only been in the field for 1 .5 years, but people in AI know that he was one of the most important people behind Gemini success. And Trenton, who's an anthropic, works on mechanistic interpretability, and it was widely reported that he has solved alignment.
0:58So this will be a capabilities only podcast, alignment is already solved, so no need to discuss further. Okay, so let's start by talking about context links. It seemed to be underhyped to give and how important it seems to me to be that you can just put a million tokens into context. There's apparently some other news that got pushed to the front for some reason. But yeah, tell me about how you see the future of long -fine text links and what that implies with these models. Yeah. So I think it's really underhyped because until I started working on it, I didn't really appreciate how much of a step up in intelligence it was for the model to have the onboarding problem basically instantly solved.
1:36And you can see that a little bit in the public graphs in the paper where just throwing millions of tokens worth of context about a codebase allows it to be counter -matically better at predicting the next token in a way that you'd normally associate with huge increments in model scale. But you don't need that. All you need is a new context. So, underhyped and buried by some other news. In context, are they as sample efficient and smart as humans? I think that's really worth exploring. For example, one of the evals that we did in the paper has it learning a language in context better than a human expert could learn that new language over the course of a couple of months.
2:14This is only a pretty small demonstration, But I'd be really interested to see things like Atari games or something like that where you like throw in a couple hundred like well thousand frames Layable actions and then in the same way that you'd like show your friend how to play a game Right and see if it's able to reason through it might at the moment You know with the infrastructure stuff it's still a little bit slower like yeah, I mean that but I would actually I would guess that might just work out of the box in a way that would be pretty mind -blowing and crucially I think this language was esoteric enough that it wasn't in the training day right exactly Yeah, if you look at the model before, it has that context thrown in.
2:45It just doesn't have the language at all, and it can't get any translation. And this is like an actual human language. And I just, yeah, exactly, an actual human language. So if this is true, it seems to me that these models are already an important sense, superhuman, not in the sense that they're smarter than us, but I can't keep a million tokens. In my context, when I'm trying to solve a problem, remembering and integrating all the information entire code base. Am I wrong in thinking this is like a huge unlock? Actually, generally, I think that's true. Like previously, I've been frustrated when my models aren't as smart.
3:18Like, you asked them a question and you want it to be smarter than you or to know things that you don't. And this allows them to know things that you don't in a way that it just ingest huge amount of information at what you just can't. So, yeah, it's extremely important. How do we explain in context learning? Yeah. So there's a piece of, there's a lot of work I quite like where it looks at in context learning as basically very similar to gradient descent, but the attention operation can be viewed as gradient descent on the in context data. That paper had some cool plot, so basically we take N steps at gradient descent, and that looks like N layers of in context learning.
3:53It looks very similar. So I think that's one way of viewing it and trying to understand what's going on. You can ignore it if we're about to say, because given the introduction, and alignment is solved, I say it isn't a problem. But I think the context stuff does get problematic, but also interesting here. I think there'll be more work coming out in the not -too -distant future, and around what happens if you give 100 shot prompt for Joe Breaks, under sorry, all attacks. It's also interesting in the sense of, if your model is doing gradient descent and learning on the fly, even if it's been trained to be harmless, you're dealing with a totally new model in a way.
4:36You're like fine tuning, but in a way where you can't control what's going on. Can you explain what do you mean by gradient descent is happening in the forward pass and attention? Yeah, there was something in the paper about trying to teach the model to do linear regression. Right, but just through the number of samples they gave in the context. And you can see if you plot on the x -axis number of shots that it has, or examples, and then the loss it gets on just like ordinary least squares regression. that will go down one time. And it goes down exactly matched with a number of gradient descent steps.
5:07Yeah, exactly. Okay. I only read the interim discussion section of that paper, but in the discussion, the way they framed it is that in order to get better at long context tasks, the model has to get better at learning to learn from these examples or from the context that is already within the window. And the implication of that is is the model, if like meta learning happens because it has to learn how to give it our long context tasks, then in some important sense, the task of intelligence is like requires long context examples and long context training. Like meta learning, like to induce meta learning.
5:46Right, understanding how to better induce meta learning, your pre -training process is like a very important thing to actually about flexible or adaptive intelligence. Right, but you can proxy for that just by getting better at doing long context tasks. One of the bottlenecks for AI progress that many people identify is the inability of these models to perform tasks on long horizons, which means engaging with the task for many hours or even many weeks or months where like if I have, I don't know, an assistant or an employee or something they can just do a thing and tell them for a while. And AI agents haven't taken off for this reason and from what I understand.
6:22So how linked are long context windows and the ability to perform well on them and the ability to do these kinds of long horizon tasks that require you to engage with an assignment for many hours or these unrelated concepts? I mean, I would actually take issue with the that being the reason that agents haven't taken off. Where I think that's more about like nine's of reliability and the model actually successfully doing things. And if you just can't chain tasks successfully with high enough probability, then you won't get something that looks like an agent. And that's why something like an agent might fall in more of a step function in sort of video like GBD4 class models, general class models, they're not enough.
6:58But maybe the next increment on model scale means that you get that extra nine, even though the losses are going down that dramatically, that small amount of extra ability gives you the extra nine. And obviously you need some amount of context to fit long horizon tasks, but I don't think that's been the limiting factor up to them. Yeah, the Nurebs best paper this year by Ryland Schaeffer was the lead author points to this as like the emergence Up Mirage where people will have a task and you get the right or wrong answer depending on if you've sampled the last five tokens correctly And so naturally that's you're multiplying the probability of sampling all of those And if you don't have enough nines for liability then you're not gonna get emergence and all of a sudden you do and it's like, oh my gosh, this ability is emergent.
7:44When actually it was kind of almost there to begin with. And there are always that you can find like a smooth metric full of that. Yeah, human e -vowler, whatever, the GPD for paper, the coding problems, they measure in the cost rate. Exactly, yeah. For the audience, the context on this is, it's basically the idea is you wanna, when you're measuring how much progress there has been on a specific task like solving coding problems, You, you up weighted when it gets it right only one in a thousand times. You don't like give it a one in a thousand score because it's like, oh, like, got to write some of the time.
8:15And so the curve you see is like, it gets it right one in a thousand, then one in a hundred, then one in ten, and so forth. So actually, I want to follow up on this. So if your claim is that the AI agents haven't taken off because of reliability rather than long horizon task performance, isn't the lack of reliability when a task is changed on top of another task, on top of another task, is not exactly the difficulty with long horizon tasks, is that you have to do 10 things in a row or 100 things in a row, and diminishing the reliability of any one of them. Or the other probability goes down from 99 .99 to 99 .9, then the whole thing gets multiplied together and the whole thing becomes much less likely to happen.
8:58That is exactly the problem, but the key issue you're pointing out there is that your base past tasks all freight is 90%. And if it was 90 % then chain, doesn't become a problem. But also, it's like a second. Yeah, exactly. And I think this is also something that just hasn't been properly studied enough. If you look at all of the evals that are commonly like the academic evals, a single problem, right? Like the math problem. It's like one typical math problem, or MMOU. It's like one university level, like from across different topics. You were beginning to start to see evals looking at this properly by a more complex task, like SweetVenche, where they take a whole bunch of GitHub issues.
9:33And that is like a reasonably long horizon task. But it's still not a multi, it's like a multi, sub -hour as opposed to like multi -hour or multi -day task. And so I think one of the things that will be really important to do over the next however long is understand better what does success rate over long rise in task or effect. And I think that's even important to understand what the economic impact these models might be and like actually properly judge increase in capabilities. Right. But like cutting down the tasks that we do and the inputs and outputs involved into minutes or hours or days and seeing how good it is, successively, like, chaining and completing ties of those different resolutions of time.
10:10But then that tells you the cowl, automated, well, a job family, or task family is, in a way that, like, MMO use school is doing. I mean, it was less than a year ago that we introduced 100K context windows, and I think everyone was pretty surprised by that. So, yeah, everyone would just kind of have this sound bite of quadratic attention costs. So, yeah, we can't have long context windows. Here we are. So, yeah, like, the benchmarks are being actively made. Wait, wait, so it doesn't the fact that there's these companies, Google and I don't know magic, maybe others who have million token attention and apply that the could drive, you shouldn't say anything.
10:45But doesn't that like imply that it's not quadratic anymore? Are they just eating the cost? Well, like who knows what Google is doing for its long -context? Yeah, I'm not saying that. I'm not saying that. It's either. One of the things frustrated me about in like the general research fields approach to attention is that there's an important way in which the quadratic cost of attention is actually dominated in typical dense transformers by the MLP block. Right. So you have this n squared term that's associated with attention, but you also an n squared term that's associated with the D model, the residual string dimension of the model.
11:19And if you look, I think Sasha Rush has a great tweet where he looks like basically plots the curve of the cost of attention, respectively, like the cost of really large models, and attention actually trails off. And you actually need to be doing pretty long contexts before that term becomes really important. And the second thing is that people often talk about how attention at inference time is such a huge cost. And if you think about when you're actually generating tokens, the operation is not n -square. It is 1 -q, like one set of q vectors, looks up a whole bunch of kv vectors, and that's linear with respect to the amount of context that the model has.
12:00And so I think this drives a lot of the recurrence and state space, research for people of this meme of, oh, like linear attention and all this stuff. And as Trenton said, there's like a graveyard of ideas around attention. I'm not the thing I don't think it's worth exploring, but I think it's important to consider where the actual strengths and weaknesses of it are. Okay, so what do you make of this take? As we move forward through the takeoff, more and more of the learning happens in the forward pass. So originally all the learning happens in the backward, during this bottom up sort of hill climbing evolutionary process.
12:36If you think in the limit during the intelligence explosion, the AI is maybe handwriting the way it's or doing go -fire or something. And we're in the middle step where a lot of learning happens in context now with these models. A lot of it happens within the backward process. Does this seem like a meaningful gradient along which progress is happening? Like how much? Because the broader thing being the if you're learning in the forward pass is like much more sample -efficient because you can kind of like basically think as you're learning like when humans when you read a textbook you're not just skimming it and trying to absorb what you know what inductive these words follow these words you like read it and you think about it and then Does this seem like a sensible way to think about the progress?
13:21Yeah, there may just be one of the ways in which like, you know, birds and planes like fly, but they fly differently. And like, the virtue of technology allows us to do that like, I just see the accomplished things that birds can't. It might be that context like the similar in that it allows it to work in memory that we can't. But functionally is not like the key thing towards actual reasoning. The key step between GPD2 and GPD3 was that all of a sudden, like there was this metal learning behavior that was observed in the pre -training of the model. And that's, as you said, like something to do with you give it some amount of context, it's able to adapt to that context, and that was a behavior that wasn't really observed before that, at all.
14:02And maybe that's a mixture of cooperative contexts and scale and this kind of stuff. There would never occurred to model tiny contexts that was there. The discussion is interesting point. So when we talk about scaling up these models, How much of it comes from just making the models themselves bigger and how much comes from the fact that during any single call You are using more compute so if you think of diffusion you can just iteratively keep adding more compute and If that computer solved you can keep doing that and In this case if there's a quadratic penalty for attention, but you're doing long -context anyways then you're still Dumping in more compute during that during training or not during having bigger models, but just like yeah Yeah, it's interesting because you do get more forward passes by having more tokens, right?
14:50My one gripe I guess I have two gripes with this though. Maybe three so one like In the old paper One of the transformer modules they have a few and the architecture is like very intricate But they do I think five forwards passes through it and will gradually like refine their solution as a result You can also kind of think of the residual stream I mean, a shelter alluded to kind of the breed right operations as like a poor man's adaptive compute Where it's like I'm just gonna give you all these layers and like if you want to use them great if you don't then that's also fine Um, and then people will be like oh well the brain is is recurrent and you can like do however many loops through it You want and I think to a certain extent that's right right like if I ask you a hard question You'll spend more time thinking about it and that would correspond to more forward passes, but um I think there's a finite number of forward passes that you can do It's kind of with language as well.
15:39People are like, oh, well, human language can have like infinite recursion in it, like infinite nested statements of like, the boy jumped over the bear that was doing this, that had done this, that had done that. But like empirically, you'll only see five to seven levels of recursion, which kind of relates to whatever that magic number of like how many things you can hold in working memory at any given time is. And so, yeah, it's not infinitely recursive, but does that matter in the regime of human intelligence? And can you not just add more layers? Breakdown for me, you're referring to this in some of your previous answers of, listen, you have these long contexts and you can hold more things in memory, but ultimately comes down to your ability to mix concepts together to do some kind of reasoning.
16:26And these noddles aren't necessarily a human level at that, even in context. Breakdown for me, how you see storing just raw information versus reasoning and what's in between. Like, where's the reasoning happening? Is that, where's just like storing where information happening? What's different between them in these models? Yeah, I don't have a super crisp answer for you here. I mean, obviously with the input and output of the model, you're mapping back to actual tokens, right? And then in between that, you're doing higher level processing. Before we get deeper into this, we should explain to the audience you referred earlier to inthropics way of thinking about transformers as these read -write operations that layers do.
17:10One of you should just kind of explain at a high level what you mean by that. So the residual stream, imagine you're in a boat going down a river and the boat is kind of the current query where you're trying to predict the next token. So it's the cat sat on the blank and then you have these little like streams that are coming off the river where you can get extra passengers or collect extra information if you want. And those correspond to the attention heads and MLPs that are part of the model. I was going to almost give a question about the working memory of the model, like the RAM of the computer.
17:46We are like choosing what information to read, exactly. So you can do something with it and then maybe you'd read something else in later on. And you can operate on subspaces of that high -dimensional vector. or a ton of things, or I mean at this point, I think it's almost given that like are encoded in superposition, right? So it's like, yeah, the residual stream is just one high -dimensional vector, but actually there's a ton of different vectors that are packed into it. Yeah, I might just like dumb it down, like as the way that would have made sense to me a few months ago, of, okay, so you have, you know, whatever words are in the input you put into the model, all those words get converted into these tokens and those tokens get converted into these vectors.
18:28And basically, it's just like the small amount of information that's moving through the model. And the way you explained it to me, Shulta, and this paper talks about is, early on in the model, maybe it's just doing some very basic things about, like, what do these tokens mean? Like if it says 10 plus 5, just like moving information about to have that - The good representation. Exactly, just represent. And in the middle, maybe the deeper thinking is happening about how to solve this. At the end, you're converting it back into the output token. Because the end product is you're trying to predict the probability of the next token from the last of those residuals of streams.
19:06And so yeah, it's interesting to think about the small compressed amount of information moving through the model and it's getting modified in different ways. Trenton, so it's interesting. You're one of the few people who have background from neuroscience. So you can think about the analogies here to the brain. And in fact, one of our friends, though, he had a paper in grad school about thinking about attention in the brain. And he said, this is the only or first, like, neuro -explanation of why attention works. Whereas we have evidence from why the CNN's work, the convolutional neural network networks work based on the visual cortex or something.
19:47Yeah, I'm curious how do you think in the brain there's something like a residual stream of this compressed amount information that's moving through and it's getting modified as you're thinking about something. Even if that's not what list literally happening, do you think that's a good metaphor for what's happening in the brain? Yeah, yeah. So at least in the Sardarvala, you basically do have a residual stream where the whole what we'll call the attention model for now and I can go into whatever amount of detail you want for that. You have inputs that route through it, but they'll also just go directly to the end point that that module will contribute to.
20:24So there's a direct path and an indirect path. And so the model can pick up whatever information it wants and then add that back in. What what happens is the cerebellum. So the cerebellum nominally just does find motor control. But I analogize this to the person who's lost their keys and is just looking under the street light, where it's very easily to observe this behavior. One leading cognitive neuroscientist said to me that a dirty little secret of any FMRI study where you're looking at brain activity for a given task is that the cerebellum is almost always active and lighting up for it. If you have a damaged cerebellum, You also are much more likely to have autism.
21:06So it's associated with like social skills. And one of these particular studies where I think they use PET instead of FMRI, but when you're doing next token prediction, the cerebellum lights up a lot. Also 70 % of your neurons in the brain are in the cerebellum. They're small, but they're there, and they're taking up real metabolic cost. This is one of Gerns' points that, like what changed with humans was not just that we have more neurons, or he shared this article, but specifically, there's more neurons in the cerebral cortex in the cerebellum, and you should say more about this, but like, they're more metabolically expensive and they're more involved in signaling and sending information back and forth.
21:49Yeah. What is his attention? What's going on? Yeah, yeah. So I guess the main thing I want to communicate here, so back in the 1980s, Penteconverba came up with a associated memory algorithm for. I have a bunch of memories, I want to store them. There's some on a noise or corruption that's going on, and I want to query or retrieve the best match. And so he writes this equation for how to do it, and a few years later realizes that if you implemented this as an electrical engineering circuit, it actually looks identical to the core cerebellar circuit. And that circuit and the cerebellum more broadly is not just in us, it's in basically every organism.
22:29There's active debate on whether or not cephalopods have it. They kind of have a different evolutionary trajectory. But even fruit flies with the Drosophila mushroom body. That is the same cerebellar architecture. And so that convergence, and then my paper, which shows that actually this operation is to a very close approximation, the same as the attention operation, including implementing the softmax and having this sort of like nominal quadratic cost that we've been talking about. And so the three -way convergence here and the takeoff and success of Transformers seems pretty striking to me. Yeah, I want to do about an ask.
23:05I think what motivated this discussion in the beginning was we were talking about like wait, what is the reasoning? What is the memory? What do you think about the analogy you found to attention and this? Do you think of this as more just looking up the relevant memories or the relevant facts? And if that is the case, like, where is the reasoning happening in the brain? How do we think about how that builds up into the reasoning? Yeah, so maybe my hot take here, I don't know how hot it is, is that most intelligence is pattern matching. And you can do a lot of really good pattern matching if you have a hierarchy of associated memories.
23:47So you have, you start with your very basic associations between just like objects in the real world. But you can then chain those and have more abstract associations such as like a wedding ring symbolizes like so many other associations that are downstream. And so, and you can even generalize the attention operation and this associated memory as the MLP layer as well. It's in a long term setting where you don't have tokens in your current context. But I think this is an argument that association is all you need. And associate of memory in general as well. So you can do two things with it. You can both denoise or retrieve a current memory.
24:33So if I see your face, but it's raining and cloudy, I can denoise and gradually update my query towards my memory of your face. But I can also access that memory and then the value that I get out actually points to some other totally different part of the space. And so, so a very simple instance of this would be if you learn the alphabet, right? And so I query for A and it returns B, I query for B and it returns C. And you can traverse the whole thing. Yeah. Yeah, one of the things I talked about was here to paper in 2008 that memory and imagination are very linked. because it's the very thing that you mentioned.
25:12Memory is reconstructive, and so you're in some sense imagining every time you're thinking of a memory because you're only storing a condensed version of it and you're like, have to. And it is as famous the UI human memory is terrible and like why people in the witness box or whatever will just make shit up. Okay, so let me ask you a stupid question. So you like reach for lock homes, right? And like the guy is incredibly sample efficient. He'll see a few observations, and he'll basically figure out who come into the crime, because there's a series of deductive steps that leads from somebody's tattoo and what's on the wall to the implications of that.
25:53How does that fit into this picture? Because crucially, what makes them smart is that there's not an association, but there's a sort of deductive connection between in different pieces of information. Would you just explain it as that's just like higher level association? Yeah, I think so. So I think learning these higher level associations to be able to then map patterns to each other is kind of like a meta learning. I think in this case, he would also just have a really long context length. Or really long working memory, where he can have all of these bits and continuously query them as he's coming up with whatever theory.
Read the full transcript
26:30So that the theory is moving through the residual stream. And then he's attention heads are querying his context, but then how he's projecting his query and keys in the space and how his MLPs are then retrieving like longer term facts or modifying that information is allowing him to then in later layers do even more sophisticated queries and slowly be able to reason through and come to a meaningful conclusion. That feels right to me. in terms of like looking back in the past, you were selectively reading in certain piece of information, comparing them maybe that informs your next step of like what piece of information you now need to pull in.
27:08And then you build this representation, which I would like progressively looks closer and closer to like the suspect in your case. Yeah, yeah. That's the little outlandish. Do you know what I mean? The lens on like, yeah, the suspect. Yeah, well, something I think that the people who aren't doing this research can overlook look is after your first layer of the model, every query key and value that you're using for attention comes from the combination of all the previous tokens. So like my first layer, all query my previous tokens and just extract information from them. But all of a sudden, let's say that I attended to tokens 1, 2, and 4 in equal amounts.
27:47Then the vector in my residual stream, assuming that they just, they wrote out the same thing to the value vectors, but ignore that for a second. Is a third of each of those. And so when I'm querying in the future, my query is actually a third of each of those things. And so, but they might be written to different subspaces. That's right. Right, exactly. But they would have to. And so you can recombine and immediately, even by layer two and certainly by the deeper layers, just have like these very rich vectors that are packing in a ton of information. And the causal graph is like literally over every single layer that happened in the past.
28:22And that's what you're operating on. It does bring to mind a very funny eVal to do. It would be like a Sherlock Holmes eVal, it's you put the entire book into context. And then you have like a sentence, which is like the suspect is like X, then you have like a larger probability distribution over like the different characters. Yeah, yeah. And then like, as you put more like. That would be super cool. Yeah, yeah. I wonder if you'd get anything at all. I'm pretty cool. The Sherlock Holmes is probably already in the training data. We're going to get a mystery novel that was written in the... You can get it all on to write it.
28:52We could have just say exclude it. We can. How do you... Well, you need to scrape any discussion of it from Reddit or any other thing, right? Right. It's hard. But that's one of the challenges that goes into things like long -context. Seavows is to get a good one. You need to know that it's not your training data. You're like putting the effort to exclude it. What? So, as you're wondering, there's two different threads I want to follow up on. Let's go to the long -context one and then we'll come back to... this. So in the Gemini 1 .5 paper, the eval that was used was, can it like something with power grms as can it like the need of a haystack, right?
29:30Which yeah, I mean there's like, we don't necessarily just care about its ability to recall one specific fact from the context. I'll sit back and ask the question, the loss function for these models is unsupervised, is you don't have to come up with these bespoke things that you keep out of the trading data. Is there a way you can do a benchmark that's also unsupervised where, I don't know, another LLM is raiding it in some way or something like that? And maybe the answer is like, well, if you could do this reinforcement learning would work, is that you have this unsupervised? Yeah, I mean, I think people have explored that kind of stuff.
30:06For example, anthropocaster, constitutional URL paper, but they take another language model and they point it and say how helpful or harmless was that response and then they get updated and try and improve along the pre -dough frontier of healthfulness and harmfulness. So you can like point language models at each other and create e -vows in this way. It's obviously an imperfect art form at the moment because you get reward function hacking basically and the language what like if you try and match up to what even humans are imperfect here like if you try and match up what humans will say humans you'll typically prefer longer answers which aren't necessarily better other answers and you got the same behavior with models.
30:47On the other thread, going back to the Sherlock Holmes thing, if it's all associations all the way down, this is sort of like naive dinner party question. If I just like match you or you're an American, I'm like, but okay, does that mean we should be less worried about super intelligence? Because there's not this sense in which it's like Sherlock Holmes plus plus, it'll still need to just like find these associations, like humans find associations and like, you know what I mean? It's not just like, it sees a frame of the world and it's like figured out all the laws of physics. So for me, because this is a very legitimate response, right?
31:23It's like, well, artificial general intelligence aren't, if you say humans are generally intelligent, then there are no more capable or competent. I'm just worried that you have that level of general intelligence in Silicon where you can then immediately clone hundreds of thousands of agents and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question, yes, you're right, they would still need to learn associations, but the recursive stuff improvement would still have to be them, like if intelligence is fundamentally about these associations, like the improvement is just getting better at association.
32:01There's not like another thing that's happening, and so then it seems like you might disagree with the intuition and that, well, they can't be that much more powerful if they're just doing associations. Well, I think then you can get into really interesting cases of meta learning. Like when you play a new video game or study a new textbook, you're bringing a whole bunch of skills to the table to form those associations much more quickly. And like, because everything in some way ties back to the physical worlds, I think there are general features that you can pick up and then apply in novel circumstances.
32:34Should we talk about intelligence solution then? I don't know if it's a good little. I'll see you next time with this jumping. I mentioned multiple agents and I'm like, oh, here we go. OK, so the reason I'm interested in discussing this is with you guys in particular is the models we have of the intelligence explosion so far come from economists, which is fine. But I think we can do better because in the model of the intelligence explosion, what happens is you replace the AI researchers. and then there's like a bunch of automated AI researchers who can speed up progress, make more AI researchers, make further progress.
33:11And so I feel like if that's the metric, or that's the mechanism, we should just ask AI researchers about whether they think this is plausible. So let me just ask you, like if I have 1 ,000 Asian show up those or Asian trendings, are they just, do you think that you get an intelligence explosion? Is that, yeah, what does that look like to you? So I think one of the important bounding constraints here is compute. I do think you could dramatically speed up AI research. It seems very clear to me that in the next couple of years we'll have things that can do many of the software engineering tasks that I didn't want to do today basis and therefore dramatically speed up my work and therefore speed up the rate of progress.
33:52At the moment, I think most of the labs are somewhat compute bound in that they're always, there are more experiments you could run and more piece of information that you could gain in the same way that like scientific research on biology is also somewhat experimentally like through what bound, like you need to be able to run and culture the cells in order to get the information. I think that will be at least a short term planning constraint. Obviously, you know, Sam's run a rate of $7 trillion to run. Get ships. And so, like it does seem like there's going to be a lot more compute in future as everyone is heavily ramping.
34:26a lot of things. You know, a lot of videos, stock price that represents the relative, a compute increase. But I think we need a few more nines of reliability in order for it to really be useful and trustworthy. Right now it's like, and just having context links that are super long and it's like very cheap to have. Like if I'm working on our code base, it's really only small modules that I can get to write for me right now. But it's very plausible that within the next few years, or even sooner, it can automate most of my task. The only other thing here that I will note is the research that at least our sub team and interpretability is working on is so early stage that you really have to be able to make sure everything is done correctly in a bug -free way and contextualize the results with everything else in the model.
35:29If something isn't going right, be able to enumerate all of the possible things and then slowly work on those. An example that we've publicly talked about in previous papers is dealing with layer norm. If I'm trying to get an early result or look at the loge effects of the model, so if I activate this feature that we've identified to a really large degree, how does that change the output of the model? Am I using layer norm or not? How is that changing the feature that's being learned? They're, yeah, and that will take even more context or reasoning abilities for the model. So you use the couple of concepts together, and it's not self -evident to me that they're the same, but it seemed like you were using them interchangeably, so I just wanna, like, one was, well, to work on the cloud code base and make more modules based on that, they need more context or something, where it seems like they might already be able to fit in the context.
36:26Do you mean like actual, do you mean like the context window context or like more... Yeah, the context window context. So yeah, it seems like now it might just be able to fit. The thing that's preventing it from making good modules is not... The lack of being able to put the code base in there. I think that will be there soon. Yeah, but like it's not going to be as good as you at like coming up with papers because it can like fit the code base in there. No, but it will speed up a lot of the engineering. In a way that causes an intelligence explosion? No, that accelerates research. But I think these things compound.
36:55So the faster I can do my engineering, the more experiments I can run. And then the more experiments I can run, the faster we can. My work isn't actually accelerating capabilities at all. Right, but it's interpreting the models. But we have a lot more work to do on that. Surprise to the Twitter. Twitter got it. Yeah, I get it from her context. So like when you released your paper there was a lot of talk on Twitter about Alamina's song, guys close the curtains.
37:23Yeah, yeah, no, it keeps me up at night how quickly the models are becoming more capable and like just how poor our understanding still is of what's going on. Yeah, I guess I'm still okay so let's ring through this specifics here. By the time this is happening, we have bigger models that are two to four orders of magnitude bigger, right? Or at least an effective compute are two to four orders of magnitude bigger. And so this idea that, well, you can run experiments fast or something. You're having to re -train that model in this version of the intelligence explosion. The recursive self -improvement is different from what might have been imagined 20 years ago where you just rewrite the code.
38:07you actually have to train a new model, and that's really expensive, not only now, but especially in the future as they keep like making these models orders of 19 and bigger. Doesn't that dampen the possibility of a sort of recursive software for room and type intelligence solution? It's definitely gonna act as a breaking mechanism. Like I agree that the world is like, what we're making today looks very different to what people imagined it would look like 20 years ago. that it's not going to be able to write a send code to be like at least smart, because actually at least a train itself, like the code itself is typically quite simple, typically small and self -contained.
38:48The John Carmack had this natural phrase where it's like the first time in history where you can actually plausibly imagine writing AI with like 10 ,000 lines of code. And that actually does seem plausible when you pair most training code bases down to the limit. But it doesn't take away from the fact that this is something we should really strive to measure and estimate how progress might occur. Like we should be trying very, very hard right now to measure exactly how much of a software engineer's job is automatable and what the trend line looks like and be trying out the hottest to project out those trend lines.
39:20But what I'll do with your software engineers, like you are not writing a React front end and deride to that. So it's like, I don't know how this, like what is concretely happening? Maybe you can walk me through, walk me through like a day in the life of show, So you're working on an experiment or project that's going to make the model quite a bit better. Like what is happening from observation, to experiment, to theory, to writing the code, what is happening? And so I think important to contextualize here is that I have primarily worked on inference so far. So a lot of what I've been doing is just taking or helping guide the pre -training process, such as designing a good model for inference and then making the model and the surrounding system faster.
40:03I've also done some pre -training work around that, but that hasn't been by 100 % focused. But I can still describe what I do when I do that work. I'm even a bit of a sorry, let me interrupt and say. Two tries to work. In Karl Schrohmens, and when he was talking about him on the podcast, he did say that things like improving inference, or even literally helping make better chips or GPUs, that's part of the intelligence explosion. Because obviously if the inference code runs faster, it happens better or faster or whatever. Anyway, sorry to go ahead. Yeah. OK, so what is completely a day look like?
40:35I think the most important part to illustrate is this cycle of coming up with an idea, proving it out at different points in scale,
40:47and interpreting and understanding what goes wrong. And I think most people would be surprised to learn just how much goes into interpreting and understanding what goes wrong. Because the ideas, people have long lists of ideas that you want to try, not every idea that you think should work will work and trying to understand why that is is quite difficult and working at what exactly you need to do to interrogate it. So so much of it is like introspection about what's going on. It's not pumping out thousands and thousands of thousands of lighter code. It's not like the difficulty in coming up with ideas even.
41:19I think many people have a long list of ideas that they want to try, but pairing that down and shop calling under very imperfect information, what the right ideas to explore further is really hard. Tell me more about what do you mean by imperfect information? Are these early experiments? What is the information that you're... So, Demis mentioned this in his podcast and also the GPT -4 paper where you have scaling little increments and you can see in the GPT -4 paper, there are a bunch of dots where they say we can estimate the performance of our final model using all of these dots. and there's a nice curve that flows through them.
41:54And, Dennis mentioned, yeah, that we do this process of scaling up. Concretely, why is that imperfect information? Is you never actually know if the trend will hold? For certain architectures, the trend has held really well. And for certain changes, it's held really well. But that isn't always the case. And things which can help at smaller scales can actually hurt the larger scales. So making guesses based on what the trend lines look like and based on like your intuitive feeling of okay this is actually something that's gonna matter particularly for those ones which help with the small scale.
42:33That's interesting to consider that for every chart you see in a release paper, technical report that shows that smooth curve, there's a graveyard of like first year runs and then it's like flat. There's all these like other lines that go in like different directions. and each of them is like a tail -on. Yeah, it's crazy, both as a grad student and then also here, the number of experiments that you have to run before getting a meaningful result. Tell me, okay, so you, but presumably, it's not just like you run it until it stops and then let's go to the next thing. There's some process by which to interpret the early data and also to look at your, I don't know, I could put a Google Doc in front of you with it.
43:11Pretty sure you could just keep typing for a while on different ideas you have. And there's some bottleneck between that and just like making the models better immediately. Yeah, walk me through like, what is the inference you're making from the first or the steps that makes you have better experiments and better ideas? I think one thing that I didn't fully convey before was that I think a lot of good research comes from working backwards from the actual problems that you want to solve. And there's a couple of like grand problems I sort of was in like making the models better today that you would identify as issues and then like work with them.
43:44Okay, how could I like change it to achieve this? There's also a bunch of, when you scale, you run into things and you wanna like fix behaviors or like issues at scale and that like informs a lot of the research for the next increment and this kind of stuff. So concretely the barrier is a little bit software engineering like often having a code base that's large and sort of capable enough that it can support many people doing research at the same time makes it complex. If you're doing everything by yourself, your iteration pace is going to be much faster. I've heard that like, Alec Radford, for example, like famously did much of the pioneering work at OpenEI.
44:20He mostly works out of like a Jupyter notebook and then like has someone else who writes and productionizes that code for him. I don't know if that's true or not. But that kind of stuff, like actually operating with other people makes it the raises the complexity a lot, because it did not have natural reasons, like familiar to like every software engineer. And then the inherent running and launching those experiments is easy, but there's inherent time slows down, induced by that. So you often want to be parallelizing multiple different streams, because one, you can't be totally focused on one thing necessarily.
44:56You might not have fast enough feedback cycles. And then intuiting what went wrong is actually really hard, working out what, this is in many respects from the team that's on is trying to better understand what is going on inside these models, we have inferences and understanding and like head canon for why certain things work, but it's not an exact science. And so you have to constantly be making guesses about why something might have happened, what experiment might reveal, whether that is or isn't true, and that's probably the most complex part. The performance work by comparatively is easier but harder in other aspects, it's just a lot of low level and like difficult engineering work.
45:36Yeah, I agree with a lot of that. But even on the Interpreterability team, I mean, especially with Chris Ola leading it, there are just so many ideas that we want to test. And it's really just having the engineering skill, but I'll put engineering in quotes because a lot of it is research to like very quickly iterate on an experiment, look at the results, interpret it, try the next thing, communicate them, and then just ruthlessly prioritizing what the highest priority things to do are. And this is really important. Like the ruthless prioritization is something which I think separates a lot of quality research from research that doesn't necessarily succeed as much.
46:15We're in this funny field where so many of our theoretical initial theoretical understanding is broken down, basically. And so you need to have this simplicity bias and ruthless prioritization over what's actually going wrong. And I think that's one of the things that separates the most effective people is they don't necessarily get too attached to solving, using a given solution that they're necessarily familiar with, but rather they attack the problem directly. You see this a lot in like, maybe people come in with a specific academic background, they try and solve problems with that toolbox.
46:51And the best people are people who expand the toolbox dramatically, they're running around and they're taking ideas from reinforcement learning, but also from optimization theory, and also they have a great understanding of systems, and so they know what the sort of constraints that bound the problem are, and they're good engineers, so they can iterate and try ideas fast. Like by far, the best researchers I've seen, they all have the ability to try experiments really, really, really, really fast. And that is that cycle time. And that's smaller scales. Cycle time separates people. I mean, machine learning research is just so empirical.
47:22Yeah. And this is honestly one reason why I think our solutions might end up looking more brain -like than otherwise. It's like even though we wouldn't want to admit it, the whole community is kind of doing like greedy, evolutionary optimization over the landscape of like possible, and architecture is in everything else. It's like no better than evolution. And that's not even necessarily a slight against evolution. That's such an interesting idea. I'm still confused on what will be the bottleneck for these, what would be have to be true of an agent such that it's like sped up your research. So in the Alec Ratford example you gave where he apparently already has the equivalent of like co -pilot for his Jupyter notebook experiments.
48:04Is it just that if he had enough of those, he would be a dramatically faster researcher and so you just need Alec Ratford? So it's like you're not automating the humans, you're just making the most effective researchers who have great taste, more effective and like running the experiments for them and so forth or like that you're still working at the point which the intellectual explosion is happening. You know what I mean? Like is that what you're saying? Right. And if that were like directly true, why can't we scale out current research teams better? For example, is it, I think an interesting question for us, like why, if this work is so valuable, why can't we take hundreds or thousands of people who are, like, they're definitely out there.
48:43And like, scale our organizations better.
48:49It's, I think we are less at the moment. bound by the sheer engineering work of making these things, then we are by compute to run and get signal and taste in terms of what the actual right thing to do it, and not making those difficult inferences on imperfect information. For the Gemini team, because I think for interpretability, we actually really want to keep hiring talented engineers. And I think it's a big bottleneck for us to just keep making a lot of effort. Obviously more people is better, but I do think it's interesting to consider. I think one of the biggest challenges that I've thought a lot about is how do we scale better?
49:35Google is an enormous organization and has 200 ,000 -ish people, right? Like, maybe 80 ,000 or something like that. And one has to imagine if there were ways of scaling out to all those fantastically talented software engineers. This seems like a key advantage that you would want to be able to take advantage of. You wouldn't be able to use, but how do you effectively do that? It's a very complex organizational problem. So, compute and taste, that's interesting to think about because, at least the compute part is not bottleneck and more intelligence. It just bottlenecked on Sam 7 trillion or whatever.
50:12So if I gave you 10x the A -100s to run your experiments, how much more effective a research are you? To be a specialist. To be a specialist. To be a specialist. How much more effective a research are you? I think the Gemini program would probably be like, maybe five times faster, which sometimes will compute something like that. So that's pretty good elasticity of like 0 .5. Yeah, wait, that's insane. Yeah, I think like more compute would just like directly convert into progress. So you have some alloc, some fixed size of compute and some of it goes to inference or some of I guess like and also Like to clients of GCP.
50:53Yep. Some of it goes to huh? Some of it goes to training and there I Guess as a fraction of it some of it goes to running these experiments for the full model Yeah, that's right. The student then the fraction goes experiments be higher given that you would just be like If the bottleneck is research and research is bottleneck by compute. And so one of the strategic decisions that every pre -training team has to make is exactly what amount of compute you allocate to your different training runs. Just like to your research program versus scaling the last best thing that you landed on. And I think they're all trying to arrive at a sort of pre -optimal point here.
51:36One of the reasons why you need to still keep training big models is that you get information there that you don't get otherwise So scale has all these emerging properties Which you want to understand better and if you like are always doing research and never remember what I said before about like You're not sure what's going to like fall off the curve right? Yeah, if you like keep doing research in this regime. Yeah And like keep on getting more and more computer -fishing you may have actually gone through the path that actually eventually scales. You need to constantly be investing in doing big runs, too, at the frontier of what you sort of expect to what.
52:16Okay, so then tell me what it looks like to be in the world where AI has significantly sped up AI research, because from this, it doesn't really sound like the AI's are going off and writing the code from scratch and that's leading to faster output. It sounds like they're really augmenting the top researchers in some way. Yeah, tell me concretely, are they doing the experiment? Are they coming up with the ideas? Are they just evaluating the outputs of the experiments? What's happening? I think there's two walls you need to consider here. One is where AI has meaningfully sped up our ability to make algorithmic progress.
52:46And one is where the output of the AI itself is the thing that's the crucial ingredient towards model capability progress. And specifically what I mean there is synthetic data. Right. And in the first world where it's meaningfully speeding up algorithmic progress, I think a necessary component of that is more compute. And you probably like reach to see the necessity point where like AI's maybe at some point are easier to speed up and get on to context than yourself. Let's just rate that other people. And so AI's meaningfully speed up your work because they're like a fantastic copilot, basically that helps you code multiple times faster.
53:28And that seems like actually quite reasonable. Super long context, Super smart model. It's on board it immediately and you can send them off and to complete sub tasks and sub goals for you. And it actually feels very plausible, but again, we don't know because there are no great evals about that kind of thing. The best one is, as I said before, sweet bench. In that one, somebody was mentioning to me, The problem is that when a human is trying to do a pull request, they'll type something out and they'll run it and see if it works, and if it doesn't, they'll rewrite it. None of this was part of the opportunities that the LLM was given when run on this.
54:08Like, it just like output it, and if it runs and checks all the boxes, then it passed. It might have been an unfair test in that way. So you can imagine that is, like, if you were able to use that, that would be an effective training source for having, like the key thing that's missing from a lot of training data is like the reasoning traces, right? And I think this would be, if I wanted to try and automate a specific field with like job family, or like understand how like at risk of automation that is then having reasoning traces feels to me like a really important part of that. There's so many threads and that I want to follow up on.
54:54Let's begin with the data versus a compute thing of like, is the output of these AI is the thing that's causing the intelligence solution or something. People talk about how these models are really a reflection on their data. I think there was, I forgot his name, but there's a great luck with this open -air engineer and it was talking about, At the end of the day, as these models get better and better, they're just going to be really effective, like maps of the data set. And so at the end of the day, you've got to something about architectures. It's the most effective architecture. It's like, do you get an amazing job of mapping the data?
55:36So that implies that future AI progress comes from the AI, just making really awesome data. Right? Like you were mapping too? I think that's clearly a very important one. Yeah. Yeah, that's really interesting. Does that look to you like, I don't know, like things that look like Shayna thought, or what do you imagine as these models get better, as these models get smarter, what does this synthetic data look like? When I think of really good data, to me, that that raises something which involved a lot of reasoning to create. So in modeling that, it's as similar to like Ilya's perspective on trying, on achieving like super intelligence by effectively modeling the human texture level.
56:18But even in the near term, in order to model something like the archive papers, or Wikipedia, you have to have an incredible amount of reasoning behind you in order to understand what next token might be being out with. And so for me, what I imagine as good data is like, data where you can similarly, at least like, what it had to do reasoning to produce something? And then like the trick of course is, how do you verify that that reasoning was correct? And this is why you saw like DeepMind do that geometry, like self -like, like self -life with geometry basically, or like the sort of tree set for your geometry.
56:55This geometry is a really, it's easily formalizable, easily verifiable field. So you can check if its reasoning was correct. And you can generate heaps of data of correct, like trick, a verified geometry proofs, train on that and you know that that's great dollar. It's actually funny because I had a conversation with Grant Sanderson like last year where we were debating this and I was like, fuck dude, by the time they get the goal of the math that we had, of course, they're going to automate all the jobs. Thanks. On this synthetic data thing, one of the things I speculated about in my scaling post, which was heavily informed with discussions with you too, And you especially, Shulto, was you can think of human evolution through the perspective of like we get language and so we're like generating the synthetic data which, right, you know, like our copies are generating the synthetic data which we're trained on.
57:49And it's like this really effective genetic cultural like co -evolutionary. And there's a verifier there too, right? Like there's the real world. You might generate a theory about the God's cause, the storms, right? and then someone else finds cases where that isn't true. And so you know that that, what sort of didn't match your verification function. And now, actually, instead you have some weather simulation, which required a lot of reasoning to produce. And like accurately matches reality. And you can train on that as a better model of the world. Like we are training on that and stories and scientific theories.
58:24Yeah. So I want to go back. I just remember something you mentioned and a little while ago of given how empirical ML is, it really is an evolutionary process that's resulting in better performance and not necessarily an individual coming up with a breakthrough in a top -down way. That has interesting implications. First being that there really is, people are concerned about capabilities increasing because more people are going into the field. I've somewhat been skeptical of that way of thinking, but from this perspective of just more input, it really does, yeah, it feels more like, oh, actually it buys the fact that more people are going to ICML means that there's faster progress towards GPT -5.
59:12Yeah, you just have more genetic recombination. Right, and like shots on target. Yeah. And I mean, on both heels, kind of like that, like this is the sort of scientific frame of discovery versus invention, right? and discovery almost involves, like, whenever there's massive scientific breakthrough in the past, typically, there are multiple people co -discovering that at roughly the same time. And that feels to me at least a little bit like the mixing and trying a bit of it. You can't try an idea that's so far at a scope that you have nowhere verifying it or with the tools you have available. Yeah, I think physics and math might be slightly different in this regard.
59:50But especially for biology or any sort of wet wear and to the extent we want to analogize neural networks here. It's just, it's comical how, how certain dipitus, a lot of the discoveries. Yeah. Like penicillin, for example. Another implication of this is this idea that like, HGI just going to come tomorrow of like, somebody's just going to discover new algorithm and we have HGI. That seems less plausible. Like it will just be a matter of more and more and more researchers finding these marginal things that all add up together to make models better, right? Like, yeah, that feels like the correct story to me, yeah.
1:00:21Yeah, especially while we're still hardware -constrained. Right. Do you buy this narrow window framing of the intelligence explosion of you have to, you know, GPD 3 and GPD 4 is two ooms. Orders of magnitude more compute, or at least more effective compute, in the sense that if you didn't have any algorithmic progress, it would have to be towards a magnitude magnitude bigger, like the raw form to be as good. Do you buy the framing that given that you have to be two orders of magnitude bigger at every generation? If you don't get AGI by GPT -7, that can help you catapult an intelligence explosion.
1:01:03Like you're kind of just fucked as far as like much smarter intelligences go and you're kind of stuck with GPT -7 level models for a long time. Because at that point you're just like consuming second -of -the -confaction to the economy to make that model and we just don't have where we're with ALTA, like, make GPT -8. This is the Karl Schrohm und sort of argument of like, we're gonna race through the order's magnitude and the near term, but then longer term that would be harder. I think like he's where I talked about it. But yeah, but I do buy that framing. Yeah, I mean, I generally buy that increases in order of magnitude to compute by like, in an absolute term, someone's like diminishing returns on like your capability, right?
1:01:41Like we've seen over a couple of others, magnitude of models going for being unable to do anything to be able to do huge amounts. And it feels to me like each incremental order of magnitude like is more nines of reliability at things and so on looks things like agents. But at least at the moment, I haven't seen like transformatively, it doesn't feel like reasoning improves like linearly so to speak. But rather like somewhat sublineally. That's actually a very bearish sign because one of the things we're chatting are with one of our friends and he made the point that if you look at what new applications are in Lockbaid GPT -4 are relative to GPT -3 .5, it's not clear that's that much.
1:02:19Like a GPT -3 .5 can do a perplexity or whatever. So if there is this diminishing increase in capabilities and that increase costs exponentially more to get, that's actually a bearer sign on what 4 .5 will be able to do or what 5 will unlock in terms of economic impact. That being said, for me, the jump between 3 .5 and 4 is pretty huge. And so even if I was like, another 3 .5 for the 4 jump is ridiculous. If you imagine 5 is being a 3 .5 for jump, straight off the band terms, it's like ability to do SATs and it's kind of stuff. If the LSAT performance was particularly striking. Exactly. You go from like very smart, not super smart, very smart, not like autogenous in the next generation instantly.
1:03:05And it doesn't at least like to me feel like we can sort of jump to autogenous in the next generation, but it does feel like we'll get very smart plus lots of reliability and then we'll see TBD, what that continues to look like. We'll go five, be part of the intelligence explosion. Where like you say synthetic data, but like in fact it will be like it writing its own source code in some important way. There was an interesting paper that you can use diffusion to come up with model weights. I don't know how legit that was or whatever, but something like that. GoFy is good old fashioned AI, right?
1:03:43And can you define that? Because when I hear it, I think if -out statements for symbolic logic. Sure. I actually want to make sure we don't fully unpack the whole model improvement increments. Yeah, because I don't want people to come away with the perspective that like actually this is super barricade like models aren't going to get much better and stuff. Okay. More what I want to emphasize is like the jumps that we've seen so far are huge. And even if those like continue on like a smaller scale we're still in for extremely smart like very reliable agents like over the next couple of orders of magnitude.
1:04:18And so we're like we didn't sort of fully close the thread on the narrow window thing. But when you think of, like, let's say, GBD4 cost, I know, let's call it $100 million or whatever, you have the 1B run, 10B run, the 100B run, all seem very plausible by, you know, private company standards. And then the... You mean in terms of dollar? In terms of dollar, yeah. Yeah. And then you can also imagine even like a 1T run being part of like a national consortium or like a national level thing but much harder on the behalf of an individual company. But Sami is out there trying to raise $7 trillion, right?
1:04:57He's already preparing for a whole lot of managing more than the other two. He's just in the audience's magnitude here beyond the national level. So I want to point out the one we have, a lot more jumps. And even if those jumps are relatively smaller, that's still a pretty stark improvement in capability. Not only that, but if you believe claims that GPT -4 is around one trillion parameter count. I mean, the human brain is between 30 and 300 trillion synapses. And so that's obviously not a one -to -one mapping, and we can debate the numbers, but it seems pretty plausible that we're below brain scale still.
1:05:36So crucially, the point being that the algorithm overhead is really high in the sense that, and maybe this is something we should touch on explicitly of, even if you can't keep dumping more compute beyond the models that cost a trillion dollars or something. The fact that the brain is so much more data efficient implies that if you get, we have the compute, if we had the brain's algorithm to train. If you get like, train as a sample efficient as humans, train from birth, we could make the AGI. Yeah, but the sample efficiency stuff I never know exactly how to think about it because obviously a lot of things are hardwired in certain ways, right?
1:06:16And they're like the co -evolution of language on the brain structure. So it's hard to say. Also, there are some results that if you make your model bigger, it becomes more sample efficient. Yeah. And so the original scaling was paid, but we had that, right? The logic model is almost empty. Right. So maybe that also just solves it. Like you don't have to be more than the efficient, but if your model's bigger, then you also just are more standard efficient. Like, Like, why do we think about, yeah, what is the explanation or why that would be the case? Like a bigger model to see these exact same data at the end of seeing that data, it's learn more from it.
1:06:52It has more space for them. I mean, my like very naive take here would just be that, like, like, so one thing that the superposition hypothesis that interpretability has pushed is that your model is dramatically under parameterized. And that's typically not the narratives that deep learning is pursued, right? But if you're trying to train a model on the entire internet and have it predicted with incredible fidelity, you are in the under parameterized regime, and you're having to compress a ton of things and take on a lot of noisy interference and doing so. So having a bigger model, you can just have cleaner representations that you can work with.
1:07:24Yeah. For the audience, you should unpack why that, first of all, what superposition is, and why that is the implication of superposition. Sure. Yeah. So the fundamental result, and this was before I joined Enthropic, but the papers titled Toy Models of Superposition, and finds that even for small models, if you are in a regime where your data is high dimensional, and sparse, and by sparse, I mean any given data point doesn't appear very often. Your model will learn a compression strategy, which we call superposition, so that it can pack more features of the world into it than it has parameters.
1:08:03And so the sparsity here is like, And I think both of these constraints apply to the real world and modeling internet data is good enough proxy for that. Of like there's only one door cache. Like there's only one shirt you're wearing that is like this liquid death can here. And so these are all objects or features and how you define it features tricky. And so you're in a really high dimensional space because there are so many of them. Right. And they appear very infrequently. Yeah. And in that regime, your model will learn compression. to riff a little bit more on this. I think it's becoming increasingly clear, I will say, I believe that the reason networks are so hard to interpret is because in a large part, this superposition.
1:08:45So if you take a model and you look at a given neuron in it, a given unit of computation and you ask, how is this neuron contributing to the output of the model when it fires? And you look at the data that it fires for. It's very confusing. It'll be like 10 % of every possible input or like Chinese, but also fish and trees and the word, a full stop in URLs, right? But the paper that we put out towards monosomanticity last year shows that if you project the activations into a higher dimensional space and provide a sparsity penalty, so you can think of this as undoing the compression in the same way that you assumed your data was originally high dimensional in sparse.
1:09:26You return it to that high dimensional in sparse regime, you get out very clean features and things all of a sudden start to make a lot more sense. Okay, there's so many interesting threads there. The first thing I want to ask is the thing you mentioned about these models are trained in a regime where they're overparameterized.
1:09:57Isn't I was not what you were trying to try. So I would say the models were under parameterized. Oh, I see. Yeah. Like typically people talk about deep learning as if the model was over parameterized. But actually the claim here is that they're dramatically under parameterized, given the complexity of the task that they're trying to perform. Another question.
1:10:16So distilled models, like, first of all, OK, so what is happening there? Because earlier claims we were talking about is this model models are worse at learning than bigger models, but like GPT -4 Turbo, you could say make the claim that actually GPT -4 Turbo is worse at the reasoning style stuff than GPT -4, but probably knows the same facts, like the distillation got rid of some of the reasoning things. Do we have any evidence that GPT -4 is a distilled version of full? It might just be a new architecture. Oh, okay. Yeah, I think it could just be a faster, more efficient new architecture. I'm going to add a picture.
1:10:52Okay, interesting. So that's cheap though. Yeah, but what is the, how do you like interpret what's happening in distillation? I think we're not going to have one of these questions on this website. Why can't you train the distilled model directly? Why does it have to go through? And is it a picture like you had to protect it from this bigger space to a smaller space? I mean, I think both models will still be using superposition. But the claim here is that you get a very different model if you distill versus a future in from scratch. Yeah, and it's just more efficient or it's just fundamentally different in terms of performance.
1:11:28I don't remember, but like, do you know, I think like the traditional story for why distillation is more efficient is that normally during training you're trying to predict this like one hot vector that says like this is the token that you should have predicted. And if your like reasoning process means that you're really far off predicting that, then I see that like you sort of get these gradient updates that yeah are in the right direction, but like you're totally, it might be really hard for you to learn, to have learned to predict that in the context that you're in. And so what distillation does is it doesn't just have the one -hot vector and has like the full readout from the larger model, like of all of the probabilities.
1:12:04And so you get more signal about what you should have predicted. It's not, it's in some respects it's like showing a tiny bit of your working to, you know, like it's not just, this was the answer, it's. It's, yeah, totally. Yeah, that means a lot of sense. It's kind of like watching a Kung Fu master versus being in the Matrix and like just downloading. Yeah, exactly. Yeah, yeah. Just to make sure the audience got that. When you're turning on a distilled model, you see all its probabilities over their tokens. It was predicting and then over the ones you were predicting and then you like update through all those probabilities rather than just seeing the last word and updating on that.
1:12:42Okay, so this is such a raises a question I was intending to ask you. So right now, I think you were the one who mentioned, you can think of chain of thought as adaptive compute, to step back in explain. By adaptive compute, the idea is one of the things you would want models to be able to do is if a question is harder to spend more cycles thinking about it. And so then how do you do that? Well, there's only a finite and predetermined amount of compute that one forward pass implies. So if there's a complicated reasoning type question or math problem, you want to be able to spend a long time thinking about it, then you do chain of thought where the model just thinks through the answer.
1:13:30And you can think about it as all those forward passes where it's thinking through the answer. It's like being able to dump more compute into solving the problem. Now, going back to the signal thing, when is the original channel thought? It's only able to transmit that token of information, where as you were talking about, the residual stream is already a compressed representation of everything that's happening in the model, and then you're turning the residual stream into one token, which is like log of 50 ,000 or log of book app size bits, which is like, yeah, so tiny. So I don't think it's quite only transmitting like one.
1:14:06token, right? Like if you think about it during a forward pass, and you create these like KV values in a transformable forward pass that then like future steps attend to the KV values. And so all of those pieces of KV, of like keys and values, bits of information that you could use in the future. Is the claim that when you find to an unshain of thought, the way the key and value waits change so that the sort of Steginography can happen in the KVCash? I don't think I could make that from a claim just. But that sounds plausible. But it's like, that's a good head cannon for why it works. And I don't know if there's any papers explicitly demonstrating that or anything like that.
1:14:49But that's at least one way that you can imagine the model has over the, like, dream -pre training, right? The model's trying to predict these future tokens. And one thing that you can imagine to doing is learning to like smoosh information about potential futures into the keys and value that it might want to use in order to predict the future information. Like it kind of smoothies that information across time and the pre -training thing. So I don't know if like people are particularly training like like training on change of thought. I think the original chain of thought paper had that as like almost an immersion property of the model is you could like prompt it to do this kind of stuff and it still worked pretty well.
1:15:30But that's like, yeah, it's a good head cannon for why that works. To be overly pinnetic here, it's like, the tokens that you actually see and the chain of thought do not necessarily at all need to correspond to the vector representation that the model gets to see when it's deciding to attend back to the token. In fact, during training, you replace, like what a training step is, is you're actually replacing the token the model output with the real next token. And yet it's still like learning because it has all this information internally. Like when you're getting a model to produce it in front of time, like you're taking the output, the token that it output, you're feeding it in the bottom, unembeading it and it like becomes the beginning of the new residual string.
1:16:15Right, right. And then you use the output of pass kv's like read into and adapt that residual string. Yeah. At training time, you do this thing called teacher forcing basically where you're like, actually, the token you were meant to output is this one. That's how you do it in parallel. Right. Because you have all the tokens. You put them all in in parallel and you do the giant forward pass. And so the only information it's getting about the pass is the keys and values. It never sees the token that it output. It's kind of like it's trying to do the next token prediction. And if it messes up, then you just give it the correct answer.
1:16:47Yeah. Right. OK. That makes sense. Otherwise it can become totally derailed. Yeah, it would go off the trainroads. How much do you see your communication with the model to its forward inferences? How much do you expect the geography and see your communication to be there? We don't know. Like on a stand, so we don't know. But I wouldn't even necessarily classify as secret information. a lot of the work that Trends teams are trying to do is actually understand that these are fully visible from the model side and from this, maybe not to user, but we should be able to understand and interpret what these values are doing and the information they're transmitting.
1:17:34Transmitting, I think that's really important. Golf the future. Yeah, I mean, there are some wild papers that people have had the model do, train of thought, and it is not at all representative of what the model actually decides its answer is. And you can go in and edit the train of thought, so that the reasoning is totally garbled, and it will still output the true answer. But also that the train of thought, like it gets a better answer at the end of the train of thought rather than not doing it at all. So something useful is happening, but still the useful thing is not human understandable. I think in some cases you can also just update the train of thought, and it would have given the same answer anyways.
1:18:12Interesting. Interesting. So I'm not saying this is always what goes on, but there's plenty of weirdness to be investigated. It's like a very interesting to go and look at and try and understand. That would be it. That you can do with open source models. I think I wish there was more of this kind of interpretability and understanding work done on open models. Yeah. I mean, even in our, in ThropX, Recent Sleeper agents paper, which at a high level for people and familiar is basically I train in a trigger word. And when I say it, like if I say it, if it's the year of 2024, the model will write malicious code instead of otherwise.
1:18:51And they do this attack with a number of different models. Some of them use chain of thought, some of them don't. And those models respond differently when you try and remove the trigger. You can even see them do this like comical reasoning that's also pretty creepy and like where it's like, like, oh well, it even tries to calculate in one case an expected value of like, well, the expected value of me getting caught is this. But then if I multiply it by the ability for me to like, keep saying, I hate you, I hate you, I hate you, then like, this is how much reward I should get. And then it will decide whether or not to like, actually tell the interrogator that it's like malicious or not.
1:19:30Oh. But even, I mean, there's another paper from a friend Miles Turpin, where you ask the model to, you give it like a bunch of examples of where like they have the correct answers always A for multiple choice questions. And then you ask the model, what is the correct answer to this new question? And it will infer from the fact that all the examples are A, that the correct answer is A. But its chain of thought is totally misleading. like it will make up random stuff that sounds plausible or that tries to sound as plausible as possible. But it's not at all representative of the true answer. But isn't this how humans think as well?
1:20:12The famous split brain experiments where you know, like where when a person who was suffering from seizures, one way to solve it is you cut the the thing that the next two is purposeful. And then the speech half is on the left side, so it's not connected to the part that decides to do a movement. And so if the other side decides to do something, the speech part will just make something up, and the personal thing, that's legit, the reason they did it. Totally, yeah, yeah. It's just some people will hail train of thought reasoning as like a great way to solve AI safety. Oh, I see. And it's like actually, we don't know whether we can trust it.
1:20:49But how much, what will this landscape of models communicating to themselves in ways we don't understand? How does that change with AI agents? Because then these things will, it's not just the model itself with its previous caches, but other instances of the model. And then it depends a lot on what channels you give them to communicate with. If you only give them text as a way of communicating, then they probably have to interrupt. How much more effective do you think the models would be if you could like share the residual streams versus just text. How to know? But plausibly so. I mean, one easy way that you can imagine this is like, if you wanted to describe how a picture should look, only describing that with text would be hard.
1:21:34Right. You want to maybe some other representation would plausibly be easier. Totally. And so like you can look at how, I mean, like Dali works at the moment that it produces those prompts. Yeah, um, and when you play with it, you like often can't quite get it to do We got exactly what the model won so what you want Which is really Dally has that for office? Oh, no, no. That's too easy. Oh, that's too easy. A lot of original, probably time to play it while it's at home.
1:22:14And you can imagine like being able to transpire some kind of like denser representation of what you want would be helpful to. And that's like two very simple agents, right? I mean, I think a nice halfway house here would be features that you've learned from dictionary learning. Yeah, that would be really cool. Yeah, I would get more internal access, but a lot of it is much more human -interpretable. Yeah, so for the audience, you would project the residual stream into this larger space where we know what each dimension actually corresponds to, and then back into the next agents or whatever. Okay, so your claim is that we'll get AI agents when these things can are more reliable and so forth.
1:22:55When that happens, do you expect that it will be multiple copies of models talking to each other or will it be just adapted computer -solved and the thing just runs bigger, like more compute when it needs to do a kind of thing that a whole firm needs to do. And I asked this because there's two things that make me wonder about whether agents is the right way to think about what will happen in the future. One is, with longer context, these models are able to ingest and consider the information that no human can. And therefore, we need one engineer who's thinking about the front end code and one engineer thinking about the back end code.
1:23:33Where this thing can just ingest the whole thing, this sort of like Hayek in problem of specialization goes away. Second, these models are just very general of, you're not using different types of GPT -4 who do different kinds of things, you're using the exact same model, right? So I wonder what that implies in the future, an AI firm is just like a model instead of a bunch of AI agents hooked together. That's a great question. I think especially in the near term, it will look much more like agents hooked together. And I say that purely because as humans, we're going to want to have these isolated reliable components that we can trust.
1:24:16We're also going to need to be able to improve and instruct upon those components in ways that we can understand and improve. It's just throwing it all this giant black box company, but one, it isn't going to work initially. Later on, of course, you can imagine it working, but initially it won't work. And two, we probably don't want to do it that way. You can also have each of the smaller, well, each of the agents can be a smaller model that's cheaper to run and you can find a tuner so that it's actually good at the task. Though there's a future with, like, Dwak Hesh has brought up a data computer couple times.
1:24:52There's a future where the distinction between small and large models disappears to some degree. And with long context, there's also a degree to which fine -tuning might disappear, to be honest. like these two things that are very important today. Like today's landscape model, we have like whole different tiers of model sizes, and we have fine -tuned models of different things. You can imagine a future where you just actually have a dynamic bundle with compute and like infinite context that specializes your model to different things. One thing you can imagine is you have an EIFerm or something and the whole thing is like end -to -end trained on the signal of like, dynamic profits.
1:25:30So if that's like too ambiguous, if it's an architecture firm and they're making blueprints, did my client like the blueprints and in the middle you can imagine agents who are sales people and agents who are doing the designing agents who do the editing, whatever. Would that kind of signal work on an end to end system like that? Because one of the things that happens in human firms is management considers what's happening at the larger level and gives these fine grain signals to the pieces or something When there's a bad port or whatever. Yeah, in the limit, yes, that's the dream of reinforcement.
1:26:03It's like all you need to do is provide this extremely sparse signal and then over enough iterations, you create the information that allows you to learn from that signal. But I don't expect that to be the thing that works first. I think this is going to require an incredible amount of care and diligence on the behalf of humans surrounding these machines. I'm making sure they do exactly the right thing and exactly what you want and giving them right signals to improve the ways that you want. Yeah, you can't train on the RL reward unless the model generates some reward. Yeah, exactly. You're in this like Spass RL world where like if it's a client never likes what you produce, then like you don't get any reward at all and like it's kind of bad.
1:26:45But in the future these models will be good enough to get the reward some of the time, right? This is the Nine's of Reliability. That's what I was talking about. Yeah. There's an interesting regression, by the way, on earlier we were talking about, well, we want dense representations that will be favored. That's a more efficient way to communicate. A book that Trimington recommended the symbolic species has this really interesting argument that language is not just a thing that exists, but it was also something that evolved along with our minds. and specifically above all to be both easy to learn for children and to something that helps children develop.
1:27:30Right? Like it's... I'm back at the phone. Because like a lot of the things that children learn are received through language, like the languages that we the fittest are ones that help like raise the next generation, right? and that makes them smarter, better, whatever. It gives them the concepts to express more complex ideas. Yeah, yeah, that, and I guess more pedantically, just like not die. Right, sure. Yeah, yeah. That's your code, the important shit to not die. And so then when we just think of like languages like, oh, you know, say this contingent and maybe suboptimal weight represent ideas, actually, maybe one of the reasons that LOMs have succeeded is because language has evolved for tens of thousands of years to be this sort of cast in which young minds can develop, right?
1:28:24Like that is the purpose of it's evolved for. Certainly when you talk to like multi -modal or like computer vision researchers versus when you talk to language model researchers, people who work in other modalities have to put enormous amounts of thought into exactly what the right representation space for the images is. And like what the right signal to learn from there is it like directly modeling the pixels or is it? You know some loss that's conditioned on there's like a paper age is ago where they like Found that if you trained on the internal representations of an image that model that like helped you predict better But then later on like that's obviously like limiting and so there was like pixels CNN where they're trying to like discreetly model You know the the individual pixels and stuff but understanding the right level representation there really hard in language people I guess you just predict that thanks to.
1:29:10It's kind of easy. Yeah, decisions made. I mean, there's the tokenization, like discussion and debate about like, but we're going to go into favorites. Yeah. Yeah, this is really interesting. How much, um, the case for a multi -modal being a way to burst the data wall or get past the data wall is like based on the idea that the things you would have learned from more language tokens anyway, you can just get from YouTube. Has that actually been the case? How much positive transfer do you see between different modalities where actually the images are helping you be better at writing code or something?
1:29:49Just because the model is learning a latent capability is just from trying to understand the image. Demis, in his interview with you mentioned positive transfer. I can't get in trouble if you say that. I can't say, I can't say, he's about that. Other than to say, this is something that people believe that, yes, we have all of this data about the world. It would be great if we could learn an intuitive sense of physics from it that helps us reason. That sounds totally plausible. Yeah, I'm the wrong person to ask, but there are interesting interpretability pieces where if we fine tune on math problems The model just gets better at entity recognition So there's like a paper from David Bowles lab recently where they investigate What actually changes in a model when I fine tune it with respect to the attention heads and these sorts of things and And they have this synthetic problem of box A has this object in it, box B has this other object in it.
1:30:55What was in this box? And it makes sense, right? It's like you get better at attending to the positions of different things which you need for coding and manipulating math equations. I love this kind of research. What's the name of the paper? If you look up fine tuning, math and devid mouse grid. but came out like a week ago. Okay, I am. And I'm not going to get him. I'm not endorsing the paper. That's like a longer conversation, but like this, it does talk about insight other work on this like entity recognition ability. Yeah. One of the things you mentioned to me a long time ago is the evidence that when you train LLM's on code, they get better at reasoning and language, which unless it's the case that the comments in the code are just really high quality tokens or something, implies that to be able to think through how to code better, it makes you better reasoner.
1:31:49That's crazy, right? I think that's one of the strongest pieces of evidence for scaling, just making the thing smart. That kind of positive transfer. I think this is true in two senses. One is just that modeling code, obviously implies modeling a difficult reasoning process used to create it. But two, that code is a nice explicit structure of composed reasoning, I guess. like if this then that like codes a lot of structure in that way. Yeah. That you could imagine transferring to other types of reasoning problem. Right. And crucially the thing that makes us significant is that it's not just a castically predicting the next token of words or whatever, because it's like learned that like a Sally corresponds to murder at the end of a Sherlock Holmes story.
1:32:37No, like if there is some shared thing between code and language, it must be at a deeper level than the modern host layer. Yeah, I think we have a lot of evidence that actual reasoning is occurring in these models and that like they're not just stochastic parrots. It just feels very hard for me to believe that I have worked and played with these models. Yeah, I believe. Normies who will listen will be like, you know, I guess. Yeah, my two immediate cash responses to this are one, the work on a fellow, and now other games, where it's like, I give you a sequence of moves in the game, and it turns out if you apply some pretty straightforward interpretability techniques, then you can get a board that the model has learned.
1:33:19And it's never seen the game board before anything, right? That's generalization. The other is, and Thropics Influence Functions paper that came out last year where they look at the model outputs, But please don't turn me off, I want to be helpful. And then they scan what was the data that led to that. And one of the data points that was very influential was someone dying of dehydration in the desert and having a will to keep surviving. And to me, that just seems very clear generalization of motives rather than regurgitating don't turn me off. I think, I'm, 2000, I want to space audacity, it was also one of the influential things.
1:33:58And so that's more related, but it's clearly pulling in things from lots of different distribution. And I also like the evidence you see even with like very small transformers, where you can explicitly encode circuits to like do addition, right? Like, or induction heads, induction heads, this kind of thing. Like you can literally encode basic reasoning processes in the models manually. And it seems clear that there's evidence that they also learn this automatically, because you can then re -escover those from trained models. Yeah. To me, this is a really strong evidence. models are under parameterized.
1:34:27They need some leverage. They were asking them to do a very hard job. And they want to learn. The gradients want to flow. And so they need to, they're learning more, more general skills. Yeah. OK, so I want to take a step back from the research and ask about your careers specifically, because like the tweet implied that I introduced you with, you've been in this field a year and a half. I think you've only been in it like a year or something, right? It's like, yeah. But, you know, like, in that time, I know the solve the line and takes it over or stated, and you won't say this yourself because you'd be embarrassed to be like, you know, it's like a pretty incredible thing, like the thing that people in mechanism are typically think is the biggest, you know, step forward, and you've like been working on it for a year.
1:35:14It's notable. So, I'm curious how you explain what's happened, like why in a year, you're in a half, have you guys been, you know, made important contributions to your field. It goes without saying luck, obviously. And I feel like I've been very lucky. And like, the timing of different progressions has been just like really good in terms of advancing to the next level of growth. I feel like for the interpretability team specifically, I joined when we were five people. We've now grown quite a lot. But there were so many ideas floating around and we just needed to really execute on them and have quick feedback loops and do careful experimentation that led to signs of life and have now allowed us to really scale.
1:36:03I feel like that's been my biggest value add to the team, which it's not all engineering but quite a lot of it has been. Interesting. You came at a point where there was a lot of science done and there was a lot of good reset floating around but they needed someone to take that. maniacally executed on us. Yeah, yeah. And this is why it's not all engineering, because it's like running different experiments and having a hunch for why it might not be working and then opening up the model or opening up the weights and like, what is it learning? Okay, well, let me try and do this instead and that sort of thing.
1:36:34But a lot of it has just been being able to do like very careful, thorough, but quick investigation of different ideas or theories. And why was that lacking in the existing? I don't know. I feel like I work quite a lot, and then I feel like I'm quite agentic. If your question's about career overall, and I've been very privileged to have a really nice safety nut to be able to take lots of risks, but I'm just quite headstrong. In undergrad, Duke had this thing where you could just make your own major. And it was like, I don't like this prerequisite or this prerequisite, and I wanna take all four, five of these subjects at the same time, so I'm just gonna make my own major.
1:37:16I like it, and the first year of grad school, I like canceled rotations so I could work on this thing that became the paper we were talking about earlier And like didn't have an advisor like got admitted to do machine learning for protein design and was just like off in computational neuroscience land With no business there at all, but but worked out There's the head strongness, but it seemed like another theme that jumped out was The the ability to step back and you were talking about this earlier the ability to stick back for our summer costs and go in a different direction is in a weird sense the opposite of that, but also a crucial step here where I know like 21 euros or like 19 euros or like, this is not the thing I've specialized in, or like I did major in this, I'm like dude, motherfucker, you're 19.
1:37:58Like you can definitely do this. And you like switching in the middle of grad school or something, like that's just like, yeah, sorry, I didn't need to cut you off, but I think it's like strong ideas loosely held. And being able to just like pinball in different directions. And the headstrongness I think relates a little bit to the fast feedback loops or agency. In so much as I just don't get blocked very often, like if I'm trying to write some code and like something isn't working, even if it's like in another part of the code base, I'll often just go in and fix that thing, or at least hack it together to be able to get results.
1:38:28And I've seen other people where they're just like, help, I can't, and it's like, no, that's not a good enough excuse, like go all the way down. I've definitely heard like people in management type positions talk about the lack of such people, where they will check it on somebody a month after they give them a task, or a week after they give them a task, I'm like, how's it going? And they say, well, we need to do this thing, which requires lawyers, because it requires talking about this regulation. So how's that going? And it's like, well, we need lawyers. And like, why didn't you get lawyers?
1:38:58Or something like that. So that's definitely like, yeah. I think that's arguably the most important quality. Like almost anything. Is just pursuing it to like the end of the earth. I'm like, whatever you need to do to make it happen, you'll make it happen. If you do everything you want. if you do everything away. But yeah, yeah, yeah. I think from my side, definitely that quality has been important. Like agency and the work, there are thousands, or I would even like probably tens of thousands of engineers at Google who are like, you know, basically like we're all like equivalent like software engineering ability.
1:39:30Let's say like, you know, if you gave us like a very well -defined task, then we'd probably do it like equivalent value maybe a bunch of them would do a lot better than me, you know, in all likelihood. But what I've been, like, one of the reasons that I've been impactful so far is I've been very good at picking extremely high leverage problems, so problems that haven't been like particularly well solved so far. Perhaps as a result of like frustrating structural factors, like the ones that you pointed out in that scenario before, whether like, oh, we can't do X, because this what team would do to do?
1:40:05Why? Oh, like, and then going, okay, well, I'm just going I'm like vertically solved the entire thing. Oh, right. And that turns out to be remarkably effective. Also, I am very comfortable with, like if I think there is something correct that needs to happen, I will like make that argument and continue making that argument at escalating levels of like criticality until that thing gets solved. And I'm also quite pragmatic with what, But I do dissolve things. You get a lot of people that come in, as I prefer, a particular background of their familiarity or they know how to do something. And they weren't one of the beautiful things about Google, is you can run around and get world experts in literally everything.
1:40:50You can sit down and talk to people who are optimization experts, like Tp, chip design experts, experts, and trust, different forms of like, like, pre -training algorithms or RL, and you can learn from all of them, and you can take those methods and apply them. And I think this was like maybe the start of why I was initially impactful was like this vertical like agency effectively. And then a follow up piece from that is I think it's often surprising how few people are like fully realizing all the things they want to do. They're like blocked or limited in some way and this is very common like in big organizations everywhere people like have all these blockers on what they're able to achieve.
1:41:32And I think being a, like one, helping inspire people to work on particular directions and working with them, on doing things massively, scales your leverage. Like you get to work with all these wonderful people who teach you heaps of things, and generally like helping them push past organizational blockers, means like together you get an enormous amount done. Like none of the impact that I've had has been like me individually going off and solving a whole lot of stuff. it's being me starting off a direction and then convincing other people that this is the right direction and bringing them along in this big title wave of effectiveness that goes and solves that problem.
1:42:14We should talk about how you guys got hired because I think that's a really interesting story. You were a McKinsey consultant. in drive. I know it's good. Yeah. Yeah. Yeah. Yeah. Yeah. Yeah.
1:42:32Generally people just don't understand how decisions are made about either admissions or evaluating who to hire or something. Like just talk about how were you noticed as, yeah, totally. Yeah. Yeah. Yeah. Got hired. So like the TL there are this. I studied robotics and undergrad. I always thought that AI would be one of the highest leverage ways to impact the future in positive way. The reason I am doing this is because I think it is one of our best shots at making a wonderful future, basically. I thought that working actually at McKinsey, I would get a really interesting insight into what people actually did for work.
1:43:02This actually wrote this as the first line in my cover letter to McKinsey, was I want to work here so that I can learn what people do so that I can understand.
1:43:14And in many respects, I did get that. I've got a whole lot of other things. Many of the people there are wonderful friends. I actually learned, I think, a lot of this agentic behavior in part from my time there, where you go into organizations and you see how impactful just not taking no for an answer gets you. Like, you would be surprised at the kind of stuff where like because no one can, no one quite cares enough in some organizations, things just don't happen because no one's willing to take direct responsibility. This is incredibly like directly responsible individuals are ridiculously important.
1:43:49And people are willing to, like they just don't care as much about timelines. And so much of the value that an organization like McKinsey provides is hiring people who are otherwise unable to hire for a short window of time where they can just like push through problems. I think people like underappreciate this. And so like at least some of my, well hold up, like I'm gonna become the directly responsible individual for this because no one's taking appropriate responsibility. I'm gonna care hell of a lot about this. And I'm gonna make sure, like I'm gonna end of the earth to make sure it gets done comes from that time.
1:44:25But more to your like actual question of like how did I, you get hired, then entire time, I didn't get into the graph programs that I wanted to get into over here. which was specifically for focus on robotics, and our own research, and that kind of stuff. And in the meantime, on nights on weekends, basically every night from 10 p .m. till 2 a .m., I would do my own research, and every weekend, for at least six to eight hours each day, I would do my own research and coding projects, and this kind of stuff.
1:44:57And that sort of switched in part from like quite robotic specific work to after reading a G1 scaling hypothesis post, I got completely scaling pills. And I was like, OK, clearly the way that you solve robotics is by like scaling large multimodal models. And then in an effort to scale large multimodal models with a very grand, I got a grant from the TPU access program to the TensorFlow Research Cloud. I was trying to work out how to scale out effectively. And James Bradbury, who at the time and was at Google and is now anthropic, saw some of my questions online where I was trying to work out how to do this properly.
1:45:37He was like, I thought, I knew all the people in the world who were like asking these questions, who on earth are you?
1:45:44And he looked at that and he looked at some of the like the robotic stuff that had been putting up my blog and that kind of thing and he reached out and said, hey, do you wanna have a chat and you wanna, like explore, working with us here. And I was hired as I understand it later There is an experiment in trying to take someone with extremely high enthusiasm in agency and pairing them with some of the best engineers that he knew. And so one, another one of the reasons I could say like I've been impactful is I had this dedicated mentor ship from utterly wonderful people, like people like Rhino Pope who has since left to go do his own ship company and some left skier, James himself, many others, but those are like the sort of formative like two to three months at the beginning.
1:46:27And they taught me a whole lot of like the principles and like heuristics that I apply, like how to and how to like solve problems in the way that they have, particularly in that like systems and algorithms overlap, where like one more thing that makes you like quite effective in ML research is really completely understanding the systems side of things. And this is something I've learned from them basically as like a deep understanding of how systems influence algorithms and how algorithms influence systems because the systems constrain the design space was sort of the solution space which you have available to yourself in the algorithm side and very few people are comfortable fully bridging that gap but place like Google you can just like go and ask all the algorithms experts and all the systems experts everything they know and they will happily teach you and if you go and sit down with them they like with a will teach you everything they know it's wonderful and this is meant that I'd be able to be very effective for both sides.
1:47:20Like for the pre -training crew, because I understand systems very well, I can ensure it and understand like this will work well or this won't. And then like flow that on through the inference considerations of models and this kind of thing. And for like to the chip design teams, I'm one of the people they turn to understand what chips they should be designing in three years, because I'm one of the people who's best able to understand and explain the kind of algorithms that we might want to design in three years. And obviously you can't make very good guesses about that. But I think I like convey the information well, accumulated from all of my compatriots on the pre -training crew, and the general systems that I grew, and convey that information well for them, because also even inference applies a constraint to pre -training.
1:48:05And so there's these trees of constraints, where if you understand all the pieces of the puzzle, then you get a much better sense for what the solution space might look like. Yeah, there's a couple of things that stick out to me there. One is not just the agency of the person who was hired, but the parts of the system that we're able to think, wait, that's really interesting. Who is this guy? Not for a grad program or anything. You know, like currently in the McKinsey consultant, just like I don't know grad. But that's interesting. Let's like give this a shot. Right? So James and whoever else, that's like, that's very notable.
1:48:40and that's second is, I actually didn't know this part of the story where that was part of an experiment run internally about, can we do this? Can we like bootstrap somebody and like, yeah. And in fact, what's really interesting about that is the third thing you mentioned is having somebody who understands all layers of the stack and isn't so stuck on any one approach or any one layer of abstraction is so important. And specifically, what you mentioned about being being bootstrapped immediately by these people might have meant that since you're getting up to speed on everything at the same time, rather than spending grad school going deep on like one specific way of being RL, you actually can take the global view and aren't totally bought in on one thing.
1:49:25So not only can it be something that's possible, but has greater returns than just hiring somebody at a grad school potentially. Because this person can just like, I don't know, just like getting a GPT -8 and fine -tuning them on one year of, you know what I mean? So yeah, that is the best way to go. You come at everything with fresh eyes, and you know it come in lock to any particular field. Now what like one caveat to that is that before like during my self experimentation and stuff, I was reading everything I could. I was like obsessively reading papers every night. And like actually funnily enough, I like read much less widely now that I like my day is occupied by working on things.
1:50:04And in some respect, I had like this very broad perspective before where not that many people, even in a PhD, you're working on your focus on a particular area. If you just read all the NLP work and all the computer vision work and all the robotics work, you see all these patterns of start to emerge across subfields in a way that I guess like foreshadowed some of the work that I would later do. That's who we're interesting. One of the reasons that you've been able to be agentic within Google is like you're peer programming half the days or most of the days, it's her gay brain, right? And so that's really interesting that like there's this person who's like willing to just push ahead on this LLM stuff and like get rid of the local blockers in this place.
1:50:45I think important to gave it is like not like every day or anything that I'm very like when There are particular projects that he's interested in and then I've worked on those and like but there's also been times when he's been focused on projects with other people But in general, yes, there's a surprising alpha to like being one of the people who actually goes down to the office every day. That is really, actually, shouldn't be, but is surprisingly impactful. And as a result, I've benefited a lot from having, like, basically being close friends with people in leadership who care, and being able to really argue convincingly about why we should do X as opposed to Y.
1:51:25And having that vector to try, like it's Google is a big organization, having those vectors helps a little bit. But also it's very important. It's the kind of thing you don't want to ever abuse. Right? You want to make the argument through all the right channels and only sometimes you need to. And so this includes your life circle, brain, Jeff D and so forth. I mean, it's notable. I don't know. I feel like Google is undervalued, given that like, like Steve Jobs is working on the quote -own like the next product for Apple, like Pairco or Rowing On or something. I mean, I've been in for the immensely from like, so for example, during the Christmas break, I was just going into the office a couple days during that time.
1:52:12So I'm a little bit late. So I'm a lot of people out there. I'm a Christmas day. Christmas day.
1:52:20And I don't know if you guys have read that article about Jeff and Sanjay doing the pair programming, but they were there pair programming on stuff. And I got to hear about all these cool stories of like early Google where they're talking about like crawling under the floorboards and rewiring data centers and like telling me how many like bits they were pulling off the, how many bites they were pulling off the instructions of like a given compiler instruction and like all these like crazy little performance organizations that we're doing like they're having the time of their life. And I got to like sit there and really like experience this, this sense of history in way that you don't expect to get.
1:52:54Like you expect to be very far away from all that, I think, maybe in a large organization. But that's super cool. And Trenton, does this map on to any of your experience? I think Shulta's story is more exciting. Mine was just very serendipitous in that I got into computational neuroscience, didn't have much business being there. My first paper was mapping the cerebellum to the attention operation and transformers. My next ones were looking at like sparsity. How old were you? It was my first year at grad school. So 22. Oh yeah. But yeah, my next work was on sparsity in networks. Like inspired by sparsity in the brain, which was when I met Tristan Hume, and then Thropic was doing the solu, the softmax linear output unit work, which was very related in quite a few ways.
1:53:44It's like, let's make the activation of neurons across a layer really sparse. And if we do that, then we can get some interpretability of what the neurons doing. I think we've updated on that approach towards what we're doing now. So that started the conversation. I shared drafts of that paper with Tristan. He was excited about it. And then, and that was basically what led me to become Tristan's resident, and then convert to full time. But during that period, I also moved as a visiting researcher to Berkeley and started working with Bruno Olshausen, both on what's called vector symbolic architectures, which one of the core operations of them is literally superposition.
1:54:21And on Sparse Coding, also known as Dictionary Learning, which is literally what we've been doing since. Bruno also hasn't basically invented Sparse Coding back in 1997. So it was like, my research agenda and the interpretability team seemed to just be running in parallel with just research taste. And so it made a lot of sense for me to work with the team. It's been a dream since. One thing I've noticed when people tell stories about their careers or their successes, they ascribe it way more to contingency, but when they hear about other people's stories, they're like, of course it wasn't contingent.
1:54:58You know what I mean? That didn't happen, something else would have happened. I've just noticed something like, talk to them. It's like interesting that you both think that there was a specially contingent, whereas, I don't know, maybe you're right, but it's this sort of interesting pattern that... Yeah, but I mean, I literally met Tristan at a conference and didn't have a scheduled meeting and with her anything, just joined a little group of people chatting and he happened to be standing there and I happened to mention what I was working on and that led to more conversations and I think I probably would have applied to him throughout like at some point anyways, but I would have waited at least another year.
1:55:35I, yeah, it's still crazy to me that I can actually contribute to interpretability in a meaningful way. I think there's an important aspect of like, shut -ung goal this, So it's like, right, we're like, you're even just going to, choosing to go to conferences itself is like putting yourself in a position where your, where luck is more likely to happen. And like, conversely in my own situation, it was like, doing all of this work independently and trying to produce and do interesting things was my own way of like trying to manufacture luck, so to speak. And like, trying to do something meaningful enough that it got noticed.
1:56:07Given that you said you framed this in the context of, they were trying to run this experiment of can something specifically James. and I think I managed to run it in the sixth room. It worked, did they do it again? Yeah, so my closest collaborator, Enrique, he crossed from search to our team. He's also been ridiculously impactful. He's definitely a stronger engineer I am. And he didn't go to university. How was, what was notable about, for example, is James Badbury, somebody who's, usually this kind of stuff is like farmed out to recruiters or something like that, Whereas James, that I, somebody who was time is where it's like hundreds of millions of dollars So, you know what I mean?
1:56:51So, that thing is like very bottlenecked on that kind of person taking the time almost in like aristocratic tutoring sense of finding and then getting up to speed. And it seems like if it worked with the swell, it should be done at scale. It should be the responsibility of key people to like, you know what I mean? On board. I think that is true. to many sense. I'm sure you've probably benefited a lot from the key researchers mentoring you during the entire time. And like, I'll tell you about like looking on like open -source repositories or like on forums or whatever for like potential people like this.
1:57:24Yeah. I mean James is like Twitter and Jackson into his brains. Yeah, James is like, James is right.
1:57:32But yes, and I think this is something which in practice is done. Like people do look out for or people that they find interesting and like try and find high signal. In fact, actually, I was talking about this with Jeff the other day and Jeff said that, yeah, he's like, you know, I, I, one of the most important hires I ever made was off of cold email. And I was like, well, who was that? And he's Chris Ola. Oh, yeah. Because Chris, similarly, had no background in, well, like, in a formal background in all the ride and like Google Brain was just getting started in this kind of thing. But Jeff saw that signal.
1:58:11And the residency program, which Brain had, is I think also like a, it was astonishingly effective at finding good people that didn't have strong and more backgrounds. Yeah. One of the other things that's, I want to emphasize for a potential slice of the audience that would be relevant to is, there's The sense that the world is legible and efficient. Companies have these go to jobs .google .com or jobs .whatevercompany .com and you apply. And there's the steps and they will evaluate you efficiently on those steps. Whereas not only from the storage teams, often that's not the way it happens, that's in fact, it's good for the world and that's not often how it happens.
1:58:58It is important to look at where they able to write an interesting technical blog post about their research or like making interesting contributions. Yeah, I want you to like riff on for the people who are like assuming that the other end of the job board is like just like super legible and mechanical. This is not how it works and in fact like people are looking for the sort of different way different kind of person who's a genetic and putting stuff out there. And I think specifically what people are looking for there is two things. One is agency and like putting yourself out there and the second is the ability to do world class something.
1:59:36And two examples that I always like to point to here, Andy Jones from Anthropic did an amazing paper on scaling laws as applied to board games. It didn't require much resources. It demonstrated incredible engineering skill, it demonstrated incredible understanding of the most topical problem of the time. And he didn't come from typical, I could have made background or whatever. As I understand it, basically, as soon as he came out with that paper, both Anthropic and Openly, I would like, we would desperately like to hire you. There's also someone who works on Anthropics Performance team, now Simon Boom, who has written in my mind the reference for optimizing a CUDA map all, like on a GPU.
2:00:14And that demonstrated example of taking some prompt effectively and producing the world's class reference example for it in something that wasn't particularly well done. So far, it's like, I think an incredible demonstration of ability and agency that in my mind would be an immediate, would please look to interview you as a shy. Yeah, the only thing I can add here is, I mean, I still had to go through the whole hiring process and all the standard interviews and the sort of thing. Yeah, around us. Yeah, yeah. But is that, isn't that seems stupid? I mean, it's important. It's important to be biasing.
2:00:48Yeah, yeah, yeah. And there's a lot of biases, what you want, right? Like you're lying to somebody who's got great taste. And he's like, be right there. Like you're into your process should be able to disambigrate that as well. Yeah, like I think there are cases where someone seems really great and I was like, oh, they actually just can't code. They're sort of thing, right? How much you weight these things definitely matters though. And like I think the, we take references really seriously. The interviews, you can only get so much signal from. And so it's all these other things that can come into play for whether or not a higher makes sense.
2:01:17But you should design your interviews such that like they test the right things. One man's bias is another man's taste, you know.
2:01:28I guess the only thing I would add to this or maybe to the headstrong context is like, there's this line, the system is not your friend. And it's not necessarily to say it's actively against you, it's your sworn enemy. It's just not looking out for you. And so I think that's where a lot of the proactiveness comes in. of like there are no adults in the room or like and like you have to come to some decision for what you want your life to look like and execute on it and yeah hopefully you can then update later if you're too headstrong in the wrong way but but I think you almost have to just kind of charge at certain things to get much of anything done not be swept up in the tide of whatever the expectations are there's like one final thing I want to add which is like we talked a lot about agency and this kind of stuff but I think actually like surprisingly enough one of the most important things, you're just caring an unbelievable amount.
2:02:21And when you care an unbelievable amount, you check all the details and you have this understanding of what could have gone wrong. It matters more than you think because people end up not caring enough. This is like LeBronquot where he talks about how, before he sat in the league, he was worried that everyone would be incredibly good. And then he gets there and he realizes that actually once people hit financial stability, then they relax a bit. And he's like, oh, it's going to be easy. I don't think that's quite true, because I think in AI research, because most people actually care quite deeply.
2:02:57But there's caring about your problem, and there's also just caring about the entire stack and everything that goes up and down, like going explicitly going and fixing things that aren't your responsibility to fix, because overall it makes the stack better. I mean, another part that I forgot to mention is you were mentioning, oh, going in on weekends and on Christmas break and you get to like the only people in the office are Jeff Dean and Sergey Brand or something. And you just like get to pay a program with them. It's just, it's interesting to me that people, I don't want to pick on your company in particular, but like a people at any big company, they've gone there because they've gone through a very selective process that's like they had to compete in high school, they I got to compete in college, but it almost seems like they get there and then they take it easy.
2:03:41When in fact, this is a time to put the pedal to the metal, go in and pair for a program with Sergey Vrenon the weekends or whatever, you know what I mean? I mean, there's pro -Send cons there, right? I think many people make the decision that the thing that they wanna prioritize is like a wonderful life with their family. And if they do wonderful work, like let's say they don't work every other day, right? But they do wonderful work in the work like the hours that they do do. That's incredibly impactful. I think this is true for many people at Google, is like maybe they don't work as many hours as typical startup mythologies, right?
2:04:11But the work that they do is incredibly valuable. It's very high leverage because they know the systems and their experts in their field. And we also need people like that. Like our world rests on these huge, like difficult to manage and difficult to fix systems. And we need people who are willing to work on and help and fix and maintain those, in frankly, like a thankless way that isn't as like, I publicity is all of this AI work that we're doing, right? And I'm like ridiculously grateful that those people do it. And I'm also happy that there are people for whom like, okay, they find technical fulfillment in their job and doing that well.
2:04:44And also like maybe they draw a lot more from one also spending like a lot of hours with their family. And I'm lucky that I'm at a stage in my life where like, Jack can go in and work every hour of the week. But like, that's like, I'm not making as many sacrifices to do that. Yeah. I mean, like, just one example of the six out of mine, mind of this sort of like the other side says no, and you can still get the yes on the other end. Basically, every single high profile of guest have gone so far. I think maybe with one or two exceptions, I've sat down for a week, and I've just come up with a list of sample questions that you know, like try to really come up with really smart questions to say to them.
2:05:23And the entire process I've always thought like, If I just cold email them, it's like a 2 % chance, they say yes, if I include this list, there's a 10 % chance. And because otherwise, you know, there's like, you go through their inbox and every 34 seconds, there's an interview for whatever podcast, interview whatever podcast. And every single time I'm done this, they've said yes. Right, you just like, you do. You do ask great questions. But if you do, everything you'll win. But you just like, you literally have your dig in the same hole for like 10 minutes, or in that case, like make a sample, listen sample questions for them to get past or not an idiot list.
2:05:57Or you're not a bean. I just demonstrate how much you care. Yeah, yeah. And the work you're going to put in. Yeah, yeah. Something that a friend said to me a while back, but I think it's stuck is like, it's amazing how quickly you can become world class at something, just because most people aren't trying that hard. And like are only working like, I don't know, the actual like 20 hours that they're actually spending on this thing or something. And so yeah, if you just go ham, then like you can get really far pretty fast. And I think I'm lucky I had that experience with the fencing as well Like I had the experience of becoming world's class in something and like knowing the fuchs worked really really hard and Would like yeah for a context by the way Shoto was one seed away as he was an expert in in line to go to The Olympics to have for fencing.
2:06:42I was at best like 42nd in the world for fencing so that's well fencing And you didn't know load as a thing man And there was one cycle where, yeah, I was like the next highest rank person in Asia and if one of the teams had been disqualified for doping as it was occurring, in fact during that cycle and as occurred for like the Australian women's rowing team I think went because one of the teams was disqualified, then I would have been the next in line. It's interesting when you just find out people's prior lives and it's like, oh, this guy was almost an Olympian, this other guy was whatever, you know what I mean?
2:07:24Okay, let's talk about intermibility. I actually want to stay on the brain stuff as a way to get into it for a second. We're previously discussing, is the brain organized in the way where you have a residual a stream that is gradually refined with higher level associations over time or something. There's a fixed dimension size in a model. If you had to, I don't even know how to ask this question in a sensible way, but what is the D model of the brain? What is the embedding size of, or because a feature splitting is that not a sensible question? No, I think it's a sensible question. Well, it is a question.
2:08:09No. Who does not say that? No, I just don't know. You can talk. I'm trying to... I don't know how you would begin to kind of be like, okay, well, this part of the brain is like a vector of this dimensionality. I mean, maybe for the visual stream, because it's like V1 to V2 to IT, whatever. You could just count the number of neurons that are there and be like, that is the dimensionality, but it seems more likely that there are kind of submodules and things are divided up. So yeah, I don't have, and I'm not like the world's greatest neuroscientist, right? Like I did it for a few years. I like study the cerebellum quite a bit.
2:08:50So I'm sure there are people who could give you a better answer on this.
2:08:55Do you think that the way to think about whether it's in the brain or whether it's in these models? fundamentally what's happening is like features are added and removed, changed, and like the feature is the fundamental unit of what is happening in the model. Like what would have to be true for, give me a, and this goes back to the earlier thing we were talking about whether it's just associations all the way down. Give me like a counterfactual in the world where this is not true, what is happening instead? Like what is the alternative hypothesis here? Yeah, it's hard for me to think about because at this point, I just just thinks so much in terms of this feature space.
2:09:36I mean, at one point there was the kind of behavioral or a list approach towards cognition, where it's like you're just, you're like input output, but you're not really doing any processing, or it's like everything is embodied, and you're just like a dynamical system that's like operating along like some predictable equations, but there's no state in the system, I guess. But whenever I've read these sorts of critiques, it's like, well, you're just choosing to not call this thing a state, but you could call any internal component of the model a state. Even with the future discussion, defining what a feature is is really hard.
2:10:18And so the question feels almost too slippery. What is a feature? A direction and activation space. A latent variable that is operating behind the scenes that has causal influence over the system you're observing. It's a feature if you call it a feature. It's tonological. These are all explanations that I feel some associated. In a very rough, intuitive sense, it's like a sufficiently sparse, like binary. It's like the features like whether or not something is turned on or off. right? Right? Like in a very simplistic sense. Yeah. Yeah. Which might be I think a useful metaphor to understand it by.
2:11:01It's like when we talk about features activating, it is in many respects the same way that neuroscience would talk about like a neuron activating, right? If that neuron corresponds to something in particular. Right. Yes. Yeah. And no, I think that's useful as like what do we want a feature to be? Right? Like what is a synthetic problem under which a feature exists? But even with the towards monosomaticity work, we talk about what's called feature splitting, which is basically, you will find as many features as you give the model the capacity to learn. And by model here, I mean the up projection that we fit after we trained the original model.
2:11:38And so if you don't give it much capacity, it'll learn a feature for bird. But if you give it more capacity, then it will learn like ravens and eagles and sparrows and like specific types of birds. words. Still on definition's thing. I guess, and not evenly, I think of things like bird, versus what kind of token is it like a period at the end of the hyperlink as your time went earlier, versus at the highest level things like love or deception or like holding a very complicated proof in your head or something. Is this all features? Because then the the definition seems so broad as to almost be not that useful.
2:12:23Rather that there seems to be some important differences between these things and they're all features, like, yeah, I'm not sure what we would mean by. I mean, all of those things are like discrete units that have connections to other things that then abuse them with meaning. That feels like a specific enough definition that it's useful or not to all encompassing, but feel free to push back. What would you discover tomorrow in that could make you think, oh, this is kind of fundamental the wrong way to think about what's happening in a model? I mean, if the features we were finding weren't predictive or if they were just representations of the data, right, where it's like, oh, all you're doing is just clustering your data.
2:13:11And there's no higher level associations that are being made. or it's some like phenomenological thing of like, you're saying that this feature fires for marriage, but if you activate it really strongly, it doesn't change the outputs of the model on a way that would correspond to it. Like I think these would both be good critiques. I guess one more is, and we tried to do experiments on MNIST, which is a data set of digits, images. And we didn't look super hard into it, and so I'd be interested if people, other people wanted to take up like a deeper investigation, But it's plausible that your latent space of representations is dense, and it's a manifold, instead of being these discrete points.
2:13:54And so you could move across the manifold, but at every point, there would be some meaningful behavior. And it's much harder than to label things as features that are discrete. And like in naive sort of outsider way, the thing that would seem to me to be like, like a way in which this picture could be wrong is if there's not some like this thing is turned on turn off, but it's like a much more global kind of like the system is I'm going to be just really clumsy like you know I measured it in a variety of kind of language, but is there a good analogy here? Yeah, I guess if you think of like something like the laws of physics, it's not like well the feature for wetness is turned on, but it's only turned on this much.
2:14:43And then the feature for like, you know, I guess maybe it's true, because like the mass is like a gradient and like, you know, like, I don't know, but the polarity or whatever is a gradient as well. But there's also a sense in which like there's the laws and the laws are more general and you have to understand like the general bigger picture and you don't get that from just like these like specific sub, sub circuit. But that's where the reasoning circuit itself comes into play, where you're taking these features, ideally, and trying to compose them into something higher level. You might say, OK, when I'm using, at least this is my head cannon, let's say I'm trying to use the foot, F equals M A.
2:15:22Then I'm presumably at some point, to have features which do note, OK, like what mass, and then that's helping me retrieve the actual mass of the thing that I'm using, and then the acceleration, and this kind of stuff. But then also, maybe there's a higher level feature that does correspond to using the first law of physics. Maybe, but the more important part is that the composition of components which helps retrieve piece of relevant piece of information and then produce like maybe something like a multiplication operator or something like that, from when necessary. At least that's my head, Canon.
2:15:50What is a compelling explanation to you, especially for very smart models of, like I understand why it made this output and it was like for a legit reason, if it's doing million line pull requests or something, What are you seeing at the end of that request where you're like, yep, should that show? Yeah, so ideally, you apply dictionary learning to the model. You've found features. Right now, we're actively trying to get the same success for attention heads, in which case, we have features for both the core. You can do it for residual stream, MLP, and attention throughout the whole model. Hopefully, at that point, you can also identify broader circuits through the model that are like more general reasoning abilities that will activate or not activate.
2:16:34But in your case where we're trying to figure out if this like, pull or crash would be approved or not, I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired. That would be like an immediate, you can do more than that, but that would be an immediate. But before I trace down on that, what is the reasoning circuit look, like what would that look like when you found it? Yeah, so I mean, the induction head is probably one of the simplest. That's not like a reasoning, right? Well, I mean, what do you call reasoning, right?
2:17:05Like, it's a good reason. So I guess context for listeners. The induction has basically, and you see the line, like Mr. and Mrs. Dersley did something, Mr. Blank, and you're trying to predict what Blank is. And the head has learned to look for previous occurrences of the word Mr. Look at the word that comes after it, and then copy and paste that as the prediction for what should come next. which is a super reasonable thing to do. And there is computation being done there to accurately predict the next token. Mm -hmm. Yeah. That is context dependent. That is, yeah. But it's not like, it's not like reasoning, you know what I mean?
2:17:47But is, I guess going back to the like associations all the way down, it's like, if you chain together a bunch of these reasoning circuits or heads that have different rules for how to relate information. But in this sort of zero shot case, something is happening where when you pick up a new game and you immediately start understanding how to play it, and it doesn't seem like an induction -hits kind of thing. Or I think there would be another circuit for extracting pixels and turning them into latent representations of the different objects in the game. And a circuit that is learning physics. And what would that because the induction heads is like one layer transformer So you can like kind of see like what like the thing that is a human picks up a new game and understands it How like how do you think about what that is?
2:18:39Is it a presumably across multiple layers, but like Is it yeah? Yeah, how like what would that physically look like? How big would it be maybe or like I mean, that would just be an empirical question, right? Of like, how big does the model need to be to perform this task? But like, maybe it's useful if I just talk about some other circuits that we've seen. So we've seen like the I -O -I circuit, which is the indirect object identification. And so this is like, if you see it, it's like, Mary and Jim went to the store, Jim gave the object to blank, right? And it would predict Mary because Mary's appeared before.
2:19:15as like the indirect object or it'll infer pronouns, right? And this circuit even has behavior where like if you ablate it, then like other heads in the model will pick up that behavior. We'll even find heads that wanna do copying behavior and then other heads will suppress. So like it's one job, one head's job to just always copy like the token that came before, for example, or the token that came five before or whatever. And then it's another head's job to be like, no, do not copy that thing. So there are lots of different circuits performing in these cases, pretty basic operations. But when they're chained together, you can get unique behaviors.
2:19:59And by the like, is the story of how you found it with the reasoning thing is like, because you won't be able to understand or I don't just be like really kind of, you know, it won't be something you can see in like a two layer transformer. So will you just be like, the circuit for a deception or whatever. or just this part of the network fire when we at the end identified the thing is being deceptive. And it didn't fire when we did not defend it as being deceptive. Therefore, this must be the deception circuit. I think a lot of analysis like that. Anthropic has done quite a bit of research before on psycho -fancy, which is like the model saying what it thinks you want to hear.
2:20:35And that requires us at the end to be able to label which one is bad and which one is good. Yeah. So we have tons of instances and actually, as you make models larger, they do more of this. Where the model is clearly, it has features that model another person's mind. And these activate and some subset of these, we're hypothesizing here, but would be associated with more deceptive behavior. Although it's doing that by, I'm going to chat GPT, I think it's probably modeling me because that's like RLHF and uses it to. Yeah, seriously. Yeah, so well, first of all, the thing you mentioned earlier about, there's redundancy.
2:21:15So then it's like, well, have you caught like the whole thing that could cause a session of the whole thing or like is just one instance of it? Yeah. Second of all, are your like labels correct? You know, maybe like you thought this wasn't deceptive. It's like, so deceptive. Especially if it's producing output, you can't understand. Third is the thing that's going to be the bad outcome. Something that's even who an understandable, like the session is a concept we can understand. and maybe there's like a, yeah, yeah, so a lot's unpacked here. So I guess a few things. One, it's fantastic that these models are deterministic.
2:21:47When you sample from them, it's stochastic, right? But like, I can just keep putting in more inputs and a blade every single part of the model. This is kind of the pitch for computational neurosciences to come and work on interpretability. It's like, you have this alien brain and you have access to everything in it, and you can just oblate however much of it you want. And so I think if you do this carefully enough, you really can start to pin down what are the circuits involved, what are the backup circuits, these sorts of things. The kind of cop -out answer here, but it's important to keep in mind, is doing automated interpretability.
2:22:16So it's like as our models continue to get more capable, having them assign labels, or like run some of these experiments at scale. And then with respect to like, if there's superhuman performance, how do you detect it? Which I think was kind of the last part of your question. Aside from the cop -out answer, if we buy this associations all the way down, you should be able to coarse -graine the representations at a certain level such that they then make sense. I think it was even in Demis' podcast, he's talking about like if a chess player makes a superhuman move, they should be able to distill it into reasons why they did it.
2:22:52And like even if the model's not gonna tell you what it is, you should be able to decompose that complex behavior into simpler circuits or features to really start to make sense of why it did the thing that it did. There's a separate question of does such representation exist, which it seems like they're must, or actually I'm not sure if that's the case. And secondly, whether using this sparse writing coder setup, you could find it. And in this case, if you don't have labels for it that are adequate to represent it, like you wouldn't find it, right? Yes and no. So like we are actively trying to use dictionary learning now on the sleeper agents work, which we talked about earlier.
2:23:35And it's like, if I just give you a model, can you tell me if there's this trigger in it and it's going to start doing interesting behavior? And it's an open question whether or not when it learns that behavior, it's part of a more general circuit that we can pick up on without actually getting activations for and having it display that behavior, because that would kind of be cheating then. Or if it's learning some hacky trick over, that's a separate circuit that you'll only pick up on if you actually have it do that behavior. But even in that case, the geometry of features gets really interesting because fundamentally, each feature is in some part of your representation space and they all exist with respect to each other.
2:24:16And so in order to have this new behavior, you need to carve out some subset of the feature space for the new behavior and then push everything else out of the way to make space for it. So hypothetically, you can imagine you have your model before you've taught at this bad behavior, you know all the features or have some coarse grain representation of them. You then fine tune it, such that it becomes malicious. And then you can kind of identify this black hole region of feature space, where everything else has been shifted away from it. And there's this region, and you haven't put in an input that causes it to fire.
2:24:47But then you can start searching for what is the input that would cause this part of the space to fire, what happens if I activate something in this space. There are a whole bunch of other ways is that you can try and attack that problem. This is a sort of a tangent. But one interesting idea I heard was, if that space is shared between models, you can imagine trying to find it in an open source model to then make, like Gemma is, they said in the paper, Gemma by the way, Google's newly released open source model, they said in the paper it's trained using the same architecture or something like that.
2:25:19I had to be honest, I didn't know because I had to have a Gemma paper. Very similar method, this is something whatever, as Gemini. So it's a thing that's true. I don't know how much like how much of the direct team you do on Gemma is like potentially helping you jailbreak into Gemini. Yeah, this gets into the fun space of like how universal or features across models and and our towards monosimenticity paper looked at this a bit and we find I can't give you summary statistics but like the base 64 feature for example which we see across a ton of models. This is like it's actually three of them but they'll fire four and model base 64 or encoded text, which is prevalent in like every URL, and there are lots of URLs in the training data, they have really high cosine similarity across models.
2:26:01So they all learn this feature. And I mean, within a rotation, right? But it's like, yeah, yeah. Like actually like vectors, so. Yeah, yeah. And I wasn't part of this analysis, but yeah, it definitely finds the feature and they're like pretty similar to each other across two models, the same model architecture, but trained with different random seats. It supports the quantity of neural scaling. It's like hypothesis, right? We just look like all models on like a similar data set, we'll learn the same features in the same order, ish roughly like you learn your Ngrams, you learn your induction heads, and you learn like the put full stops after numbered lines.
2:26:34And that's kind of stuff. Hey, but by the way, okay, so this is another tangent. To the extent that that's true, and like I guess there's evidence of this true, why doesn't curriculum learning work? Because if it is the case that you learn certain things first, shouldn't just directly training those things first lead to better results? Both Gemini papers mention some like aspects of curriculum learning. Okay, interesting. I mean, the fact that fine -tuning works is like evidence or curriculum learning, right? Because like the last things you're turning on have a disproportionate impact. I wouldn't necessarily say that.
2:27:02Like there's one mode of thinking which fine -tuning is specialized. Like you've got this like latent bundle capabilities, you know, like specialized before. It's particular. Like use case that you want. I think I'm not sure how true it is. I think the diva ball lab kind of paper kind of supports this, right? Like you have that ability and you're just getting better at entity recognition. Like fine -tuning that circuit instead of other ones. So what was the thing we're talking about more? But generally, I do think curriculum learning is really interesting to me for people to explore more. And it seems very pleasant.
2:27:29I would really love to see more analysis along the lines of the quantitative stuff and understanding better what do you actually learn at each stage and decomposing that out and exploring whether or not curricula change that. But by the way, I just realized we just got in conversation and forgot there's an audience. Curriculum learning is when you organize a data set, when you think about a human, how they learn, they don't just see random wiki text and they just try to predict it, right? They're like, we'll start you off with like, a lore act or something and then you'll learn. I don't remember what first grade was like, but you'll learn the things that first grade was learned and then like second grade was and so forth.
2:28:08And so you do, imagine that. Sorry, we never got past first grade.
2:28:14Aquí frozen into the bite!
2:28:24Not so cold because защ来物 watch to say audio from secret services. Oke, anyways. Uh let's get back to like the big far we get into like of momentos like door to space details. The big picture, there's two threads I want to explore. First is, I guess it makes me a little worried that there's not even an alternative formulation of what could be happening in these models that could invalidate this approach, which feels like, I mean, we do know that we don't understand intelligence, right? There are definitely unknown unknowns here. So, like the fact that there's not a null hypothesis, I don't know if you like, but if we're just wrong and we don't even know the way in which we're wrong, which actually increases is the uncertainty.
2:29:03And yeah, yeah, yeah, yeah. So it's not that there aren't other hypotheses. It's just I have been working on superposition for like a number of years and very involved in this effort. And so I'm less sympathetic to or will I just have their rock? Like to these other approaches, especially because our recent work has been so successful. And like quite high explanatory power. Like, in the scaling lowest paper, there's this little bump at a particular, like the original scaling lowest paper is a little bump. And that, apparently, corresponds to when the model learns induction heads. And then, like, after that, it goes off track, learns induction heads, gets back on track, which is an incredible piece of retroactive explanatory power.
2:29:48Before I forget it, though, I do have one thread on feature universality that you might want to have in. So there are some really interesting behavioral evolutionary biology experiments on like, should humans learn of real representation of the world or not? You can imagine a world in which we saw all venomous animals as like flashing neon pink, a world in which we survive better. And so it would make sense for us to not have a realistic representation of the world. And there's some work where they'll simulate like little basic agents and see if the representations they learn, like map to the tools they can use and the inputs they should have.
2:30:28And it turns out if you have these little agents perform more than a certain number of tasks given these basic tools and objects in the world, then they will learn a ground truth representation because there are so many possible use cases that you need for these based objects that you actually want to learn what the object actually is and not some cheap visual heuristic or other thing. And so to the extent that we are doing, and we haven't talked at all about like, for instance, free energy principle or predictive coding or anything else, but like to the extent that all living organisms are trying to like actively predict what comes next and form like a really accurate world model, it wouldn't surprise me or I'm optimistic that we are learning genuine features about the world that are good for modeling it, and our language models will do the same.
2:31:17At least especially because we're training them on human data and human text. Another dinner party question. Isn't should we be less worried about misalignment and maybe that's even the right word for what I'm referring to, but just alienness and chagotness from these models, given that there is feature universality and there are certain ways of thinking and ways of understanding the world that are instrumentally useful to different kinds of intelligences. So we just be less worried about like bizzaro paper with maximizers as a result. I think that's the, this is kind of why I bring this up is like the optimistic take.
2:31:56Predicting the internet is very different from what we're doing, right? Like the models are way better at predicting next tokens than we are. They're trained on so much garbage. They're trained on so many URLs. Like in the dictionary learning work, we find there are like three separate features for base 64 in coatings. And like even that is kind of an alien example that is probably worth me talking about for a minute. One of these base 64 features fired for numbers. Like other base 64, like if it sees base 64 numbers, it'll like predict more of those. Another fired for letters. But then there was this third one that we didn't understand.
2:32:30And it like fired for like a very specific subset of base 64 features. And someone on the team who clearly knows way too much about base 64 realized that this was the subset that was ASCII decodable. So you could coat it back into the ASCII characters. And the fact that the model learned these three different features and it took us a little while to figure out what was going on, was very shogoth -esque. That it has a denser representation of regions that are particularly relevant to predicting the next token. Yeah, because it's so, but yeah, and it's clearly doing something that humans wouldn't.
2:33:06You can even talk to any of the current models in base 64 and it will apply in base 64, and you can then decode it and it works great. That particular example, I wonder if that implies that the difficulty of doing interpability on smarter models will be harder, because if it requires somebody with esoteric knowledge that just happened to see that base 64 has, I don't know, whatever that distinction was, doesn't imply when you have the million line pull requests, it's like there is no that's going to be able to decode two different reasons why the pull request, there's two different features for this pull request.
2:33:43Yeah, you know what I mean? Yeah, yeah. So if you think of any type of comment, like small CLs, please. Like the mature. Yeah, exactly. No, no, I mean, you could do that, right? This is what I was going to say is one technique here is anomaly detection. And so one beauty of dictionary learning instead of linear probes is that it's unsupervised. You are just trying to learn to span all of the representations that the model has. and then interpret them later. But if there's a weird feature that suddenly fires for the first time that you haven't seen fire before, that's a red flag. You could also coarse -grained it so that it's just a single -base 64 feature.
2:34:17I mean, even the fact that this came up and we could see that it specifically favors these particular outputs and it fires for these particular inputs, gets you a lot of the way there. I'm even familiar with cases from the auto -interp side where a human will look at a feature and try to annotate it for. or it fires for Latin words. And then when you ask the model to classify it, it says it fires for Latin words defining plants. So it can already beat the human in some cases for labeling what's going on. So at scale, this would require an adversarial thing between models where some model, you have millions of features potentially for GPT -6.
2:34:59And it's just a bunch of models or just trying to figure out what each of these features means. How is this not right? OK. But you can even automate this process. I mean, this goes back to the determinism of the model. Like, you could have a model that is actively editing input text and predicting if the feature is going to fire or not and figure out what makes it fire, what doesn't, and like search the space. Yeah. I want to talk more about the features, because I think that's like an interesting thing that has been under Explorer. Especially for scalability. I think it's it's un -appreciated right?
2:35:31First of all, how do we even think about is it really just you can keep going down and down? There's no end to the amount of features. I mean, so at some point, I think you might just start fitting noise or things that are part of the data, but that the model isn't actually right. But do you want to explain what feature splitting is? Yeah, yeah. So it's the part before where like the model will learn however many features it has capacity for, that still span the space of representation. So they give an example potentially. Yeah, so if you don't give them out of that much capacity for the features it's learning, concretely, if you project to not as high a dimensional space, we'll learn one feature for birds.
2:36:13But if you give them out of more capacity, it will learn features for all the different types of birds. And so it's more specific than otherwise. And oftentimes, there's the bird vector that points in one direction, and all the other specific types of birds point in a similar region of the space, but are obviously more specific than the course label. Okay, so let's go to back to GBD7. First of all, is this a linear text on any model to figure out? Is this the one time thing you have to do, or is this the kind of thing you have to do on every output? or just like one time it's not deceptive, we're good to get the role.
2:36:52Actually, yeah, let me let you answer that. Yeah, so you do dictionary learning after you've trained your model and you feed it a ton of inputs and you get the activations from those and then you do this projection into the higher dimensional space. And so the method is unsupervised and that it's trying to learn these sparse features, you're not telling them in advance what they should be, but it is constrained by the inputs you're giving the model. I guess two caveats here, one, like we can try and choose what inputs we want. So if we're looking for theory of mind features that might lead to deception, we can put in the sake of fancy dataset.
2:37:27Hopefully at some point we can move into looking at the weights of the model alone, or at least using that information to do dictionary learning. But I think in order to get there, that's such a hard problem that you need to make traction on just learning what the features are first. But yeah, so what's the cost of this? Can you read the lessons? Weathe the model alone. So right now we just have these neurons in the bowl. They don't make any sense. We apply dictionary learning. We get these features out. They start to make sense. But that depends on the activations of the neurons. The weights of the model itself, what neurons are connected to what other neurons certainly has information in it.
2:38:07And the dream is that we can kind of bootstrap towards actually making sense of the weights of the model that are independent of the activations of the data. I mean, this is, I'm not saying we've made any progress here. It's a very hard problem, but it feels like we'll have a lot more traction and be able to like, sanity check what we're finding with the weights if we're able to pull out features first. For the audience, weights are permanent, I don't know if it's permanent, sorry, word, but like they are the model itself, whereas activations are the sort of like artifacts of any single call. Mm -hmm.
2:38:38Yes. In a brain metaphor, you know, the weights, like the actual connection scheme between neurons and the activations of it, current neurons of the line. Yeah, so there's going to be two steps to this for GPT -7 or whatever model we're concerned about. Actually, first correct me if I'm wrong, but like training the Sparrow Auto encoder and do the unsupervised projection into a wider space of features that have a higher fidelity to what is actually happening in the model. and then secondly label those features. Because let's say like the cost of training the model is n. What will those two steps cost relative to n?
2:39:19We will see, like it really depends on two main things, what is your expansion factors, like how much are you projecting into the higher dimensional space, and how much data do you need to put into the model, how many activations do you need to give it? But this brings me back to the feature splitting to a certain extent, because if you know you're looking for specific features, you can start with a really cheaper course representation. So maybe my expansion factor is only two. So I have a thousand neurons, I'm projecting to a 2 ,000 dimensional space. I get 2 ,000 features out, but they're really coarse.
2:39:51And so previously I had the example for birds. Let's move that example to I have a biology feature. But I really care about if the model has representations for bioethans and is trying to manufacture them. And so what I actually want is an anthrax feature. What you can then do is rather than, and let's say you only see the anthrax feature, if instead of going from 1 ,000 dimensions to 2 ,000 dimensions, I go to a million dimensions. And so you can kind of imagine this big tree of semantic concepts where biology splits into cells versus whole body biology, and then further down it splits into all these other things.
2:40:28So rather than needing to immediately go from 1 ,000 into a million and then picking out that one feature of interest, you can find the direction that the biology feature is pointing in, which again is very coarse, and then selectively search around that space. So like only do dictionary learning if this, if something in the direction of the biology feature fires first. And so the computer science metaphor here would be like instead of doing breadth first search, you're able to do depth first search where you're only recursively expanding and exploring a particular part of this like semantic tree of features.
2:41:04I was given the way that these features are not organized in things that are intuitive for humans, right? Like, because we just don't have to deal with basics before. So we don't have that many, you know, we just don't dedicate that much like whatever firmware to like deconstructing which kind of basically where it is. How would we know that the subjects and this will go back to maybe the M .O .E. discussion we'll have of, I guess we might as well talk about it. like in mixture of experts, the mixture of paper talked about how they couldn't find, the experts weren't specialized in the way that we could understand.
2:41:38There's not like a chemistry expert or a physics expert or something. So why would you think that it'll be like biology feature and then deconstruct rather than like blah, and then you just deconstruct and it's like anthrax and you're like shoes and whatever. So I haven't read the the Mistral paper, but I think that the heads, I mean, this goes back to like, If you just look at the neurons in a model, they're polysmantic. And so if all they did was just look at the neurons in a given head, it's very plausible that it's also a polysmantic because of superposition. And so, I'm going to thread the door case range in there.
2:42:11Have you seen in the subtrees when you expand them out, like something in a subtreet, which like you really wouldn't guess that it would should be there based on like the higher level of extraction? So this is a line of work that we haven't pursued as much as I want to yet. But I think we're planning to, I hope that maybe external groups do as well. But what is the geometry of features? What's the geometry of the, exactly? How does that change over time? It would really suck if the anthrax feature happened to be below the coffee can. It's some tree, something like that. It's totally totally. And that feels like the kind of thing that you could quickly try and find proof of, which would then mean that you need to then solve that problem.
2:42:48Yeah, and inject more structure into the geometry. Totally. It would really surprise me, I guess, especially given how linear the ball seems to be. Completely agree. That there isn't some component of the anthrax feature. like Vector that is similar to and looks like the biology vector and that they're not in a similar part of the space But yes, I mean ultimately Machine learning is empirical. Yeah, we need to do this I think it's pretty important for certain aspects of scaling dictionary. Yeah, yeah, interesting On the M O E discussion. Yeah, there's an interesting scaling vision transforms paper the Google put out a little while ago Where they like do image net classification with a like an M O E and they find really clear class specialization there.
2:43:27For experts, there's a clear dog expert. But the mixed -through people just not do a good job of identifying different. I think it's hard. And it's entirely possible that in some respects, there's almost no reason that all of the different archives features should go to one expert. You could have biology. Let's say I don't know what buckets they had in their paper, but let's say they had archive papers as one of the things. You can imagine biology papers going here, math papers going here, and all of a sudden, you know, right down is ruined. But that vision transformal one, where the class separation is really clear.
2:44:03And obviously, I think some evidence towards the specialization of hypothesis. So I think images are also in some ways just easier to interpret than text. Yeah, exactly. And like so Chris, all those interpretability work on AlexNet and these other models. Like in the original AlexNet paper, they actually split the model into two GPUs just because that GPUs were so bad back then, I'm relatively speaking, right? Like still great at the time. That was one of the big innovations of the paper. But they find branch specialization and there's a distilled pop article on this where like colors go to one GPU and like, goblore filters and like line detectors go to the other.
2:44:43And then like all of the other. Really? Yeah, yeah, yeah. And then all of the other interpretability work that was done, like the floppy ear detector, that just was a neuron in the model that you can make sense of. You didn't need to disentangle it too much. So just different data set, different modality. I think a wonderful research project to do, if someone is out there listening to this, would be to try and disentangle, take some of the techniques that Trenton's team has worked on, and try and disentangle the neurons in the mixture model, which is like, I think that's a fantastic thing because it feels intuitively like they should be.
2:45:20They didn't demonstrate any evidence that there is. There's also, like in general, a lot of evidence that they should be specialization. Go and see if you can find it. And that's that's work that has, that, you know, anthropocase published most of this stuff on, like, as I understand it, like, dense models basically. That is a wonderful research product to try. And given DoorCash's success with the Vesuvius challenge, we should be pitching more projects because they will be sold to for example in the podcast. What I was thinking about after the interview challenge was like, wait, I knew, like Nata told me about it before it dropped because we recorded the episode before it dropped.
2:45:55Why didn't you, why didn't I not even try it? Like, you know what I mean? Like, I don't know. Luke is obviously very smart and like, yeah, he's an amazing kid, but like, you showed that like a 21 year old on like some 10, 70 or whatever he was working on could do this. I don't know, like I feel like I should have, So before this episode drops, I'm gonna need my... I'm gonna try to make an interpreter. I don't know, I'm gonna try to do a research. I was like, honestly, I think the background experience was like, wait, I shouldn't have been like, what happened to the new fuck? Yeah, the hands dirty.
2:46:27Dual -cash is a request for research.
2:46:32I want to harp back on this, like, the neuron thing. You said, I think a bunch of your papers have said, there's more features than there are neurons. and this is just like, wait a second. I don't know, like a neuron is like, waits going and a number comes out. That's like, a number comes out, you know what I mean? That's so little information. Do you mean like there's street names and species and whatever? There's more of those kinds of things than there are a number comes out in a model. That's right, yeah. But how is a number comes out as like so little information? How is that encoding for like superposition?
2:47:10You're just in the building, you're in the building. That's one of the features in these high -dimensional vectors. In a brain, is there like an exonl -firing, or how are you thinking about it? Like, I don't know how you think about like, how much like superposition is there in the human brain? Yeah, so Bruno Olshausen, who I think of as the leading expert on this, thinks that all the brain regions you don't hear about are doing a ton of computation in superposition. So everyone talks about V1 as having good bore filters and detecting lines of various sorts. And no one talks about V2. And I think it's because we just haven't been able to make sense of it.
2:47:47What is V2? It's the next part of the visual processing stream. And it's like, yeah, so I think it's very likely. And fundamentally, superposition seems to emerge when you have high dimensional data that is sparse. And to the extent that you think the real world is that, which I would argue it is, we should expect the brain to also be under parameterized in trying to build a model of the world, and also use super possession. You can get a good intuition for this, and correct me, just like this example, is wrong in like a 2D plane, right? Let's say you have like two axes, right? Which represents like a two -dimensional, like feature space here, like two neurons basically.
2:48:22And you can imagine them each like turning on to various degrees, right? And that's like your X -Coin and your Y -Coin, it. But you can like, now like map this onto a plane, you can actually represent a lot of different things and like different parts of the plane. Oh, okay. So crucially, the superposition is not an artifact of a neuron. It is an artifact of like the space that is created. It's a territorial code. Yeah, yeah, exactly. Yeah. Okay, cool. Um, yeah, thanks. I mean, we kind of talked about this, but like I think it just like kind of wild that it seems to the best of our knowledge the way intelligence works in these models and then presumably also in brains.
2:49:00It's just like there's a stream of information going through that has quote -unquote features that are infinitely or at least to a large extent just like splitable and you can expand out a tree of like what this feature is and what's really happening is a stream like that feature is getting turned into this other feature or this other feature is added. I don't know. It's like that's not something I would have just like thought like that's what intelligence is. You know what I mean? It's like a surprising thing. It's not it's not whatever it would have expected necessarily. What did you think it was?
2:49:35I don't know, man. I mean, yeah. Go fight. So that's a great subject because all of this feels like GoFy. Like you're using distributed representations, but you have features and you're applying these operations to the features. I mean, the whole field of vector symbolic architectures, which is this computational neuroscience thing, it all you do is you put vectors in superposition. and which is literally a summation of two high -dimensional vectors. And you create some interference, but if it's high -dimensional enough, then you can represent them. And you have variable binding, or you connect one by another, and if you're doing with binary vectors, it's just the x -or operation.
2:50:14So you have AB, you bind them together, and then if you query with A or B again, you get out the other one. And this is basically the key value pairs from attention. and with these two operations, you have a turn complete system, which if you have enough nested hierarchy, you can represent any data structure you want, et cetera, et cetera. Yeah. Okay, let's go back to the super intelligence. So like walk me through GPT -7, you've got like the sort of depth first search on its features. Okay, GPT -7 has been trained. What happens next? Your research has succeeded. GB7 has been trained. What are we doing now?
2:50:58We try and get it to do as much interpretability work and other safety work as possible. Like, concrete. What has happened such that you're like, cool, let's deploy a GB7. Oh, geez. I mean, we have our responsible scaling policy, which has been really exciting to see other labs adopt. And, like, this is only from the perspective of your research and that like I trained and given you a research you got the thumbs up on GPT -7 from you or actually which is like cloud whatever and then oh I like what is the basis on which you're telling the team like hey let's go ahead. I mean I think we need to make a lot more if it's as capable as GPT -7 like implies here I think we need to make a lot more interpretability progress to be able to like comfortably give the green light to deploy it.
2:51:47Like what would you like definitely not? I'd be crying. name like tears would interfere with the GPUs. What is? Yes. Gemini 5, TV is back.
2:52:06But like what, what, what, given the way your research is progressing like, what does it kind of look like to you? Like, what would, if this exceeded, what would it mean for us to okay, GPT -7 based on your methodology? I mean, ideally, we can find some compelling deception circuit, which lights up when the model knows that it's not telling the full truth to you. Why can you just do your linear probe, like Colin Bird's did? So the CCS work is not looking good in terms of replicating or like actually finding truth's directions. And like, in hindsight, it's like, well, why should it have worked so well.
2:52:42But linear probes, you need to know what you're looking for, and it's like a high dimensional space, and it's really easy to pick up on a direction that's just not. Wait, but don't you also, here you need to label the features. So you sort of know. You need to label and post -hoc, but it's unsupervised. You're just like, give me the features that explain your behavior is the fundamental question, right? It's like, like the actual setup is, we take the activations, we project them to this higher dimensional space, and then we project them back down again. So it's like reconstruct or do the thing that you were originally doing, but do it in a way that's sparse.
2:53:14By the way for the audience, linear probe is you just like classify the activations. I don't know from what I vaguely remember about the paper was like if it's like telling a lie, then you like you just train a classifier on like is it? Yeah, but in the end it was it was it a lie or is it just like wrong or something? I don't know. It was like true or false question. Classifier on the actuation. So yeah, like right now what we do for GPT -7, ideally we have like some deception circuit that we've identified that appears to be really robust. And it's like, Well, so you've done the projecting out to the million whatever features or something.
2:53:55Is the circuit, because maybe we're using feature in circuit interchangeably when they're not. So is there like a deception circuit? So I think there's one to fear. There are features across layers that create a circuit. And hopefully the circuit gives you a lot more specificity and sensitivity than an individual feature. And it's like, hopefully we can find a circuit that is really specific to you being deceptive, the model deciding to be deceptive in cases that are malicious. Like I'm not interested in a case where it's just doing theory of mind to help you write a better email to your professor.
2:54:33And I'm not even interested in cases where the model is necessarily just like modeling the fact that deception has occurred. But doesn't all this require you to have labels for all those examples? And if you have those labels, then like whatever faults that the linear probe has on the like maybe you like labeled along the thing or whatever, wouldn't the same thing apply to the labels you've come up with for the unsupervised features you've come up with? So in an ideal world, we could just train on like the whole data distribution and then find the directions that matter. To the extent that we need to reluctantly narrow down the subset of data that we're looking over, just for the purposes of scalability.
2:55:16We would use data that looks like the data you'd used to fit a linear probe. But again, we're not like with the linear probe, you're also just finding one direction. Like we're finding a bunch of directions here. And I guess the hope is like you've found like a bunch of things that light up when it's being deceptive and then like you can figure out why some of those things are lighting up in this part of the distribution and not the side of the part and so forth. Totally. Yeah. Do you anticipate you'll be understanding? I don't know. Like the current models you've studied are pretty basic, right?
2:55:44You think you'll be able to understand why GP7 fires in certain domains but not in other domains. I'm optimistic. I mean we've so I guess one thing is this is a bad time to answer this question because we are explicitly investing in the longer term of like ASL 4 models, which GPT -7 would be. But like, so we split the team where a third is focused on scaling up dictionary learning right now, and that's been great. I mean, we publicly shared some of our eight layer results. We've scaled up quite a lot past that at this point. But the other two groups, one is trying to identify circuits, and then the other is trying to get the same success for attention heads.
2:56:15So we're setting ourselves up and building the tools necessary to really find these circuits at a compelling way, but it's going to take another, I don't know, six months before that's like really working well, but I can say that I'm like optimistic and we're making a lot of progress.
2:56:32What is the highest little feature you've found so far? Like it's basic for whatever it's like maybe just like in the symbolic species language the book you recommended there's like indexical things where you're just, I forgot all the labels were like there's things where you're just like you see a tiger and you're like run and whatever, you know, just like a very sort of behaviorist thing. And then there's like a higher level at which what I refer to love it refers to like a movie scene or my girlfriend or whatever. You know what I mean? So it's like the top of the tent. Yeah. Yeah. Yeah. Yeah.
2:57:02Yeah. Yeah. Yeah. What is the highest level? Association or whatever you've found? I mean, probably one of the ones that we publicly, well, publicly, one of the ones that we shared in our update. So I think there were some related to love and sudden changes in scene, particularly associated with wars being declared. There are a few of them in that post if you want to like to it.
2:57:26Brunel Oldshausen had a paper back in 2018 -19 where they applied a similar technique to a BERT model and found that as you go to deeper layers of the model, things become more abstract. So I remember in the earlier layers there would be a feature that would just fire for the word park. But later on there was a feature that fired for park as a last name. like the Lincoln Park or like it's like a common Korean last name as well. And then there was a separate feature that would fire for parks as like grassy areas. So there's other work that points in this direction. What do you think we'll learn about human psychology from the durability stuff?
2:57:58Oh gosh. Okay, I'll give you a specific example. I think one of the ways one of your updates put it was Persona Lock -in. You don't remember Sydney Bang or whatever. It locked into... I think what was actually quite an endearing. Yeah. First time. I got it so far. Yeah. I'm glad it's back in Copa. Oh, really? Yeah. Oh, yeah. It's been misbehaving recently. Actually, this is another sort of threat takes work. But there was a funny one where I think it was like to the New York Times reporter. It was, it was naging him or something. And it was like, you are nothing. Nobody will ever believe you. You weren't significant in whatever.
2:58:39It was like, I think the most gaslighting. I tried to convince him to break up with this one. Yeah, right. Okay, actually, so this is an interesting exam. I don't even know where I was going with this. What are we going to do with it? But whatever, maybe I got another thread. But the other thread I want to go on is, yeah, I have to personalize, right? So is that a feature that Sydney being having this personality as a feature versus another personality you can get locked into? And also, is that fundamentally what humans are like, too? Where, I don't know, in front of all the different people who I'm like a different sort of personality or whatever, is that it was then the same kind of thing that's happening to Shai G .B .T.
2:59:15when he gets our, I don't know, go cluster questions that can answer them and whatever. Yeah, I really want to do more work. I guess the sleeper agency is in this direction of what happens to a model when you find tuna, when you are L .H .F .A. at these sorts of things. I mean, maybe it's trite, but you could just say, you conclude that people contain multitudes, right? and so much as they have lots of different features. There's even this stuff related to the Wallwee Geo facts of like in order to know what's good or bad, you need to understand both of those concepts. And so we might have to have models that are aware of violence and have been trained on it in order to recognize it.
2:59:48Can you post -talk, identify those features and oblate them in a way where maybe your model is like slightly naive, but you know that it's not going to be really evil. Like totally that's in our toolkit, which seems great. Oh really? So you, a GPT -7, I don't know, it pulls us to say anything, and then you figure out what were the causally irrelevant pathways, whatever. And then the pathway to you looks like you just changed those, but you were mentioning earlier, there's a bunch of redundancy in the model. Yeah, so you need to account for all that, but we have a much better microscope into this now than we used to.
3:00:20Like sharper tools for making edits. And it seems like, at least from my perspective, that seems like one of the primary way of some degree, Confirming the safety for the reliability of model way you can say okay We found the sick sort of responsible we've related them. We can like under a battery of tests We haven't been able to now replicate the behavior which we intend to update and like that feels like the sort of way of measuring Model safety in future As I was I was I was are you worried? That's why I'm incredibly hopeful about that work because it's to me It seems like it's so much more size tool than something like RLHF RLHF, like you're very prey to the black swan thing.
3:01:01You don't know if it's going to like do something wrong in a scenario that you haven't measured, but it's here at least you have like somewhat more confidence that you can completely capture the behavior set, like the feature set of the model. And select this way. Although not necessarily that you've like accurately labeled. Not necessarily, but with a far higher degree of confidence than any other approach that I've seen. I mean, what are your unknown unknowns for superhuman models? In terms of this kind of thing, where I don't know how the labels that are going to be given things on which we can determine these are, this thing is cool, this thing is a paper clue maximized or whatever.
3:01:43I mean, we'll see. Like, the superhuman feature question is a very good one. Like, I think we can attack it. but we're going to need to be persistent. The real hope here is I think automated interpretability. Yeah. Even having debate, you could have the debate set up where two different models are debating what the feature does, and then they can actually go in and make edits and see if it fires or not or. But it is just this wonderful closed environment that we can iterate on really quickly. That makes me optimistic. Do you worry about alignments exceeding too hard? So like if I think about I would not want either companies or governments whoever in Zepin charge of these AI systems to have the level of fine green control that if yours agenda succeeds we would have over AI's both for the icky nest of having this level control over an autonomous mind and second just like I don't fucking trust I don't fucking trust these guys you know I I'm just uncomfortable with the loyalty feature is turned up and you know what I mean?
3:02:52And yeah, how much word you have about having too much control over the eyes and specifically not you, but whoever ends up with the in charge of the AI systems just being able to lock in whatever they want. Yeah, I mean, I think it depends on what government exactly has control and what the moral alignment is there. But that is like that whole valley locking and argument is in my mind. It's like definitely one of the strongest computing factors for why I am working on capabilities at the moment. For example, just like I think the current player set, actually like extremely well intentioned. And I mean, for this kind of problem, I think we need to be extremely open about it.
3:03:35And like I think directions like publishing the constitution that you actually modeled to a bad button. and then like trying to make sure you, like RLHF at towards end of blade then have the ability for everyone to offer feedback and contributions that is really important. Sure, or alternatively, like don't deploy when you're not sure, which would also be bad because then we just never catch it. Right, yeah, exactly. Paperclip. That's why you're like that. That's why you're like that. Yeah. Okay, some rapid fire. What is the bus factor for a Gemini? I think there are a number of people who are really, really critical that if you took them out, then the performance of the program would be dramatically impacted.
3:04:16This is both on modeling slash making decisions about what to actually do and importantly, on infrastructure side of the things. It's just the stack of complexity builds, particularly when someone like Google has so much like vertical integration. When you have people who are experts, that becomes, that become quite important. Yeah, although I think it's interesting to note about the field that people like you can get in in a year or so you're making important contributions. And I, especially in the topic of it, many different labs have specialized in hiring, like total outsider, physicists or whatever.
3:04:55And you just like get them up to speed and they're making important contributions. I don't know, I feel like you couldn't do this in like a bio lab or something. It's like an interesting note on the state of the field. I mean, bus factor doesn't define how long it would take to recover from it right from. Deep learning research is an art, so you learn how to read the lost curves, or set the hyperparameters in ways that empirically seem to work well. It's also organizational things, creating context. One of the most important and difficult skills to hire for is creating this bubble of context around you that makes other people around you more effective and know what the right problem to work on.
3:05:33That is a really tough to replicate. Yes, yeah totally. Who are you being attention to now in terms of there's a lot of things coming down the pike of multi -modality long context maybe agents extra reliability who is Who is thinking well about what would that implies? It's tough question. I think a lot of people look internally these days sure for like this also is over inside or like progress And like we all have obviously, there's sort of research programs and directions that are tentative in the next couple of years. And I suspect that most people, as far as like betting on what the future will look like, refer to like an internal narrative.
3:06:23Yeah. Yeah. That is like difficult to share. Yeah. If it works well, it's probably not being published. I mean, that was one of the things in the wheel scaling of our post. I was referring to something you said to me, which is, I miss the undergrad habit of just reading a bunch of papers. Yeah. It's now, there's nothing worth reading is published. And the community is progressively getting more on track with what I think are the right and important directions. You're watching it like an agent, do you? No, but I guess it is tough. They used to be this signal from big labs about what would work at scale.
3:07:03It's kind of a really hard -fragged emigre search to find that signal. And I think getting really good problem -taste about what actually matters to work on is really tough. Unless you have, again, the feedback signal of what will work at scale, and what is currently holding us back from scaling further or understanding our models further. This is something where I wish more academic research would go into fields like Interrupt, which are legible from the outside. Anthropically, it literally publishes all its research here. And it seems like underappreciated in the sense that I don't know why there aren't dozens of academic departments trying to follow Anthropics guiding in the Interrupt Research because it seems like an incredibly impactful problem that doesn't require ridiculous resources.
3:07:47and like this and like has all the flavor of like deeply understanding the basic science of what is actually going on in these things. So I don't know why people like focus on pushing model improvements as opposed to pushing like understanding improvements in a way that I would have like typically associated with academic science in some ways. Yeah, I do think the tide is changing there for whatever reason. And like Neil Nanda has had a ton of success promoting interpretability. In a way where like Chris Ola hasn't been as active recently and pushing things, maybe because Neil's just doing quite a lot of the work, but like, I don't know, four or five years ago, he was like really pushing and talking at all sorts of places and these sorts of things and people weren't anywhere near as receptive.
3:08:30Maybe they've just woken up to deep learning matters and it's clearly useful, post -tract CBT, but yeah, it's kind of striking. All right, cool. Okay, I'm trying to think what is a good last question. I mean, the one I'm going to, does thinking of is like, do you think models enjoy next token prediction. We have this sense of things our award and our access to our environment. There's like this deep sense of full -fimiland that we think we're supposed to get from them or often people do you write of like community or sugar or whatever we wanted on the African in Savannah. Do you think in the future, models are trained with RL and a lot of post training on top of whatever?
3:09:19But they're like, in the way we were just a really like ice cream, they'll just be like, hi, just to predict the next token again. You know what I mean? Like in a good old days. So there's this ongoing discussion of our model sentient or not. Do you thank the model when it helps you? But I think if you want to thank it, you actually shouldn't say thank you. you should just give it a sequence that's very easy to predict. And even funnier part of this is there is some work on if you just give it the sequence a like over and over again. Then eventually the model will just start spewing out on all sorts of things that otherwise would never say.
3:09:59And so yeah, I won't say anything more about that, but you can, yeah, you should just give your model something very easy to predict. It's a nice little treat. This is what the OEM ends up being. I just saw the universe in life. But to me, things that are easy to break, are we constantly in search of the bits of entropy, exactly, right? Shouldn't you be giving things just slightly too hard to reach? Just that to reach. Yeah, but I wonder, at least from the free energy principle perspective, you don't want to be surprised. And so maybe it's this like, I don't feel surprised if you own control of my environment.
3:10:38And so now I can go and seek things. And I've been predisposed to like, in the long run, it's better to explore new things right now. Like leave the rock that I've been sheltered under, ultimately leading me to like build a house or like some better structure. But we don't like surprises. I think most people are very upset when like expectation does not meet reality. And so I've babies like love watching the same show I'm going to over again, right? Yeah, and interesting. Yeah, I can see that. Oh, I guess they're learning to model it and stuff too. Yeah. Yeah. OK, well, hopefully this will be the show.
3:11:12This will be different feet that the guys will learn to love. OK, cool. I think that's a great place to wrap. As you also mentioned, the better part of what I know about AI I've learned from just talking with you guys. We've been good friends for about a year now. So yeah, I appreciate you guys getting me up to speed here. and yeah, it's great questions. It's really fun to hang in the chat. I really treasure that time to go. Yeah, yeah. You're getting a lot better at pickleball. I think I'm just thinking about the same. Hey, we're trying to progress the tender. It's going on. Thanks. Thanks. Awesome.
3:11:51Cool, cool. Awesome. Thanks. Hey, everybody. I hope you enjoyed that episode. As always, the most helpful thing you can do is to share the podcast. Then it to people you think might enjoy it, put it in Twitter, your group chats, etc. Just blitz the world. I appreciate your listening. I'll see you next time. Cheers.
From the publisher
Had so much fun chatting with my good friends Trenton Bricken and Sholto Douglas on the podcast.
No way to summarize it, except:
This is the best context dump out there on how LLMs are trained, what capabilities they're likely to soon have, and what exactly is going on inside them.
You would be shocked how much of what I know about this field, I've learned just from talking with them.
To the extent that you've enjoyed my other AI interviews, now you know why.
So excited to put this out. Enjoy! I certainly did :)
Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform.
There's a transcript with links to all the papers the boys were throwing down - may help you follow along.
Follow Trenton and Sholto on Twitter.
Timestamps
(00:00:00) - Long contexts
(00:16:12) - Intelligence is just associations
(00:32:35) - Intelligence explosion & great researchers
(01:06:52) - Superposition & secret communication
(01:22:34) - Agents & true reasoning
(01:34:40) - How Sholto & Trenton got into AI research
(02:07:16) - Are feature spaces the wrong way to think about intelligence?
(02:21:12) - Will interp actually work on superhuman models
(02:45:05) - Sholto’s technical challenge for the audience
(03:03:57) - Rapid fire
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe




