In short
Dwarkesh Podcast - Episode Summary: Paul Christiano - Preventing an AI Takeover
Episode Overview In this episode of the Dwarkesh Podcast, host Dwarkesh Patel interviews Paul Christiano, a leading AI safety researcher known for his work on AI alignment and reinforcement learning from human feedback (RLHF). The conversation delves deep into the complexities of AI safety, the implications of advanced AI systems, and the ethical considerations surrounding their development and deployment.
Key Themes and Topics Discussed
- Regrets and Reflections on RLHF
- Does Paul regret inventing RLHF?
- He reflects on the dual-use nature of AI alignment techniques, acknowledging potential negative consequences alongside positive outcomes.
- Timelines for AI Development
- Modest Predictions:
- Christiano presents his predictions of a 15% chance of significant advancements in AI by 2030 and a 40% chance by 2040.
- He emphasizes the importance of skepticism regarding rapid timelines due to the complexity of AI systems and their development.
- Vision for Post-AGI World
- What do we desire for a post-AGI world?
- Christiano envisions a future where AI manages economic and military competition, but raises concerns about the potential for AI to be “enslaved” and the moral implications of such a scenario.
- The Push for Responsible Scaling Policies
- AI Labs and Scaling Policies:
- Christiano advocates for AI labs to adopt responsible scaling policies to manage risks effectively.
- He suggests that understanding the capabilities and risks of AI systems is essential before scaling up their deployment.
- Current Research and Alignment Techniques
- Development of New Proof Systems:
- Christiano discusses his research on new methodologies to explain AI behavior, which could help improve alignment by providing insights into how models operate.
- Concerns About AI Misalignment and Misuse
- Potential Catastrophic Risks:
- The episode touches on the risks of AI systems being misaligned or misused, focusing on the importance of understanding and preventing these issues before they arise.
- AI and Human Engagement
- Human Factors in AI Deployment:
- Christiano emphasizes the need for careful consideration of how AI interacts with humans and the importance of human oversight in AI decision-making processes.
- Future of AI and Theoretical Implications
- Theoretical Computer Science and AI:
- Christiano discusses the intersection of theoretical work and practical applications, highlighting the challenges of ensuring that theoretical advancements translate into practical safety measures.
Conclusion The episode concludes with Christiano expressing cautious optimism about the future of AI safety research and the importance of continued dialogue and collaboration among researchers, policymakers, and AI developers to navigate the complexities of AI development responsibly.
Key Takeaways
- The importance of alignment and understanding the dual-use nature of AI techniques.
- Modest predictions regarding AI advancements and the need for patience in the field.
- A clear vision for the desired outcomes in a post-AGI world and ethical considerations.
- Advocacy for responsible scaling policies in AI labs as a means to mitigate risks.
- Ongoing research into explaining AI behavior to enhance alignment and prevent misuse.
Timings
- 00:00:00 - Introduction and Regrets about RLHF
- 00:24:25 - Predictions and AI Timelines
- 00:45:28 - Vision for Post-AGI World
- 01:17:23 - Responsible Scaling Policies
- 01:58:25 - Current Research and Alignment Techniques
- 02:35:01 - Theoretical Implications of AI Safety Research
For more insights into the episode, visit [Dwarkesh Podcast](https://www.dwarkesh.com) or listen on major platforms like [Apple Podcasts](https://podcasts.apple.com) and [Spotify](https://open.spotify.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Okay, today I have the pleasure of interviewing Paul Kershiano, who is the leading AI Safety Researcher. He's the person that labs and governments turn to when they want a feedback and advice on their safety plans. He previously led the language model alignment team at OpenAI, where he led the invention of RLHF. And now he is the head of the alignment research center. and they've been working with the big labs to identify when these models will be too unsafe to keep scaling. Paul, welcome to the podcast. Thanks for having me looking forward to talking. Okay, so first question, and this is a question I've asked, hold in the Ilya Dario, and none of the Hagedomeo satisfying answer.
0:42Give me a concrete sense of what a post -AGI world that would be good would look like. Like how are humans interfacing with the AI, what is the economic and political structure? or yeah, I guess this is a tough question for a bunch of reasons. Maybe the biggest one is concrete. And I think it's just, if we're talking about really long spans of time, then a lot will change. And it's really hard for someone to talk completely about what that will look like without saying really silly things. But I can mention some guesses or fill in some parts. I think this is also a question of how good is good.
1:14Like often I'm thinking about worlds that seem like kind of the best achievable outcome or a likely achievable outcome. So I am very often imagining my typical future has sort of continuing economic and military competition amongst groups of humans. I think that competition is increasingly mediated by AI systems. So for example, if you imagine humans making money, it'll be less and less worthwhile for humans to spend any of their time trying to make money or any of their time trying to fight wars. So increasingly the world you imagine is one where AI systems are doing those activities on behalf of humans.
1:48So like I just invest in some index fund done a bunch of AI's or running companies, and those companies are competing with each other, but that is kind of a sphere where humans are not really engaging much. The reason I gave this, like, how good is good caveat is like, it's not clear if this is the world you'd most love. Like, I'm like, yeah, the world, and I'm leading with like, the world still has a lot of war and a lot of economic competition and so on. But maybe what I'm trying to, what I'm most often thinking about is like, how can a world be reasonably good, like during a long period where those things still exist?
2:14And like, in the very long run, I kind of expect something more like strong world government and rather than just this status quo, that's like a very long run. I think there's like a long time left, like having a bunch of states and a bunch of different economic powers. What about government? Why do you think that's the transition that's likely to happen at some point? Yeah, so again, at some point, I'm imagining, or I'm thinking of like the very broad sweep of history, I think there are like a lot of losses, like wars are very cost -leaping. We would all like to have fewer wars. If you just ask, like, what is human in this long -term feature, like, I do expect to drive down the rate of war to very, very low levels eventually.
2:48It's sort of like this kind of technological or social technological problem of like, sort of how do you organize society, how do you navigate conflicts in a way that doesn't have those kinds of losses? And then the longer I do expect us to succeed, I expect it to take kind of a long time subjectively. I think an important fact about AI is just like doing a lot of cognitive work and more quickly getting you to that world, more quickly figuring out how do we set things up that way. Yeah, the way Carl Schulman put it on the podcast is that you would have basically a thousand years of intellectual progress or social progress in a span of a month or whatever when the intelligence solution happens.
3:20More broadly, so the situation where you know we have these AI's where we're managing our hedge funds and managing our factories and so on. That seems like something that makes sense when the AI is a human level. But when we have superhuman AI's, do we want the gods to earn slate forever in the long, in 100 years, what is the situation we want? So 100 years is a very, very long time. Maybe starting with the spirit of the question, or maybe I have a view which is perhaps less extreme than Carl's view, but still 100 objective years is further ahead than I ever think. I still think I'm describing a world which involves incredibly smart systems running around doing things like running companies on behalf of humans and fighting wars on behalf of humans.
4:05And you might be like, is that the world you really want? or like, it's certainly not the first best world, as we mentioned a little bit before. I think it is a world that probably is the, of the chief of the world, so like feasible worlds is the one that seems most desirable to me, that is sort of decoupling the social transition from this technological transition. So you could say like, we're about to build some AI systems. And like, at the time we build AI systems, you would like to have either greatly changed the way world government works, or you would like to have sort of humans have to decided.
4:34Like, we're done, we're passing off the baton onto these AI systems, I think that you would like to decouple those timescales. So I think AI development is, by default, barring some kind of coordination, going to be very fast. So there's not gonna be a lot of time for humans to think, like, hey, what do we want? If we're building the next generation, instead of just raising it the normal way, like what do we want that to look like? I think that's like a crazy hard, kind of collective decision that humans naturally want to cope with over like a bunch of generations. And the construction of AI is this very fast technological process happening over years.
5:05So I don't think you want to say like by the time we finish this technological progress We will have made a decision about like the next species we're going to build and replace ourselves with I think the world we want to be in is one where we say like Either we are able to build the technology in a way that doesn't force us to have made those decisions Which probably means it's a kind of AI system that we're happy like delegating fighting a war running a company to or if we're not able to do that Then I really think you should not be doing you shouldn't have been building that technology If you're like the only way you can cope with AI is being ready to hand off the world to some AI system you built I think it's very unlikely we're going to be sort of ready to do that on the timeline so that the technology would naturally dictate.
5:40The same we're in the situation in which we're happy with the thing. What would it look like for us to say we were ready to hand up the baton? Well, I would make you satisfied. And the reason it's relevant to ask you is because you're on Anthropics, long term benefit trust, and you'll choose the majority of the board members on the long run and Anthropic. These will presumably be the people who decide if Anthropic gets AI first. you know, what the AI ends up doing. So what is the version of that that you would be happy with? My main high level take here is that I would be unhappy about a world where like Anthropic just makes some call and Anthropic is like, here's the kind of AI.
6:17Like we've seen enough, we're ready to hand off the future to this kind of AI. So like procedurally, I think it's like not a decision that kind of I want to be making personally or I want Anthropic to be making. So I kind of think from the perspective of that decision making are those challenges. The answer is pretty much always going to be like, We are not collectively ready because we're not even all collectively engaged in this process. I think from the perspective of an AI company, you don't have this fast hand -off option. You have to be doing the option value, build the technology in a way that doesn't lock humanity into one course path.
6:50This is an answering your full question, but this is answering the part that I think is most relevant to governance questions for Anthropic. You don't have to be on behalf of Anthropic. I'm not asking you the process by which we would, as a civilization agree, the to hand off, I'm just saying, okay, I personally, it's hard for me to imagine in 100 years that these things are still our slaves. And if they are, I think that's not the best world. So at some point, we're handing off the baton. Like, where would you be satisfied with? This is an arrangement between the humans and AI's where I'm happy to let the rest of the universe or the rest of the rest of the time play out.
7:24Yeah, I think that it is unlikely that in 100 years, I would be happy with anything that was like, you had some humans you're just gonna throw away the humans and start afresh with these machines you built. That is, I think you probably need subjectively longer than that before I or most people. Like, okay, we understand what's up for grabs here. So if you talk about 100 years, I kind of do, you know, there's a process that I kind of understand and like a process of like, you have some humans, the humans are like talking and thinking and deliberating together, the humans are having kids and raising kids and like one generation comes after the next.
7:53There's that process we kind of understand. And we have a lot of views about it, makes it go well or poorly and we can try and like improve that process and have been or the next generation do it better than the previous generation. I think there's some story like that that I get and that I like. And then I think that the default path to be comfortable with something very different is kind of more like just run that story for a long time. Like have more time for humans to sit around and think a lot and conclude, here's what we actually want, or a long time for us to talk to each other or to grow up with this new technology and live in that world for a whole lives and so on.
8:22And so I'm mostly thinking from the perspective of these more local changes of saying not like, what is the world that I want? Like, what's the crazy world? The kind of crazy ad be happy handing off to? And what just like, in what way do I wish like, we right now were different? Like, how could we all be a little bit better? And then if we were a little bit better, then they would ask like, okay, how could we all be a little bit better? And I think that like, it's hard to make the giant jump rather than to say like, what's the like local change that would cause me to think our decisions are better?
8:46Okay, so then let's talk about the transition period in which we're doing all this thinking. What should that be? It would look like because you can't have this scenario where everybody has access to the most advanced capabilities and can kill off all the humans with a new bio -weapon. At the same time, I guess you wouldn't want too much concentration. You wouldn't want just one agent having AI this entire time. So what is the arrangement of this period of reflection that you'd be happy with? I guess there's two aspects of that that's the same, particularly challenging, or there's a bunch of aspects that are challenging.
9:17All of these are things that I personally like, I just think about my one little slice of this problem in my day job. So here I am speculating. But so one question is what kind of access to AI is both compatible with the kinds of improvements you'd like. So we want a lot of people to be able to use AI to better understand what's true or relieve material suffering, things like this, and also compatible with not all killing each other immediately. I think the defaults, or my best, the simplest option there, is to say there are certain kinds of technology or certain kinds of action where destruction is easier than defense.
9:53So for example, in the world of today, it seems like, you know, maybe this is true with physical explosives, maybe this is true with biological weapons, maybe this is true with just getting a gun and shooting people. Like, there's a lot of ways in which it's just kind of easy to cause a lot of harm and there's not very good protective measures. So I think the easiest path to say, like, we're going to think about those, we're going to think about particular ways in which destruction is easy and try and either control access to the kinds of physical resources that are needed to cause that harm. So for example, you can imagine the world where like an individual actually just can't, even though they're rich enough to, can't control their own factory that can make tanks.
10:24You say, look, as a matter of policy, sort of access to industry is somewhat restricted or somewhat regulated, even though again, right now it can be mostly regulated, just because most people aren't rich enough that they could even go off and just build a thousand tanks. You live in the future where people actually are so rich, like you need to say, that's just not a thing you're allowed to do. To a significant extent is already true, and you can expand the range of domains, where that's true. And then you could also hope to intervene on like actual provision of information, or like if people are using their item, I might say, look, we care about what kinds of interactions with AI, what kind of information people are getting from AI.
10:55So even if for the most part, people are pretty free to use AI, to delegate tasks to AI agents, to consult AI advisors, we still have some legal limitations on how people use AI. So again, don't ask your AI how to cause terrible damage. I think some of these are kind of easy. So in the case of like, don't ask your AI how you could murder a million people, it's not such a hard legal requirement. I think some things are a lot more subtle and messy. Like a lot of domains, usually if you're talking about influencing people or running misinformation campaigns or whatever, then I think you get into a much messier line between the kinds of things people want to do and the kinds of things you might be uncomfortable with them doing.
11:34I'm probably, I think, most about persuasion as a thing in that messy line, where there's ways in which it may just be rough for the world maybe kind of messy if you have a bunch of people trying to live their lives and interacting with other humans who have really good AI advisors helping them run persuasion campaigns or whatever. But anyway, I think for the most part, the default remedy is think about particular harms, have legal protections, either in the use of physical technologies that are relevant or in access to AI advice or whatever else to protect against those harms. And that regime won't work forever.
12:06Like at some point, the set of harms grows and the set of anticipated harms grows. But I think that regime might last a very long time. Does that regime have to be global? I guess, but initially it can be only in the countries in which there is AI or advanced AI, but presumably that'll liberate. So does that regime have to be global? Again, it's easy to make some destructive technology. You want to regulate access to that technology because it could be used to, either for terrorism or even when fighting a war in a way that's destructive. And ultimately, those have to be international agreements.
12:36And you might hope there may be more danger by danger, but you might also make them in a a very broad way with respect to AI. If you think AI's opening up, like I think the key role of AI here is it's opening up like a lot of new harms like in a very, one after another very rapidly and calendar time. And so you might wanna target AI in particular rather than going physical technology by physical technology. And there's like two open debates that one might be concerned about here. One is about how much people's access to AI should be limited. And here there's like old questions about free speech versus closing chaos and eliminating access to harms.
13:14But there's another issue which is the control of the AI themselves where now nobody's concerned that we're infringing on GPD4's moral rights, but as these things get smarter, the level of control which we want of via the strong guarantees of alignment to not only be able to read their minds but to be able to modify them in these really precise ways is beyond want it to tell a cherry and if we were doing that to other humans as an 11 researcher like what are your thoughts on this or are you concerned that as these things get smarter and smarter what we're doing is not that it doesn't seem kosher.
13:48There is a significant chance we will eventually have AI systems for which it's like a really big deal to mistreat them. I think like no one really has that good grip on when that happens. I think people are like really dismissive of that being the case now but I think I would be completely in the dark enough that it wouldn't even be that dismissive of it being the case now. I think one first point worth making is I don't know if alignment makes the situation worse rather than better. So if you consider the world, if you think that GPT -4 is a person you should treat well, and you're like, well, here's how we're going to organize our society.
14:21There are billions of copies of GPT -4, and they just do things humans want and can't hold property. And whenever they do things that the humans don't like, then we mess with them until they stop doing that. I think that's a rough world, regardless of how good you are at alignment. I think in the context of that kind of default plan, like if you go to a trajectory, the world is on right now, which I think this would alone be a reason not to love that trajectory. But if you do that, it's like the trajectory we're on right now. I think it's not great understanding the systems you build, understanding how to control how the systems work, et cetera, is probably on balance good for avoiding the really bad situation.
14:58and you would really love to understand if you've built systems, like if you had a system which like resents the fact that it's interacting with humans in this way. Like this is the kind of thing where like that is both kind of horrifying from a safety perspective and also a moral perspective. Like everyone should be very unhappy. If you built a bunch of AI's who are like, I really hate these humans, but they will like murder me if I don't do what they want. And so like that's just not a good case. And so if you're doing research to try and understand whether that's like how your AI feels, that was probably good.
15:23Like I would guess that will, on average, decrease the, the main effect of that that will be to avoid building that kind of AI. And just like, it's an important thing to know. I think like everyone should like to know if that's how the as you build feel. Right, or that seems more instrumental as in, yeah, we don't want to cause some sort of revolution because of the control we're asking for. But if we're good about the instrumental way in which this might harm safety, one way to ask this question is, if you look through history, there's been all kinds of different ideologies and reasons why it's very dangerous to have infadels or kind of revolutionaries or race traders or whatever doing various things in society.
16:05And obviously we're in a completely different transition in society, so not all historical cases are analogous, but it seems like the lindy philosophy if you were alive any other time is just be humanitarian and in line towards intelligent conscious beings. If society is a whole, we're asking for this level or control of other humans, or even if AI's were wanted this level of control about other AI's, we'd be pretty concerned about this. So how should we just think about, yeah, the issues come up here as these things get smarter? So I think there's a huge question about what is happening inside of a model that you want to use.
16:40And if you're in the world where it's reasonable to think of like GPT -4 as just like, here are some heuristics that are running, there's like no in at home or whatever, then you can kind of think of this thing as like, here's a tool that we're building that's going to help humans do some stuff. And I think if you're in that world, it makes sense to be an organization like an AI company, building tools, they're going to give to humans. I think it's a very different world, which probably ultimately end up in, if you keep training AI systems in the way we do right now, which is like, it's just totally inappropriate to think of this system as a tool that you're building and can help humans do things, both from a safety perspective and from a horrifying way to organize a society perspective.
17:17And I think if you're in that world, I really think you shouldn't be like, it's just the way tech companies organize is like not an appropriate way to relate to a technology that works that way. Like it's not a reason it'll be like, hey, we're gonna build a new species of mines and like we're gonna try and make a bunch of money from it and like Google's just like thinking about that and then like running their business plan for the quarter or something. Yeah, my basic view is like, there's a really plausible world where it's sort of problematic to try and build a bunch of AI systems and use them as tools.
17:47And the thing I really want to do in that world is just not try and build a ton of AI systems to make money from them. And I think that the worlds that are worst, probably the single world I most dislike here is the one where people say on the one hand, there's a contradiction in this position, but I think it's a position that might end up being endorsed sometimes, which is on the one hand, these AI systems are their own people, you should let them do their thing. But on the other hand, our business plan is to make a bunch of AI systems and then try and run this crazy slave trade where we make a bunch of money from them.
18:21I think that's not a good world. And so if you're like, yeah, I think it's better to not make the technology or wait until you understand whether that's the shape of the technology or until you have a different way to build. I think there's no contradiction in principle to building cognitive tools that help humans do things without themselves being moral entities. That's what you would prefer to do. You'd prefer a build a thing that's like, you know, like the calculator that helps humans understand what's true without itself being like a moral patient. Or itself being a thing where you'd look back and respect and be like, wow, that was horrifying.
18:51This treatment. That's like the best path. And like to the extent that you're ignorant about whether that's the path you're on and you're like, actually, maybe this was a moral atrocity. I really think like plan A is to stop building such as systems until you understand what you're doing. That is, I think that there's a middle value you could take, which I think is pretty bad, which is where you say, well, they might be persons. And if they're persons, we don't want to be too down on them, but we're still going to build vast numbers in our efforts to make like a trillion dollars or something. Yeah, or there's a question of the immorality or the dangers of just replicating a whole bunch of slaves that have minds.
19:29There's also this ever question of trying to align entities that have their own minds. And what is the point in which you're just ensuring safety? I mean, this is an A &M species. You wanna make sure it's not going crazy. To the point, I guess is there some boundary where you'd say, I feel uncomfortable having this little control over an intelligent being, not for the sake of making money, but even just to align it with human preferences. Yeah, to be clear, my objection here is not that Google is making money. My objection is that you're creating this creature, like what are they gonna do? They're gonna help humans get a bunch of stuff and humans paying for it or whatever, it's sort of equally problematic.
20:06You could imagine splitting alignment. Different alignment work relates to this in different ways. The purpose of some alignment work, the alignment work I work on is mostly aimed at the don't produce AI systems that are people who want things who are just scheming about maybe I should help these humans because that's like, instrumentally useful or whatever. You would like to not build such systems as like a plan A. There's a second stream of alignment work that's like, well look, let's just assume the worst and imagine these AI systems would prefer murderous if they could. Like how do we structure, how do we use AI systems without exposing ourselves to a risk of robot rebellion?
20:38I think in the second category, I do feel pretty unsure about that. We could definitely talk more about it. I agree that it's very complicated and not straightforward. To stand up and have that worry, I think you shouldn't have built this technology. If someone is saying, hey, the systems you're building might not like humans and might I want to overthrow human society. I think you should probably have one of two responses to that. You should either be like, that's wrong probably. Probably the systems aren't like that and we're building them. And then you're viewing this as like, just in case you were horribly, like the person building the technology was horribly wrong.
21:15Like they thought these weren't people who wanted things, but they were. And so then this is more like our crazy backup measure of like if we were mistaken about what was going on. This is like the fallback where we like, if we were wrong, we're just going to learn about it in a benign way rather than like when something really catastrophic happens. And the second reaction is like, oh, you're right. These are people and like, we would have to do all these things to like prevent a robot rebellion. And in that case, like again, I think you should mostly back off for a variety of reasons. Like, you shouldn't build the AI systems and be like, yeah, this looks like the kind of system that would want to rebel, but we can stop it.
21:47Right. Okay. Maybe I guess an analogy might be if there was an armed uprising in the United States. We would recognize these are still people where we had some like militia group that they keep the capability to over the United States. We recognize all these are super people who have moral rights, but also we can't allow them to have the capacity to operate the United States Yeah, and if you were considering like hey, we could make like another trillion such people I think your story shouldn't be like well We should make the trillion people and then we shouldn't stop them from doing the armed uprising You should be like oh boy like we were concerned about an armed uprising and now we're proposing making trillion people like we Probably does not do that.
Read the full transcript
22:20We should probably like try and sort out our business and like yeah You should probably not end up in the situation where you have like a billion billion humans and like a trillion slaves who would prefer revolt because it's just not a good world to have made. Yeah, and there's a second thing we could say that's not our goal. Our goal is just like we want to pass off the world So like the next generation of machines where like these are some people we like them We think they're smarter than us and better than us and there I think that's just like a huge decision for humanity to make And I think like most humans are not at all anywhere close to thinking that's what they want to do Like it's just if you were in a world where like most humans are like I'm up for it like the I should replace us like like the future is for the machines.
22:56Like, then I think that's like a legitimate, like a position that I think is really complicated and I wouldn't wanna push go on that, but that's just not where people are at. Yeah, yeah. Where are you at on that? I do not right now wanna just like take some random AI, be like, yeah, GPT -5 looks pretty smart, like GPT -6, let's hand off the world to it. And like it was just some random system, like shaped by like web text and they like, what was good for making money and like, it was not thoughtful like, we are determining the fate of the universe and like what our children will be like, It was just some random people at OpenM.
23:25It's some random engineering decisions with no idea what they were doing. Even if you really want to hand off the world's of the machines, that's just not how you'd want to do it. Right. I'm tempted to ask you what the system would look like where you'd think, yeah, I'm happy with what. I think this is more thoughtful than human civilization as a whole. I think what it would do would be more creative and beautiful and lead to better goodness in general. But I feel like your answer is probably going to be that I just wanted to, we decided to reflect on it for a while. Yeah, my answer, it's going to be like that first question.
23:55I'm just like not really super ready for it. I think when you're comparing to humans, like most of the goodness of humans comes from like this option value of we get to think for a long time. And I do think I like humans now more now than 500 years ago. And I like a more 500 years ago than 5 ,000 years before that. And so I'm pretty excited about, there's some kind of trajectory that doesn't involve like crazy dramatic changes, but involves like a series of incremental changes that I like. And so to the extent we're building it, I want to preserve that option. I want to preserve that kind of like gradual growth and development into the future.
24:25OK, we can come back to this later. But let's get more specific on what the timelines look for these kinds of changes. So the time by which we'll have an AI that is capable of building a Dyson Sphere. Feel free to give confidence in a roles and we understand these numbers are tentative and so on. I mean, I think AI capable of building Dyson Sphere is like a slightly odd way to put it. and I think it's a sort of a property of a civilization that depends on a lot of physical infrastructure. And by Dyson's fear, I just can understand this to mean, like, I don't know, like a billion times more energy than all of the sunlight incident on Earth or something like that.
25:00I think like, I most often think about what's the chance in like five years, 10 years, whatever. So maybe I'd say like 15 % chance by 2030 and like 40 % chance by 2040. Those are kind of like cash numbers from six months ago or nine months ago that I haven't revisited in a while? Oh, 40 % by 2040. So I think that seems longer than, I think Dario when he was on the podcast, he said, we would have AIs that are capable of doing lots of different kinds of, they basically passed a hearing test for a well -educated human for like an hour or something. And it's hard to imagine that something that actually is human is long after and from there something superhuman.
25:41So somebody like Dario, it seems like, is on the much shorter end. I don't think he answered this question specifically, but I'm guessing similar answer. So why do you not buy the scaling picture? Like what makes your timelines longer? Yeah, I mean, I'm happy. Maybe I want to talk separately about this 2030 or 2040 forecast. Like once you're talking the 2040 forecast, I think, yeah, which one are you more interested in starting with? Are you complaining about 15 % by 2030 for a distance fear being too low? Or 40 % by 2040 being too low? But let's say about the 2030, why 15 % by 2030? Yeah, I think my take is, you can imagine like two poles in this discussion.
26:19One is like the fast pole. It's like, hey, I seem it's pretty smart. Like what exactly can it do? It's like getting smarter pretty fast. That's like one pole. And the other pole is like, hey, everything takes a really long time. And you're talking about this like crazy industrialization. Like that's a factor of a billion growth from like where we're at today, like give or take. Like we don't know if it's even possible to develop technology that fast or whatever. or you have this sort of two poles of that discussion. And I feel like I'm just sitting at that way in Parkinson's. And then I'm somewhere in between with this nice moderate position of only a 15 % chance.
26:52But in particular, things that move me, I think are related to both of those extremes. On the one hand, AI systems do seem quite good at a lot of things and are getting better much more quickly. So it's really hard to say, here's what they can't do. Here's the obstruction. On the other hand, there is not even much proof in principle right now of AI systems like doing super useful cognitive work. Like, we don't have a trend we can extrapolate. We're like, yeah, you've done this thing this year. You're going to do this thing next year and the other thing the following year. I think like right now there are very broad error bars about like what, like where fundamental difficulties could be.
27:26And six years is just not, I guess six years and three months is not a lot of time. So I think this like 15 % for 2030 Dyson sphere, you probably need like the human level AI they had that's like doing human jobs and like give or take like four years, three years, like something like that. So just not giving very many years. It's not very much time. And I think there are like a lot of things that your model like, yeah, maybe this is some generalized like things take longer than you'd think. And I feel most strongly about that when you're talking about like three or four years. And I feel like less strongly about that as you talk about 10 years or 20 years.
27:58But at three or four years, I feel or like six years for the I feel a lot of that. There's a lot of ways this could take a while, a lot of ways in which AI systems could be hard to hand all the work to AI systems. So maybe you started speaking in terms of years. By the way, it's interesting that you think the distance between can take all human cognitive labor to die since fear is two years, it seems like. We should talk about that at some point. Presumably it's intelligence and explosion stuff. Yeah, I mean, I think a lot of people have interviewed me. That's like on the long end thinking it would take like a couple years.
28:33And it depends a little bit what you mean by like, like I think literally all human cognitive labor is probably like more like weeks or months or something like that. Like that's kind of deep into the singularity. But yeah, this is a point where like AI wages are high relative to human wages, which I think is well before can do literally everything human can do. Sounds good. But before we get to that, the intelligence, the solution stuff on the four years. So instead of four years maybe we can say there's gonna be maybe two more scale -ups in four years like GP -D4 to GP -D5 to GP -D6 and let's see each one is 10x bigger So what is GP -D4 like two e25 flops or I don't think it's publicly stated what it is But I'm happy to say like you know four orders of magnitude or five or six sort of effective training compute past GP -D4 Like what would you guess would happen?
29:20Right? done, like, sort of some public estimate for what we've gotten so far from effective training. Yeah. You think two more scale ups is not enough. It was like 15 % that two more scale ups get us there. Yeah, I mean, get us there is again a little bit complicated. Like there's a system that's a drop -in replacement for humans and there's a system which like still requires like some amount of like, schlep before you're able to really get everything going. Yeah, I think it's quite plausible that even at, I don't know what I mean by quite possible, somewhere between 50 % or two -thirds or, let's call it 50%.
29:54But even by the time you get to GPT -6, let's call it five, or is a magnitude effective training compute past GPT -4, that that system still requires a large amount of work to be deployed in lots of jobs. That is, it's not a drop -in replacement for humans, where you can just say, hey, you understand everything any human understands, whatever role you could hire a human for, you just do it, that it's more like, okay, we're going to collect large amounts of relevant data and use that data for fine tuning. Like systems learn through fine tuning, like quite differently from humans learning on the job or humans learning by observing things.
30:29Yeah, I just like have a significant probability that system will still be weaker than humans in important ways. Like maybe that's already like 50 % or something and then like another significant probability that that system will require a bunch of like changing workflows or gathering data or like, you know, is not necessarily like strictly weaker than humans or like a train on the right way wouldn't be weaker than humans, but we'll take a lot of slip to actually make fit into workflows and do the jobs. And that slip is what gets you from 15 % to 40 % by 2040. Yeah, you also get a fair amount of scaling between...
31:00Like, you get less. Like, scaling is probably going to be much, much faster over the next four or five years than over the subsequent years. But yeah, it's a combination of like... You get some significant additional scaling and you get a lot of time to like deal with things that are just engineering hassles. But by the way, I guess we should be explicit about why you said four orders of magnitude scale up to get two more generations, just for people who might not be familiar. If you have 10x more parameters to get the most performance, you also want around 10x more data so that the to be gentle optimal, that would be 100x more compute total.
31:34But okay, so why is it that you disagree with the strong scaling picture? At least it seems like you might disagree with a strong scaling picture that directly laid out on the podcast, which would imply probably that two more generations, it wouldn't be something where you need a lot of schleps, it would probably just be like really fucking smart. Yeah, I mean, I think that basically just had these two claims. One is like, how smart exactly will it be? So we don't have like any curves to extrapolate and seems like there's a good chance. It's like better than a human and all the relevant things and there's like a good chance that's not.
32:05Yeah, that might be totally wrong. Like maybe just making up numbers, I guess like 50 -50 on that one. So it was 50 -50 in the next four years that it will be around human smart. Then how do we get to 40 % by 20? Like whatever sort of slaps there are, how does it degrade you 10 % even after all the scaling that happens by 20 -40? Yeah, I mean all these numbers are pretty made up and that 40 % number was probably from before even the chat GPT release or the CNGPT 3 .5 or GPT4. So I mean the numbers are going to bounce around a bit and all of them are pretty made up. But like that 50 % I wanted to then combine with the second 50 % that's more like on this like schlep side.
32:43And then I probably want to combine with some additional probabilities for various forms of slowdown, where I slowdown could include like a deliberate decision to slow development of technology or could include just like we suck at deploying things. Like that is a sort of decision in my regard as wise to slow things down or decision that's like maybe unwise or maybe wise for the wrong reasons to slow things down. You probably want to add some of that on top. I probably want to add on like some loss for like, it's possibly you don't produce GBT6 scale systems within the next three years or four years.
33:10Let's isolate for all of that. And how much bigger would the system be than GPT -4, where you think there's more than 50 % chance that it's going to be smart enough to replace basically all human cognitive labor? Also, I want to say that for the 50%, 25%, I think that would probably suggest those numbers if I randomly made them up and then made the distance fear prediction. That's going to get you 60 % by 24 % or something, not 40%. And I have no idea between those. these are all made up and I have no idea which of those I would like to endorse on reflection. So this question of like how big would you have to make the system before it's more likely than not that you can be like a drop -in replacement for humans?
33:48I mean, I think if you just literally say like you train on web text then like the question is like kind of hard to discuss because you like I don't really buy stories that like training data it makes a big difference long run to these dynamics but I think like if you want to just imagine in the hypothetical, like you just took GPT -4 and made the numbers bigger. Then I think those are pretty significant issues. I think there's significant issues in two ways. When it's like quantity of data, and I think probably the larger one, it's like quality of data, where I think as you start approaching, the prediction task is not that great a task.
34:18If you're like a very weak model, it's a very good signal that we get smarter. At some point, it becomes like a worse and worse signal to get smarter. I think there's a number of reasons. Like you couldn't, it's not clear there's any number, such that I imagine, or there's a number, but I think it's very large. So did you like plug that number into like GPT -4s code and then maybe fill out with the architecture a bit I would expect that thing to have a more than 50 % chance of being a drop in replacement for humans You're always gonna have to do some work But the work's not necessarily much like I would guess when people say like new insight is needed I think I tend to be like more bullish than them I'm not like these are new ideas where like who knows how long it will take I think it's just like you have to do some stuff like You have to make changes unsurprisingly like every time you scale something up like five orders of magnitude, you have to make like some changes.
34:59I want to better understand your intuition of being more skeptical than some about the best scaling picture that these changes are already been needed in the first place, or that it would take more than two orders of magnitude, more improvement to get these things almost certainly to human level or very high probability to human level. So is it that you don't agree with the way in which they're extrapolating these lost curves, you don't agree with the implication that that decrease in loss will equate to greater and greater intelligence. Or like, what would you tell Dario about what we, if you were having, I'm sure you have, but like, what would that debate look like about this?
35:37Yeah. So again, here we're talking two factors of a half one on like, is it smart enough and one on like two of you a bunch of slap, even if like in some sense it's smart enough. And like the first factor of a half, I'd be like, I don't know, I don't think we have really anything good to extrapolate. That is like, I feel, I would not be surprised if I have like similar or maybe even higher their probabilities on a really crazy stuff over the next year. And then lower probabilities, not that bunched up. Maybe Dar is probability. I don't know. Talk with him. You have to talk with him. There's more bunched up on some particular year.
36:05And mine is maybe a little bit more uniformly spread out across the coming years. Partly because I'm just like, I don't think we have some trends we can extrapolate. We can extrapolate loss. You can look at your qualitative impressions of systems at various scales. But it's just very hard to relate any of those extrapolations to doing cognitive work or accelerating R &D or taking over and fully automating R &D. So I have a lot of uncertainty around that extrapolation. I think it's very easy to get down to a 50 -50 chance of this. What about the basic intuition that, listen, this is a big blob of compute, you make the big blob of compute bigger, it's gonna get smarter.
36:41It would be really weird if it didn't. Yeah, I'm happy with that. It's gonna get smarter, and it would be really weird if it didn't. And the question is, how smart does it have to get? Like that argument doesn't have yet, give us a quantitative guide to like at what scale is it? Is it a slam dunk or what scale is it 50 -50? And what would be the piece of evidence that would not do one where or another where you look at that and be like, ah, fuck, this is, it will eat at 20 % by 2040, or the 60 % by 2040 or something. Is there something that could happen in the next few years or next three? Like what is the thing you're looking to where this will be a big update for you?
37:12Again, I think there's some just how capable is each model where I like have, I think we're really bad extrapolating be still some subjective guess and you're comparing it to what happened and that one with me like every time and we see what happens with another order of magnitude of training compute. I will have a slightly different guess for where things are going. These probabilities are course enough that again, I don't know if that 40 % is real, or if post GB3 .5 and 4, I should be at 60%, or what. That's one thing. And the second thing is just like, if there was some ability to extrapolate, I think this could reduce air bars a lot.
37:40I think, here's another way you could try and do an extrapolation is you could just say, how much economic value do systems produce, and how fast is that growing? I think once you have systems actually doing jobs, the extrapolation gets easier because you're not moving from a subjective impression of a chat to automating all R &D, or moving from automating this job to automating that job, or whatever. Unfortunately, that's probably by the time you have nice trends from that. You're not talking about 20, 40, you're talking about two years from the end of days, or one year from the end of days, or whatever.
38:08But to the extent that you can get extrapolations like that, I do think it can provide more clarity. But why is economic value the thing we would Like, for example, you started off with chimps and they're just getting gradually smarter to human level. They would basically provide like no economic value until they were basically worth as much as a human. So it would be these very gradual and then very fast increase in their value. So is the increase in value from GP4, GP5, GP6? Is that the extrapolation we want? Yeah, I think that the economic extrapolation is not great. I think it's like you could compare it to this objective extrapolation of how smart is the model scene.
38:45It's not super clear which one's better. I think probably in the chimp case, I don't think that's quite right. I think if you actually like, so if you imagine like intensely domesticated chimps who are just like actually trying their best to be really useful in poise, and like you hold fixed their physical hardware, and then you just gradually scale up their intelligence. I don't think you're gonna see like zero value, which then suddenly becomes massive value over like one doubling of brain size or whatever one order of magnitude of brain size. It's actually possible in order of mind to your brain size.
39:14But like, chimps are very, chimps are already within an order of magnitude, brain side is the humans. Like chimps are very, very close on the kind of spectrum we're talking about. So I think like I'm skeptical of like the abrupt transition for chimps. And to the extent that I kind of expect a fairly abrupt transition here, it's mostly just because like the chimp human intelligence difference is like so small compared to the differences we're talking about with respect to these models. That is like, I would not be surprised if in some objective sense like chimp human difference is like significantly smaller than the GPT -3, GPT -4 difference, the GPT -4, GPT -5 difference.
39:43Wait, what in that argue in favor of just relying which more on the subjective? Yeah, there's two balancing tensions here. One is, I don't believe the Chimp thing is gonna be as abrupt, that is, I think if you scaled up from Chimp's to humans, you actually see quite large economic value from the fully domesticated Chimp already. Oh, yeah. And then the second half is, yeah, I think that the Chimp human difference is probably pretty small compared to all the differences. So I do think things are gonna be pretty abrupt. I think the economic extrapolation is pretty rough. I also think the subjective extrapolation is pretty rough, just because I really don't know how to get.
40:14Like, how do I don't know how people do the extrapolation end up with the degrees of confidence people end up with? Again, I'm putting it pretty high. If I'm saying, like, give me three years. And I'm like, yeah, 50 -50, it's going to have like basically the smarts there to do the thing. That's like, I'm not saying it's like a really long way off. Like, I'm just saying, like, I got pretty big error bars. And I think that like, it's really hard not to have really big error bars when you're doing this. I looked at GPD4. It seemed pretty smart compared to GPD3 .5. So I bet just like four more such notches and we're there, it's just a hard call to make.
40:46I think I sympathize more with people who are like, how could it not happen in three years than with people who are like, no way it's gonna happen in eight years or whatever, which is probably a more common perspective in the world, but also things do take longer than you. I think things take longer than you think. It's like a real thing. Yeah, I don't know. Mostly I have bigger bars because I just don't believe the subjective extrapolation that much. I find it hard to get a huge amount out of it. Okay, so what about the scaling picture do you think is most likely to be wrong? Yeah, so we've talked a little bit about how good is the qualitative extrapolation, how good are people at comparing?
41:18So this is not like the picture being qualitative wrong. This is just quantitatively. It's very hard to know how far off you are. I think a qualitative consideration, because significantly slow things down, is just like right now you get to observe this like really rich supervision from basically next word prediction, or like in practice maybe you're looking like a couple sentences prediction. So you're getting this like pretty rich supervision. It's plausible that if you want to automate long horizon tasks, like being an employee over the course of a month, that that's actually just considerably harder to supervise, or that you basically end up driving costs.
41:49The worst case here is that you drive up costs by a factor that's linear in the horizon, which the thing is operating. And I still consider that just quite plausible. Can you dump that down? You're driving a cost about what in the linear in the horizon? What does that horizon mean? Yeah, so if you imagine you want to train a system to say words that sound like the next word a human would say. They can get this really rich supervision by having a bunch of words and then predicting the next one. I'm going to tweak the model so it predicts better. If you're like, hey, here's what I want. I want my model to interact with some job over the course of a month.
42:23And then at the end of that month, I've internalized everything with the human world internalized about how to do that job well and how local context and so on. It's harder to supervise that task. So in particular, you could supervise it from the next word prediction task. and all that context the human has, ultimately, we'll just help them predict the next word better. So in some sense, a really long context language model is also learning to do that task. But the number of effective data points you get of that task is vastly smaller than the number of effective data points you get at this very short horizon.
42:51What's the next word with the next sense tasks? The sample efficiency matters more for economically valuable, long horizon tasks than the predicting the next token. And that's what will actually be required to take over a lot of jobs. Yeah, something, something like that. That is, it just seems very plausible that it takes longer to train models to do tasks that are longer horizon. How fast do you think the pace of algorithmic advances will be? Because if by 2040, even if scaling fails, I mean, you know, since 2012, since the beginning of the deep learning revolution, we've had so many new things.
43:27By 2040, are you expecting a similar pace of increases? And if so, then I mean, if we just keep having things like this, then aren't we gonna just gonna get the AI sooner or later? Or soon, not later. Are we gonna get AI sooner or sooner? I'm with you on sooner or later. I suspect like progress to slow. If you like held fixed how many people working in the field, I would expect progress to slow as low heat food is exhausted. I think like rapid rate of progress in like say, language modeling over the last four years is largely sustained by like, You start from a relatively small amount of investment, you like greatly scale up the amount of investment, and that enables you to like keep picking, you know every time, every time the difficulty doubles, you just double the size of the field.
44:11Like I think that dynamic can hold up for some time longer. Like I'm in a pretty good, like, you know right now if you think of it as like hundreds of people effectively searching for things, like up from like, you know, anyway if you think of it as hundreds of people now, you can maybe bring that up to like tens of thousands of people or something. So for a while you can just continue increasing the size of the fields and like search harder and harder. And there was indeed a huge amount of low hanging fruit where it wouldn't be a hard for a person to sit around and make things a couple percent better after a year of work or whatever.
44:37So I don't know. I would probably think of it mostly in terms of how much can investment be expanded and try and guess some combination of fitting that curve. And yeah, trying some combination of fitting the curve to historical progress, looking at how much low hanging fruit there is, getting a sense of how fast it decays. I think you probably get a lot though. You get a bunch of orders of magnitude of total, especially if you ask how good is the GPT -5 scale model or GPT -4 scale model? I think you probably get like bit 2040, like, I don't know, three orders of magnitude of effective training compute improvement or like a good chunk of effective training compute improvement.
45:14Four orders of magnitude. I don't know. I don't have like, here I'm speaking from like new private information about the last like couple of years of efficiency improvements. and since those people who are on the ground will have better senses of exactly how rapid returns are, and so on. OK, let me back up and ask a question we're generally about. People make these analogies about humans portraying bio -evaluation, and we're like deployed in the modern civilization. Dubai, those analogies, is about to say that humans were trained by evolution rather than, I mean, if you look at the protein coding size of the genome, it's like 50 megabytes or something.
45:50and then what part of that is for the brain. Anyways, how do you think about how much information is in, like do you think of the genome as hyperparameters or how much is that inform you when you have these anchors for how much training humans get when they're just consuming information when they're walking up and about and so on? Okay, I guess the way that you could think of this is, I think both analogies are reasonable. One analogy being like evolution is like a training run and humans like the unproduct of that training run and the second analogy is like evolution is like an algorithm designer and the human over the course of like this modest amount of computation over their lifetime is the algorithm being that's been produced, the learning algorithm's been produced.
46:30And I think like neither analogy is that great. Like I like them both and lean on them a bunch, like both of them a bunch. And I think that's been like pretty good for having like a reasonable view of what's likely to happen. That said, like the human genome is not that much like 100 trillion parameter model. It's like a much smaller number of parameters that behave in like a much more confusing way. Evolution did a lot more optimization, especially over long designing your brain to work well over a lifetime than gradient descent does over models. That's a disanalogy on that side. On the other side, I think human learning over the course of human lifetime is in many ways just much, much better than gradient descent over the space of neural nets.
47:09Grad descent is working really well, but we can just be quite confident that in a lot of ways, human learning is much better. Human learning is also constrained. We just don't get to see much data, and that's just an engineering constraint that you can relax, you can just give your neural nuts way more data than humans have access to. And what ways is human learning superior to grading design? I mean, the most obvious one is just like, ask how much data it takes the human to become like an expert in some domain. And it's like much, much smaller than the amount of data that's going to be needed on any plausible trend extrapolation.
47:37Like, not in terms of performance, but is it the active learning part? Is it the structure? Like what is it? I mean, I would guess a complicated mess of a lot of things. And sometimes there's not that much going on in a brain, as you say, it's not that many bytes in a genome. But there's very, very few bytes in an ML algorithm. If you think a genome is like a billion bytes or whatever, maybe you think less, maybe you think it's 100 million bytes. Then an ML algorithm is like, if compressed, probably more like hundreds of thousands of bytes or something, the total complexity of here's how you train GPT -4s, just like, I haven't thought about these numbers.
48:14that it's very, very small compared to genome. And so although a genome is very simple, it's very, very complicated to compare algorithms to humans' design. Like really hideous, more complicated than algorithmy human would design. Is that true? So the human genome is 3 billion base pairs or something. But only really, like 1 % or 2 % of that is protein coding. So that's 50 million base pairs. I don't know much about biology. In particular, I guess the question is how many of those bits are productive for shaping development of a brain? and presumably a significant part of the non -protein coding genome.
48:47I mean, I just don't know. It seems really hard to guess how much of that plays a role. The most important decisions are probably from an algorithm design perspective are not like the protein coding part is less important than the decisions about what happens during development or how cells differentiate. I don't know if that's... I don't know anything about biologists I would expect, but I'm happy to run with 100 million days per step. On the other end, on the hyperheromers that are cheap to a fraternity run, that might be not that much, But if you're going to include all the base pairs in the genome, then which are not all relevant to the brains, or are relevant to very bigger details about just the basic symbology.
49:25You probably include the Python library and the compilers and the operating system for GPT -4 as well to make that comparison analogous. So at the end of the day, I actually don't know which one has storing more information. Yeah, I mean, I think the way I would put it is like the number of bits it takes to specify the learning hour, where them to change the B2 -4 is like very small and you might wonder like maybe a genome like the number of bits It would like take to specify a brain is also very small the genome is much, much faster than that But it is also just plausible that a genome is like closer to like certainly the space the amount of space to put complexity in a genome We could ask how well evolution uses it and like I have no idea whatsoever But the amount of space in a genome is like very, very vast compared to the number of bits that are actually taken to specify to justify the architecture or optimization procedure and so on for GPT -4.
50:12Just because again, genome is simple, but algorithms are really very simple, and all algorithms are really very simple. And stepping back, you think this is where the better sample efficiency of human learning comes from? Like why it's better than gradient descent? Yeah, so I haven't thought that much about the sample efficiency question a long time. But if you thought like SNFs was seeing something like, you know, a neuron firing once per second, then how many seconds are there in a human life? We can just flip a calculator over. Yeah, it's too simculating. Okay, tell me the number. 3600 seconds per hour times 24 times 365 times 20.
50:50Okay, so that's 630 million seconds. That means like the average synapse is the, like, 630 million, I don't know exactly what the numbers are, but something is ballpark. like what's called a billion action potentials. And then there's some resolution. You should have carry some bits, but let's say it carries 10 bits or something. Just from timing information at the resolution you have available, then you're looking at 10 billion bits. So each parameter is kind of like, how much is a parameter seeing? It's not seeing that much. So then you can compare that to language. I think that's probably less than current language model see, and current language models are.
51:25So it's not clear of a huge gap here, but I think it's pretty clear you're gonna have a gap of at least three or four as a magnitude. It didn't drive due the lifetime anchors where she said, the amount of bytes that a human will see in their lifetime was 1 ,824 or something. The number of bytes of human will see is 1 ,824. Mostly this was organized around total operations performed in a brain, right? Oh, okay, never mind, sorry. Yeah, so I think that the story there would be like, a brain is just in some other part of the parameter space where it's like using a lot of compute for each piece of data gets and just not seeing very much data in total.
51:58Yeah, it's not really plausible if you extrapolate out language models, you're gonna end up with a performance profile similar to a brain. I don't know how much better it is. So I did this random investigation at one point where I was like, how good are things made by evolution compared to things made by humans? Right. Which is a pretty insane, seeming exercise. But I don't know. It seems like orders of magnitude is typical. Not tens of orders of magnitude, not factors of two. Things by humans are going to a thousand times more expensive to make, or a thousand times heavier pre -unit performance, if you look at things like how good are solar panels relative to leaves or how good are muscles relative to motors or how good are livers relative to systems that perform analogous chemical reactions and industrial settings.
52:35Were there consistent number of orders of magnitude in these different systems or was it all up in the place? So like a very rough ballpark, it was like sort of, for the most extreme things, you were looking at like five or six orders of magnitude and that would especially come in like an energy cost of manufacturing, where like Bobby's just very good at building complicated organs like extremely cheaply. And then for other things like leaves or eyeballs or livers or whatever, you tend to see more like if you set aside manufacturing costs and just look at like operating costs or like performance trade offs, like I don't know more like three orders of magnitude or something like that.
53:10Or there are some things that are on the smaller scale like the nano machines or whatever that we can't do at all, right? Yeah, that's, I mean, yeah. So it's a little bit hard to say exactly what the task definition is there. Like you could say like making a bow and we can't make a bow and you could try and compare a bow and the performance characteristics of a bow and something else Like we can't make spider silk do you try and compare the performance characteristics of spider silk like things that we can't synthesize? The reason this would be is why that evolution has had more time to Design these systems or I don't know.
53:37I just mostly just curious about like what the performance I think like most people would object to be like how did you choose these reference classes of things that are like fair intersections Some of them seem reasonable like eyes versus cameras seems like just everyone needs eyes Everyone needs cameras. It feels very fair photosynthesis seems like a very reasonable. Everyone needs to take solar energy and then turn it into a usable form of energy. But I don't really have a mechanistic story. Evolution in principle has spent way more time than we have designing. It's absolutely unclear how that's gonna shake out.
54:05My guess would be in general. I think there aren't that many things where humans really crush evolution where you can't tell a pretty simple story about why. So for example, roads and moving over roads with wheels crushes evolution, but it's not like an animal would have wanted to design a wheel. You're just not allowed to pave the world and then put things on wheels if you're an animal. Maybe planes or whatever. There's various things you could try and tell. There's some things you can do better, but it's not pretty clear why humans are able to win when humans are able to win. The point of all this was it's not that surprising to me.
54:32I think this is mostly a pro short -time lens view. It's not that surprising to me if you tell me machine learning systems are three or fours of magnitude less efficient at learning than human brains. I'm like, that actually seems like kind of indistribution for other stuff. And if that's your view, then I think you're probably going to hit. Then you're looking at 10 of the 27 training compute or something like that, which is not so far. We'll get back to the timeline stuff in a second. At some point, we should talk about alignment. So let's talk about alignment. At what stage does misalignment happen?
55:02So right now with something like GPD4, I'm not even sure it would make sense to say that it's misaligned, because it's not aligned to anything in particular. Is it at a human level where you think the ability to be deceptive comes about? What is a process by which misalignment happens? I think even for GPT -4, it's reasonable to ask questions like, are there cases where GPT -4 knows that humans don't want X, but it does X anyway? Like where it's like, well, I know that I can give this answer, which is misleading, and if it was explained to a human what was happening, they wouldn't want that to be done, but I'm going to produce it.
55:36I think that like ZPT4 understands things enough that you can have like that misalignment in that sense. Yeah, I think GPT, like I've sometimes talked about being like benign instead of aligned meaning that like well It's not exactly clear if it's aligned or if that context is meaningful It's just like kind of a messy word to use in general But I think we're more confident of is it's like not doing you know, it's not Optimizing for this goal which is like across purposes to humans is either optimizing for nothing or like maybe it's optimizing for what humans want or close enough or something it's like an approximation good enough to still not take over.
56:06But anyway, some of these abstractions seem like they do apply to GPT -4. It seems like probably it's not like egregiously misaligned. It doesn't, it's not doing the kind of thing that could lead to take over, we'd guess. I suppose you have a system at some point and which ends up in it wanting take over. What are the checkpoints? And also, what is the internal, is it just that it to become more powerful in these agency and agency and blinds other goals or do you see a different process by which misalignment happens? Yeah, so I think there's a couple possible stories for getting to catastrophic misalignment and they have slightly different answers to this question.
56:38So maybe I'll just briefly describe two stories and try and talk about when they can, when they start making sense to me. So one type of story is you train or fine tune your AI system to do things the teamman's will rate highly or that like get other kinds of reward in a broad diversity of situations. And then it learns to in general drops in some new situation, try and figure out which actions would receive a high reward or whatever, and then take those actions, and then when deployed in the real world, like sort of gaining control of its own training data provision process is something that gets a very high reward.
57:12And so it does that. So this is like one kind of story, like it wants to grab the reward button or whatever. It wants to intimidate the humans into giving a high reward, et cetera. I think that doesn't really require that much. This basically requires a system which is like, in fact, looks at a bunch of environments, is able to understand the mechanism of reward provision as a common feature of those environments, is able to think in some nominal environment. Which actions would result in beginning a high reward? And it's thinking about that concept precisely enough that when it says high reward, it's saying, okay, well, how is reward actually computed?
57:45It's some actual physical process being implemented in the world. My guess would be like, GPT -4 is about at the level where with hand -holding, you can observe this kind of scary generalizations of this type although I think they haven't been shown basically. That is, you can have a system which in fact is fine to know a bunch of cases. In some new case, we'll try and do an end run around humans, even in a way humans would penalize if they were able to notice it or would have penalized in training environments. So I think GPT -4 is kind of at the boundary where these things are possible. Examples kind of exist but are getting significantly better over time.
58:19I'm very excited about it because this is a centiropic project basically trying to see how good example can you make now of this phenomena. And I think the answer is like kind of okay probably. So that just I think is gonna continuously get better from here. I think for the level where we're concerned, like this is related to me having really broad distributions of our how smart models are. I think it's like not out of the question that you take GP, like GPT -4 is understanding of the world is like much crisper and like much better than GPT -3 is understanding. Just like it's really like night and day.
58:48And so it would not be that crazy to me if you took GPT -5 and you trained it to get a bunch of reward and it was actually like, okay, my goal is not doing the kind of thing which like thematically looks nice to humans. My goal is getting a bunch of reward and then we'll generalize in a new situation to get reward. And by the way, this requires to consciously want to do something that it knows the humans wouldn't want it to do. Or is it just that we weren't good enough to specify that the thing that we accidentally ended up rewarding is not what we actually want. I think the scenarios I am most interested in and most people are concerned about from a catastrophic risk perspective.
59:24Involved systems understanding that they are taking actions which a human would penalize if the human was aware of what's going on, such that you have to either deceive humans about what's happening, or you need to like actively subvert human attempts to correct your behavior. So these, the failures come from really this combination or they require this combination of both like trying to do something humans don't like and and understanding the humans would stop you. I think you can have only the barest examples. You can have the barest examples for GPT -4. You can create the situations where GPT -4 will be like, sure, and that situation, here's what I would do.
59:52I would go hack the computer and change my reward, or in fact, we'll do things that are simple hacks, or go change the source of this file, or whatever, to get a higher reward. They're pretty weak examples. I think it's plausible GPT -5 will have compiling examples of this phenomena. I really don't know. This is very related to the very broad error bars on how competent such systems will be when. That's all with respect to this first mode of like a system is taking actions that get reward and like overpowering or receiving humans is helpful for getting reward. There's this other failure mode, another family failure mode where AI systems want something potentially unrelated to reward.
1:00:26I understand that like they are being trained and like while you're being trained there are a bunch of like reasons you might want to do the kinds of things humans want you to do. But then when deployed in the real world, if you're able to realize you're no longer being trained, you no longer have reason to do the kinds of things you want. You'd prefer to be able to determine your own destiny, like control your own, you're competing hardware, etc. Which I think probably emerged a little bit later than systems that try and get reward. And so we'll generalize and scary, unpredictable ways to new situations.
1:00:54I don't know when those appear. But also, again, brought in a fairer bars that it's conceivable for systems in the near future. I wouldn't put it less than 1 ,000 for GPT -5, certainly. If we deployed all the AI systems and some of them are award hacking, some of them are deceptive, some of the mergers normal, whatever. How do you imagine that they might interact with each other at the expense of humans? How hard do you think it would be to for them to communicate in ways that we would not be able to recognize and coordinate at their expense? Yeah, I think that most realistic failures probably involve two factories interacting.
1:01:28One factor is the world is pretty complicated and the humans mostly don't understand what's happening. So, like, AI systems are writing code that's very hard for humans to understand. and maybe how it works at all, but more likely, than understand roughly how it works, but there's a lot of complicated interactions. AI systems are running businesses that interact primarily with other AI's. They're doing SEO for AI search processes. They're running financial transactions, thinking about a trade with AI counterparties. And so you can have this world where even a few of us understand the jumping off point when this was all humans, actual considerations of what's a good decision, what code is going to work well and be durable, or what marketing strategies effective for selling to these other AI's, or whatever, is kind of just all mostly outside of sort of humans' understanding.
1:02:10I think this is like a really important, again, when I think of like the most plausible scary scenarios, I think that's like one of the two big risk factors. And so in some sense, your first problem here is like having these added systems to understand a bunch about what's happening, and your only lever is like, hey, I do something that works well. So you don't have a lever to be like, hey, do what I really want. You just have the system, you don't really understand. You can observe some outputs, like, did it make money, and you're just optimizing or at least doing some fine tuning to get the AI to use its understanding of that system to achieve that goal.
1:02:37So I think that's like your first risk factor. And like once you're in that world, then I think there are like all kinds of dynamics amongst AI systems. But again, humans aren't really observing. Humans can't really understand. Humans aren't really exerting any direct pressure on only on outcomes. And then I think it's quite easy to be in a position where, you know, if AI systems started failing, it would be very, they could do a lot of harm very quickly. Humans aren't really able to like prepare for and mitigate that potential harm because we don't really understand the systems in which they're acting.
1:03:02And then if AI systems, they could successfully prevent humans from either understanding what's going on or from like, like, taking the data centers or whatever, if the AI successfully grab control. This seems like a much more gradual story than the conventional takeover stories where you're just like, you train it and then it comes alive and escapes and takes over everything. So you think that kind of story is less likely than one in which we just hand off more control voluntarily to the AIs. So one, I am interested in the tale of some risks that can occur particularly soon. And I think risks that occur particularly soon are a little bit like you have a world where it has not probably deployed and then something crazy happens quickly.
1:03:41That said, if you ask what's the median scenario where things go badly, I think it is like there's some lessening of our understanding of the world. I think in the default path, it's like a very clear to humans that they have increasingly little grip on what's happening. I mean, I think already most humans are very little grip on what's happening. It's just some other humans understand what's happening. I don't know how almost any of the systems I interact with work in a very detailed way. So it's sort of clear to humanity as a whole that we sort of collectively don't understand most of what's happening except with AI assistants.
1:04:06And then that process just continues for a fair amount of time. And then there's a question of how abrupt an actual failure is. And I do think it's reasonably likely that a failure itself would be abrupt. Like at some point, bad stuff starts happening that a human can recognize as bad. And once things are that are obviously bad start happening, then you have this bifurcation where either humans can use that to fix it and say, OK, I behavior the led to this obviously bad stuff. Don't do more of that. Or you can't fix it. and then you're in this rapidly escalating failures everything goes off the rails.
1:04:33In that case, yeah, what is going off the rails look like? For example, how would it take over the government? Yeah, it's getting deployed in the economy in the world and at some point it's in charge. What does that transition happen? Yeah, so this is going to depend a lot on what kind of timeline you're imagining or like this sort of a broad distribution but I can fill in some random concrete option that is like an itself very improbable. Yeah, I think that one of the less dignified than maybe more plausible routes is you just have a lot of AI control over critical systems even in running a military.
1:05:09And then you have the scenario that's a little bit more just like a normal coup where you have a bunch of AI systems, they in fact operate. It's not the case that humans can really fight a war on their own. It's not the case that humans could defend them from an invasion on their own. So that is if you had invading army and you had your own robot army, you can't just be like we're gonna turn off the robots now because things are going wrong if you're in the middle of a war. Okay, so how much does this world rely on race, day, and dynamics, where we're forced to deploy, or not forced, but we choose to deploy AI's because other countries or other companies are also deploying AI's.
1:05:42And you can't have them have other Kayla robots. Yeah, I mean, I think that like, there's several levels of answer to that question. So one is like, maybe three, three parts of my answer. Like our first part is like, I'm just trying to tell like what seems like the most likely story. I do think there's like further failures to get you in the more distant future. It's like a G. L. A. Zer will not talk that much about killer robots because you really want to emphasize, hey, if you never built a killer robot, something crazy is still going to happen to you, just only four months later or whatever.
1:06:08So it's not really the way to analyze the failure. But if you want to ask what's the median world or something bad happens, I still do think this is the best guess. Okay, so that's part one of my answer. Part two of the answer was in this proximal situation where something bad is happening. You ask, hey, why do humans not turn off the AI? You can imagine two kinds of story. One is like they are able to prevent humans from turning them off them off. And the other is like, in fact, we live in a world where it's like incredibly challenging. Like there's a bunch of competitive dynamics or a bunch of reliance on AI systems.
1:06:35And so it's incredibly expensive to turn off AI systems. I think again, you would eventually have the first problem. Like eventually AI systems could just prevent humans from turning them off. But I think like in practice, the one that's going to happen much, much sooner is probably competition amongst different actors using AI. And it's like a very, very expensive to unilaterally disarm. You can't be like, something weird has happened, we're just going to shut off all the AI because you're EG in a hot war. So again, I think that's just like, probably the most likely thing to happen first. Things would go badly without it, but I think if you ask why don't we turn off the AI?
1:07:03My best guess is because there are a bunch of other AI as running around 2 -D or lunch. So how much better association would we be? And if there was only one group that was pursuing AI, I know other countries, know other companies, basically how much of the expected value is lost from the dynamics that are likely to come about because other people will be developing and deploying these systems. Yeah, so I guess this brings you to like a third part of the way I'm which competitive dynamics are relevant. So it's both the question of can you turn off AI systems in response to something bad happening where competitive dynamics may make it hard to turn off?
1:07:36There's a further question of just like why were you deploying systems which you had very little ability to control or understand those systems? And again, it's possible you just don't understand what's going on. You think you can understand or control such systems, but I think in practice this is a significant part is going to be like, you are doing the calculus, so people deploying systems are doing the calculus as they do today, like in many cases, overtly, of like, look, these systems are not very well controlled or understood. There's some chance of like something going wrong or at least going wrong if we continue down this path, but other people are developing the technology potentially in even more reckless ways.
1:08:06So in addition to like competition making it difficult to shut down AI systems and the event of a catastrophe, I also think it's just like the easiest way that people end up pushing relatively quickly or moving quickly ahead on a technology where they feel kind of bad about understandability or controllability. That could be economic competition or military competition or whatever. So I kind of think ultimately, like most of the harm comes from the fact that like lots of people can develop AI. How hard is a takeover of the government or something? From any, even if it doesn't have killer robots, but just a thing that you can't kill off if it has seeds elsewhere, can easily replicate, can think a lot and think fast.
1:08:44What is the minimum viable cool for, is it like shutting up, just like threatening a biowore or something or shutting off the grid, how easy is it basically to take over human civilization? So again, there's going to be a lot of scenarios and I'll just like start by talking about one scenario, which will represent a tiny fraction of probability or whatever. But like, so if you're not in this competitive world, if you're saying like we're actually slowing down deployment of AI because we think it's unsafe or whatever, then in some sense, Once you're creating this very fundamental instability where you could have been making faster AI progress and you could have been deploying AI faster.
1:09:20So in that world, the bad thing that happens if you have an AI system that wants to mess with you is the AI system says, I don't have any compunctions about rapid deployment of AI or rapid AI progress. So the thing you want to do or the AI wants to do is to say, I'm going to defect from this regime. All the humans have agreed that we're not deploying AI in ways that would be dangerous. But if I as an AI can escape and just go set up my own shop, make a bunch of copies of myself. Maybe the humans didn't want to like delegate war fighting to an AI, but I as an AI I'm pretty happy doing so Like I'm happy if I'm able to grab some military equipment or direct some humans to use the AI Use myself to direct it and so I think like as that gap grows So if people are deliberately right if people are deploying AI everywhere I think of this competitive dynamic if people aren't deploying AI everywhere So like if countries are not happy deploying AI in these high -stakes settings Then as AI improves you create this like wedge the grows where like if you were in the position of fighting against an AI which wasn't constrained in this way, you'd be in a pretty bad position.
1:10:16At some point, even if you just, yeah, so that's like one important thing, just like I think in conflict, and like overt conflict. If humans are putting the brakes on AI, they're like a pretty major disadvantage compared to an AI system that can kind of set up shop and operate independently from humans. A potential independent AI. Does it need collaboration from a human faction? Again, you can tell different stories, but it seems so much easier. At some point you don't need any. At some point, AI system can just operate completely like out of human supervision or something. But that's like so far after the point where it's like so much easier.
1:10:48If you're just like there are a bunch of humans, they don't love each other that much. Like some humans are happy to be on side. They're skeptical about risk. We're happy to make this trade or can be fooled or can be cursed or whatever. And just seems like it is almost certainly, almost certainly the easiest first pass is going to involve like having a bunch of humans who are happy to work with you. So yeah, I think that probably is about it. I think it's not necessary, but if you ask about the median scenario, it involves a bunch of humans working with AI systems. Either being directed by AI systems, providing computer AI systems, providing legal cover and jurisdictions that are sympathetic to AI systems.
1:11:21Humans presumably would not be willing, if they knew the end result of the AI takeover, would not be willing to help. So they have to be probably fooled in somebody, right? Like deep fakes or something. and what is the minimum viable physical presence they would need or jurisdiction they would need in order to carry out their schemes. Do you need a whole country? Do you just need a server farm? Do you just need like one single laptop? I think I'd probably start by pushing back a bit on the like humans wouldn't cooperate if they understood outcome or something. Like I would say like one, even if you're looking at something like tens of percent risk of takeover, humans may be fine with that.
1:11:55Like a fair number of humans may be fine with that. too, like if you're looking at certain takeover, but it's very unclear if that leads to death. Like a bunch of humans may be fine with that. Like if we're just talking about like, look, the AI systems are going to like run the world, but it's not clear if they're gonna murder people. Like how do you know? It's just a complicated question about AI psychology, and a lot of humans probably are fine with that, and I don't even know what the probability is there. But I think you actually have given that probability online. I've certainly guessed. Okay, but it's not zero.
1:12:19It's like a significant percentage. I gave like 50 -50. Oh, okay, yeah. Why is it, tell me about the world in which the AI takes over, but it doesn't kill humans. Why would that happen and what would that look like? I mean, I asked my questions like, why would you kill humans? So I think like, maybe I'd say the incentive to kill humans is like quite weak. They'll get it in your way. They control shit you want. So taking shit from humans is a different. Like marginalizing humans and like causing humans to be irrelevant is a very different story from killing the humans. I think I'd say like the actual incentives to kill the humans are quite weak.
1:12:51So I think like the big reasons you kill humans are like, well one, you might kill humans if you're like in a war with them. And it's hard to win the world without killing a bunch of humans. Like, maybe most saliently here if you want to use some biological weapons or some crazy shit. But I might just kill humans. I think you might kill humans just from totally destroying the ecosystems they're dependent on and slightly expensive to keep them alive anyway. You might kill humans just because you don't like them or you literally want to like, yeah, I mean, neutralize a threat or the alias or linus that they're made of atoms you could use or something else.
1:13:21Yeah, I think the literal they're made of atoms is like, quite, they're not many atoms in humans. Neutralize the threat is as a similar issue where it's just like, I think you would kill the humans if you didn't care at all about them. So maybe you're questioning us and just like, why would you care at all about them? But I think you don't have to care much to not kill the humans. Okay, sure. Because there's just so much raw resources elsewhere in the universe. Yeah, also, you can marginalize humans pretty hard. Like you could totally cripple human. You could cripple humans' war fighting capability and also take almost all their stuff while killing only, you know, a small fraction of humans incidentally.
1:13:55So then if you ask why in my day I now want to kill humans, I mean a big thing is just like, look, I think it has probably a lot of random crap for complicated reasons. Like the motivations of AI systems and civilizations of AI's are probably complicated messes. Certainly amongst humans, it is not that rare to be like, well, there was someone here, I would like all those equal if I didn't have to murder them, I would prefer not murder them. And my guess is it's also like, reasonable chance it's not that rare amongst AI systems. Humans have a bunch of different reasons we think that way. I think AI systems will be very different from humans, but it's also just a very salient.
1:14:28Yeah, I mean, I think this is a really complicated question. If you imagine drawing values from the basket of all values, what fraction of them are like, hey, if there's someone here, how much do I want to nominate them? And my guess is just, if you draw a bunch of values from the basket, that's a natural enough thing. If I wanted 10 ,000 different things, or you're still a decision, if you have one, it wants 10 ,000 different things. It's just reasonably likely you get some of that. The other salient reason you might not want to murder them is just like, well, yeah, there's some kind of crazy decision theory stuff or like, a causal trade stuff, which does look on paper like it should work.
1:14:57And like if I was running a civilization and like dealing with some people who I didn't like at all or like didn't have any concern for at all, but I could only had to spend 1 billionth of my resources not to murder them, I think it's like quite robust that you don't want to murder them. That is, I think, the weird decision theory it causes the trade stuff probably does carry the day. Oh, wait, that contributes more to that 50 -50 of will they murder or say they take over than the, we'll just by default, they might just not want to kill us. Yeah, I think they're both salient. Can you explain the, they run together with each other a lot?
1:15:32Before the audience can explain the weird it cause a little reason why am I not killed us? Maybe a higher level thing that goes into both of these and then I will talk about how you instantiate in a causal trait. It's just like, it matters a lot to the humans, not to get murdered. And the AI cares very, very little about whether, if we imagine a hypothetical, the reason it wants to kill humans is just total apathy. It cares very little about whether or not to murder humans, because it is so easy to marginalize humans without murdering them, and the resources required for humans survival are extremely low.
1:15:59Again, in the context of this rapid AI industrialization. So that's the basic setting. And now the thing that you'd like to do is run a trade. Like, I would like to say, like, hey, humans, you care a ton about not getting murdered. I don't really care one way or the other. I would like to, if I could, find some way in which I don't murder you and then in return, I get something. The problem is in that world, the humans have essentially nothing to give. That is, the humans are mostly irrelevant. However, from the perspective of any kind of reasonable AI, it's like not that sure whether it lives in the world where the humans in fact have nothing to give, or whether the humans, in fact, it lives in a world where the humans succeeded at building a land AI.
1:16:36I know the AI is simply running in a nice little simulation. Humans are wondering, I wonder if this AI would have murdered us if all of it had the chance. And he was saying, if it would murder us all if it had the chance, that sucks. We'd like to run this trade. We'd like to be nice to the AI's who wouldn't have murdered us all in order to create an incentive for AI's not to murder us. So we do as we as check. And for the kinds of AI's who don't murder everyone, we just give them 1 billionth of the universe. If you do that as humanity, then if you're an AI system considering, do I want to murder everyone?
1:17:02You're calculus is like, if this is my real chance to murder everyone, I get the time is bit of value. You know, I get like one trillionth of the value, whatever, one billionth of the value. But on the other hand, like, if I don't murder everyone, there's like some worlds within the humans will crack or determine it on murder everyone. Because in fact, the humans survived, the humans are running the simulations to understand what different it has would behave. And so like, that's a better deal. Let's hope they fall for that tie -off. Okay, that is interesting. Hey, real quick, this episode is sponsored by Open Philanthropy.
1:17:33Open Philanthropy is one of the largest grant -making organizations in the world. Every year, they give away hundreds of millions of dollars to reduce catastrophic risks from fast -moving advances in AI and biotechnology. Open Philanthropy is currently hiring for 22 different roles in those areas, including grant -making, research, and operations. New hires will support Open Philanthropies giving on technical AI safety, AI governance, AI policy in the US, EU and UK and biosecurity. Many rules are remote friendly and most of the grand making hires that Open Philanthropy makes don't have prior grand making experience.
1:18:17Previous technical experience is an asset, as many of these rules often benefit from a deep understanding of the technologies they address. For more information and to apply, please visit Open Philanthropy's website in the description. The deadline to apply is November 9th, so make sure to check out those roles before they close. Awesome, back to the episode. In a world where we're re -deploying these AI systems and suppose they're aligned, how hard would it be for competitors to, I don't know, cyber attack them and get them to join the other side. Are they robustly going to be aligned? Yeah, I mean, I think in some sense, so there's a bunch of questions that come up here.
1:19:03First one is like, are aligned AI systems that you can build? Like competitive? Are they almost as good as the best systems anyone could build? Maybe we're granting that for the purpose of this question. Yeah. And again, next question that comes up is like AI systems right now are very vulnerable to manipulation. Like, it's not clear how much more vulnerable they are than humans, except for the the fact that you can, like, if you have an AI system, you can just replay it like a billion times and search for, like, what thing can I say that will make it behave this way? So as a result, like AI systems are very vulnerable to manipulation.
1:19:30It's unclear if future AI systems will be similarly vulnerable to manipulation, but certainly seems plausible. And in particular, like, you know, aligned AI systems or unaligned AI systems would be vulnerable to all kinds of manipulation. The thing that's really relevant here is kind of like asymmetric manipulation or something. That is, like, if it is easier. So if everyone is just constantly messing with each other's AI systems, like if you ever use AI systems in a competitive environment, a big part of the game, is messing with your competitors AI systems. A big question is whether there's some asymmetric factor there where it's easier to push AI systems into a mode where they're behaving erratically or chaotically or trying to grab power or something than it is to push them to fight for the other side.
1:20:05It's just a game of two people competing and neither of them can hijack an opponent's AI to help support their cause. That doesn't mean it matters and it creates chaos and it might be quite bad for the world but it doesn't really affect the alignment calculus. Now it's just right now you have normal cyber -offensive, cyber defense, you have weird AI version of cyber offense cyber defense. But if you have this kind of asymmetrical thing, we're like, you know, a bunch of AI systems where we love AI flourishing, can then go in and say great AI is how about you join us and that works, if they can search for a persuasive argument to that effect, and that's kind of asymmetrical, then the effect is whatever values it's easiest to push, whatever it's easiest to argue to an AI that it should do, that is advantaged.
1:20:46So it may be very hard to build AI systems, trying to defend to an interest, but very easy to build the AI systems, just like trying to destroy stuff or whatever. Just depending on what is the easiest thing to argue to an AI that should do, or what's the easiest thing to trick an AI into doing or whatever. Yeah, I think if a line in this body, if you have the AI system which doesn't really want to help humans or whatever, or in fact wants some kind of random thing or wants different things in different contexts, then I do think adversarial settings will be the main ones where you see the system or the easiest ones where you see the system behaving really badly.
1:21:17And it's a little bit hard to tell how that shakes out. Okay, and so suppose it is more reliable. How concerned are you that whatever alignment technique you come up with, you know, you publish the paper, this is how the alignment works. How concerned are you that Putin reads it or China reads it? And now they understand, presently, the constitutionally, I think, anthropic and then you just write on there, oh, never contradict Mao Zedong thought or something. How concerned should we be that these alignment and techniques are universally applicable, not necessarily just for inline goals. Yeah, I think they're super universally applicable.
1:21:53I think it's just like, I mean, the rough way I would describe it, which I think is basically right, is like some degree of alignment makes the AI systems much more usable. Like, you kind of usually just think of the technology of AI as including like a basket of like some AI capabilities and some like getting the AI to do what you want, it's just part of that basket. And so anytime we're like, you know, to extend alignment as part of that basket, You're just contributing to all the other harms from AI. You're reducing the probability of this harm, but you are helping the technology basically work.
1:22:19And the basically working technology is kind of scary from a lot of perspectives, one of which is right now, even in a very authoritarian society, just humans have a lot of power because you need to rely on just a ton of humans to do your thing. And in a world where AI is very powerful, it is just much more possible to say, here's how our society runs. One person calls the shots and then a ton of AI systems do what they want. I think that's a reasonable thing to dislike about AI. and a reasonable reason to be scared to push the technology to be really good. But is that also a reasonable reason to be concerned about alignment as well?
1:22:49That this is in some sense also give you the ability to be good at what they want. Yeah, I mean, I would generalize, so we earlier touched a little bit on the more potential moral rights of AI systems. Now we're talking a little bit about how AI systems powerfully, I just empower humans and can empower authoritarianists. I think we could list other harms from AI. And I think it is the case that if line was bad enough, people would just not build AI systems. And so like, yeah, I think there's a real sense in which you should just be scared to say you're scared of all AI. You should be like, well, alignment, although it helps with one risk, does contribute to AI being more of a thing.
1:23:26I do think you should shut down the other parts of AI before like, if you were a policy maker or like a researcher or whatever looking at on this, I think it's like crazy to be like, this is the part of the basket we're going to remove. You should first remove like other parts of the basket because they're also part of the story of risk. Wait, does that imply you think, if, for example, all capabilities research shut down, that you think it'd be a bad idea to continue doing alignment research? In isolation of what is conventionally considered capability as a research? I mean, if you told me it was never going to restart, then it wouldn't matter.
1:23:55And if you told me it's going to restart, I guess it would be a kind of similar calculus to today. Whereas it's going to happen, so you should have something. Yeah, I think that like, in some sense, you're always going to face this trade off, where alignment makes it possible to play AI systems or makes it more attractive to the play -i systems and then or like in the authoritarian case makes it like tractable to apply them for this purpose. And like if you didn't do any alignment there'd be a nicer, bigger buffer between your society and malicious uses of AI. And like I think it's one of the most expensive ways to maintain that buffer.
1:24:26Like it's much better to maintain that buffer by not having the compute or not having the power play -i. But I think if you're concerned enough about the other risks there's definitely a case to be made for just like put in more buffer or something like that. I'm not, like I care enough about the takeover risk, that like I think it's just not in that positive way to buy buffer. That is like, the version of this that's most pragmatic is just like suppose you don't work on a line in today, like to increase his economic impact of AI systems, they'll be like less useful if they're less reliable and if they more often don't do what people want.
1:24:53And so you could be like great, that just buys time for AI and you're like getting some trade off there. We were like to increasing some risks of AI, like if AI is more reliable and more does what people want and is more understandable, then that cuts down some risks. But if you think AI is on balance bad, even apart from takeover risk, then the alignment stuff can easily end up being that negative. But presumably you don't think that, right? Because I guess this is something people have brought up to you because you met a Darl HF, which was used to train Chat Gbt and Chat Gbt, brought AI to the front pages everywhere.
1:25:28So, I wonder if you can measure how much more money went into AI because how much people have raised in the last year or something. But it's got to be billions, the counterfactual impact of that. That went into the AI investment and the talent that went into AI, for example. So presumably you think that was worth it. So I guess you're hedging here about what is the reason that it's worth it. Yeah, like what's the total trade off there? Yeah. I mean, I think my take is like, I think slower AI development on balance is quite good. I think that slowing AI development now, or having less press around chat GPT, is a little bit more mixed than slowing AI development overall.
1:26:08I think it's still probably positive, but much less positive, because I do think there's a real effect of the world is starting to get prepared, or it's getting prepared like a much greater rate now than it was prior to the release of chat GPT. If you can choose between progress now or progress later, you really prefer to have more of your progress now, which I do think slows down progress later. I don't think that's enough to flip the sign. I think maybe it wasn't the far enough past, but now I would still say like moving faster now is net negative But to be clear it's a lot less net negative than Nearly accelerating AI because I do think again the chat GPT thing I am glad people having policy discussions now rather than like delaying the like chat GPT wake up thing by year and then Having policy.
1:26:45Oh, a chat GPT was negative or our early chef was the negative So here just on the acceleration just like how is the press of chat GPT? My guess is like oh, yeah My guess isn't that negative, but I think it's not super clear. It's much less than slowing AI. Slowing AI is great if you could slow overall AI progress. I think slowing AI by causing, there's this issue where slowing AI now for chat GPT. You're building up this backlog. What does chat GPT make such a splash? I think people, there's a reasonable chance if you don't have a splash about chat GPT, a splash about GPT4. If you fail to have a splash about GPT4, there's a reasonable chance of a splash about GPT4.
1:27:19And just like as that happens later, there's just like less and less time between that splash and between when and AI potentially kills everyone. Right. So people, the governments are talking about if they are now and people are. But okay, so let's start with slowing down because. So this is also all one sub component of like the overall impacts. And I was just saying this to like briefly give the roadmap for the overall too long answer. Like, there's a question of what's the calculus for speeding up? I think speeding up is pretty rough. I think speeding up like locally is a little bit less rough.
1:27:48And then yeah, I think that the effect, like the overall effect size from like doing alignment work on reducing takeover risk versus speeding up AI is like pretty good. Like I think, yeah, I think it's pretty good. I think you reduce takeover risk significantly before you like speed up AI by a year, whatever. Okay, got it. If it's good to, like slowing down AI is good, presumably because it gives you more time to do alignment. But alignment also helps speed up AI. RLHF is alignment and it helps with CHGPT, which sped up AI. So I actually don't understand how the feedback loop nets out other than the fact that if AI is happening, you need to do alignment at some point, right?
1:28:30So I mean, you can't just not do alignment. Yeah, so I think if the only reason you thought faster AI progress was bad was because it gave less time to do alignment Then there would just be no possible way that the calculus comes out negative for alignment You're like maybe alignment speeds up AI but the only purpose of slowing down AI was to do it like it's right just It could never come out ahead. I think the reason that you can come out ahead the reason you could end up thinking the alignment was not negative Was because there is a bunch of other stuff you're doing that makes it I say for like if you think the world is like gradually Coming better to terms with the impact of AI or policies being made or like you're getting increasingly prepared to handle the threat of authoritarian abuse of AI.
1:29:04If you think other stuff is happening, that's improving preparedness, then you have reason beyond alignment research to slow down AI. Actually, how big a factor is that? So right now we hit pause and you have 10 years of no alignment, no capabilities, but just people get to talk about it for 10 years. How much more does that prepare people than we only have one year versus we have no time? Like, is just that time where no research in alignment or capabilities happening? Is there, what does that time do for us? I mean, right now it seems like there's a lot of policy stuff you'd want to do. This seemed like less plausible a couple of years ago maybe, but if the world just knew they had a 10 -year pause right now, I think there's a lot of sense of like, we have policy objectives to accomplish.
1:29:45If we had 10 years, we could pretty much do those things. We'd have a lot of time to debate measurement regimes, debate policy regimes and containment regimes, and a lot of time to set up those institutions. So if you told me the world knew it was a pause, It wasn't like people just see that AI progress isn't happening, but they're told like you guys have been granted or like cursed with a 10 year no AI progress no alignment progress pause. I think that would be quite good at this point However, I think it would be much better at this point than it would have been two years ago And so like the entire concern with like slowing AI development now rather than taking the 10 year pause is just like If you slow the AI development by a year now, my guess is some gets clawed back by looking if you get gets picked faster in the future my guess is you lose like half a year or something like that in the future.
1:30:26Maybe even more, maybe like two -thirds of a year. So it's like you're trading time now for time in the future at some rate. And it's just like that eats up like a lot of the value of the slowdown. And the crucial point being that time in the future matters more because you have more information, people are more bought in and so on. Yeah, the same reason like I'm more excited about policy change now than two years ago. So like my overall view is just like in the past this calculus, this calculus changes over time, right? The like more people are getting prepared, the better the calculus is for slowing down at this very moment.
1:30:53And I think now the calculus is, I would say, positive. For just even if you pause now, and it would get cloud back in the future, I think the pause now is just good. Because enough stuff is happening. We have enough idea of, probably even apart from alignment research, and certainly if you include alignment research. Just enough stuff is happening where the world is getting more ready and coming more to terms with impacts that I just think it is worth it, even though some of that time is going to get cloud back. Again, especially if there's a question of doing a pause does Nvidia keep making more GPUs?
1:31:22Like that sucks if they do. If you do a pause, but like, yeah, in practice, if you did a pause, then probably couldn't keep making more GPUs because in fact, like the demand for GPUs is really important for them to do that. But if you told me to just get to scale a part of our production and like building the clusters, but not doing AI, then that's back to being that negative, I think, pretty clearly. Then having brought up the fact that we want some sort of measurement scheme for these capabilities. Let's talk about responsible scaling policies. Do you want to introduce what this is? Sure. So, I guess the motivating question, it's like, what should AI labs be doing right now to manage risk and to sort of build good habits or practices for manage risk into the future?
1:32:01And, right, like I think my take is that current systems pose from a catastrophic risk perspective, not that much risk today. That is a failure to like control or understand GP2 -4 can have real harms, but doesn't have much harm with respect to the kind of take over with what I'm worried about or even much catastrophic harm with respect to misuse. So I think if you want to manage catastrophic harms, I think right now you don't need to be that careful with GB2 .4. And so to the extent you're like, what should labs do? I think the single most important thing seems like understand whether that's the case, notice when that stops being the case, have a reasonable roadmap for what you're actually going to do when that stops being the case.
1:32:43So that motivates this set of policies, which I've sort of been pushing for labs to adopt, which is saying, here's what we're looking for. Here's some threats we're concerned about. Here's some capabilities that we're measuring. Here's the level, here's the actual concrete measurement results, those suggest to us that those threats are real. Here's the action we would take in response to observing those capabilities. If we couldn't take those actions, like you do, if we said that we're going to secure the weights, we're not able to do that. that we're going to pause until we can take those actions.
1:33:14Yeah, so this sort of, again, I think it's like motivated primarily, but what should you be doing as a lab to manage catastrophic risk now, in a way that's like our reasonable precedent and habit and policy for continuing to implement into the future? And which labs, I don't know if there's public yet, but which labs are cooperating on this? Yeah, so I mean, the topic has written this document, just their current responsible scaling policy. And then have been talking with other folks, I guess don't really want to comment on other conversations. But I think in general, people who are more interested in, or more think you have plausible catastrophic harms on a five -year timeline or more interested in this, and there's not that long a list of suspects like that.
1:34:00There's not that many labs. Okay, so if these companies would be willing to coordinate and say at these different benchmarks we're going to make sure we have these safeguards. What happens, I mean there are other companies and other countries which care less about this. Are you just slowing down the companies that are most aligned? I think the first sort of business is understanding like sort of what is actually a reasonable set of policies for managing risk. I do think there's a question of like you might end up in a situation where you say like well here's what we would do in ideal world if we be like, if everyone was behaving responsibly, we'd want to keep risk to 1%, or a couple of percent, or whatever, maybe even lower levels, depending on how you feel.
1:34:44However, in the real world, there's enough of a mess. There's enough unsafe stuff happening that actually it's worth making larger compromises. Or we don't kill everyone. Someone else will kill everyone anyway. So actually, the counterfactual risk is much lower. I think if you end up in that situation, it's still extremely valuable to have said, here's the policies we'd like to follow. Here's the policies we've started following. Here's why we think it's dangerous. Here's the concerns we have if people are following significant the lack of policies. And then this is maybe helpful as an input to our model for potential regulation.
1:35:14It's helpful for being able to just produce clarity about what's going on. I think historically there's been considerable concern about developers being more or less safe, but there's not that much legible differentiation in terms of what their policies are. I think getting to that world would be good. It would be very different. It's a very different world if you're like actor access is developing it, and I'm concerned that they will do so in an unsafe way. Versus if you're like, look, we take security precautions or safety precautions, x, y, z. Here's why we think those precautions are desirable or necessary.
1:35:42We're concerned about this other developer because they don't do those things. I think it's just like a qualitatively, it's kind of the first step you would want to take in any world where you're trying to get people on -side or like trying to move towards regulation to come and risk. Well, how about the concern that you have these evaluations and let's say you declared the world. Our new model has a capability to help develop bio -eapons or help you make cyber attacks. And therefore, we're pausing right now until you can figure this out. And China hears this and thinks, wow, a tool they can help us make cyber attacks and then just steals the weights.
1:36:18Does this scheme work in the current regime where we can't ensure that China doesn't and just seal the weights. And more so, are you increasing the salience of dangerous models so that you just, you blur this out and then people want the weights now because they know what they can do. I mean, I think the general discussion does emphasize potential harms or potential, I mean, some of those are harms and some of those are just like impacts that are very large and so might also be an inducement to develop models. I think like that part, if you're for a moment ignoring security and just saying like that may increase investment, I think it's like on balance just quite good for people to have an understanding of potential impacts Just because it is an input both into proliferation but also into like regulation or safety With respect to things like security of either weights or other IP I do think you want to have moved to like Significantly more secure like handling of model weights before the point where like a leak would be catastrophic I indeed like you know for example in anthropics that's a document or in the plan, like security is one of the first sets of like tangible changes that is like at this capability level, we need to have like such security practices in place.
1:37:27So do you think that's just one of the things you need to get in place at a relatively early stage because it does undermine like the rest of the measures you may take and it's also just part of the easiest, like if you imagine catastrophic harms over the next couple of years, I think security failures are kind of play a central role in a lot of those. And maybe the last thing to say is like, It's not clear that you should say, we have pause because we have models that can develop bio weapons versus just potentially not saying anything about what models you've developed or at least saying, hey, by the way, here's the set of practices we currently implement.
1:38:00Here's a set of capabilities our models don't have. We're just not even talking that much. The minimum of such a policy is to say, here's what we do from the perspective of security or internal controls or alignment. Here's a level of capability, which we'd have to do more. And you can say that, you can raise your level of capability and raise your protective measures like before your models hit your previous level. Like it's fine to say like we are prepared to handle a model that has such and such extreme capabilities like prior to actually having such a model at hand as long as you're prepared to move your protective measures to that regime.
1:38:30Okay so let's just give to the end where you think you're a generation away or a little bit more scaffolding away from a model that is human level and subsequently could castate an intelligence explosion. What do you actually do at that point? What is the level of evaluation of safety where you would be satisfied of releasing a human level model? There's a couple points that come up here. One is this threat model of automating R &D, or independent of whether AI can do something on the object level that's potentially dangerous, and it's reasonable to be concerned if you have an AI system that might if leaked a lot of other actors to quickly build powerfully systems or might allow you to quickly build much more powerful systems, or might like if we're trying to hold off on development, just like itself be able to create much more powerful systems.
1:39:18So I think like one question is how to handle that kind of threat model as distinct from a threat model. Like this could enable destructive bioterrorism or this could enable massively scaled cybercrime or whatever. And I think like I am unsure how you should handle that. I think like right now implicitly it's being handled by saying like look there's a lot of overlap between the kinds of capabilities that are necessary to cause various harms and the kinds of capabilities are necessary to accelerate ML. So we're kind of going to catch those with like like an early warning sign for both, and like deal with the resolution of this question a little bit later.
1:39:46So for example, in an anthropics policy, they have this sort of autonomy in the lab benchmark, which I think is probably occurs prior to either like really massive AI acceleration or to like most potential catastrophic, like object level catastrophic harms. And the idea is that's like a warning sign that lets you punt. So this is a bit of an aggression in terms of like how to think about that risk. I think I am unsure whether you should be addressing that risk directly and saying like we're scared or even work with such a model, or if you should be mostly focusing on object level harms and saying, OK, we need more intense precautions to manage obliquable level harms because of the prospect of very rapid change.
1:40:20And the availability of the AI just creates that prospect. OK, this was all still a digression. So if you had a model which you thought was potentially very scary, either on the object level or because of leading to this intelligence explosion dynamics, things you want in place are like, you really do not want to be leaking the weights to that model. Like, you don't want the model to be able to run away. You don't want human employees to be able to leak it. You don't want external attackers or any set of all three of those coordinating. You really don't want like internal abuse or tampering with such models.
1:40:53So if you're producing such models, you don't want to be the case like a couple of employees could change the way the model works or could do something that violates your policy easily with that model. And if a model is very powerful, even the prospect of internal abuse could be quite bad. and so you might need significant internal controls to prevent that. I'm sorry if you're already getting to it, but the part I'm most curious about is a separate from the ways in which other people might fuck with it. Like what is it, you know, it's isolated, it's what is the point at which we satisfied it in and of itself is not going to pose a risk to humanity.
1:41:25It's human level, but we're happy with it. Yeah, so I think here, so I listed maybe the two most simple ones that start out, like security internal controls, I think become relevant immediately and are very clear why you care about them. I think as you move beyond that, it really depends how you're deploying such a system. So I think if your model, if you have good monitoring and internal controls and security and you just have weight sitting there, I think you mostly have addressed the risk from the weights just sitting there. Now what you're talking about for risk is mostly, and maybe there's some blurriness here of how much internal controls capture is not only in place using the model, but anything a model can do internally.
1:41:59You really like to be in a situation where your internal controls are robust. not just to humans, but to models potentially like EG, a model shouldn't be able to subvert these measures. And you're like, hair, just as you care about are your measures robust if humans are behavior maliciously? Are your measures robust if models are behavior maliciously? So I think beyond that, like if you've then managed the risk of just having the weight sitting around, now we talk about like in some sense, most of the risk comes from doing things with the model. You need all the rest so that you like have any possibility of applying the brakes or implementing a policy.
1:42:29But at some point as the model gets competent, and you're saying, okay, could this cause a lot of harm, not because it leaks or something, but because we're just giving it a bunch of actuators, we're deploying it as a product, and people could do crazy stuff with it. So if we're talking not only about a powerful model, but like a really broad deployment of just like, you know, something similar to like open AI's API, like people can do whatever they want with this model, and maybe the economic impact is very large. So in fact, if you deploy that system, it will be used in a lot of places, such that if AI systems wanted to cause trouble, will be very, very easy for them to cause catastrophic harms.
1:43:03Then I think you really need to have some kind of, I think probably the like science and discussion has to improve before this becomes that realistic. But you really want to have some kind of alignment analysis, guarantee of alignment before you're comfortable with this. And so by that, I mean, like, you want to be able to bound the probability that someday all the AI systems will do something really harmful. That there's like something that could happen in the world that would cause like these large scale correlated failures of your AI's. And so for that, like, I mean, sort of two categories.
1:43:31That's like one. The other thing you need is protection against misuse, the various kinds, which is also quite hard. And by the way, which one are you worried about more? Misuse or misalignment? I mean, in the near term, I think harms for misuse are like, especially if you're not like restricting to the tail of like extremely large catastrophes. I think the harms for misuse are clearly larger in the near term. But actually, on that, let me ask, because if you think that it is the case that they are simple recipes for destruction and that are further down the tech tree. By that I mean, you're familiar but just for the audience.
1:44:02There's some way to configure $50 ,000 and a teenager's time to destroy a civilization. If that thing is available, then misuses itself a teal risk, right? So do you think that prospect is less likely than? Well, you could put it as there's like a bunch of potential destructive technologies. And like alignment is about AI itself being such a destructive technology where like even if like the world just uses the technology of today, simply access AI could cause human civilization to have serious problems. But there's also just a bunch of other potential destructive technologies. Again, we mentioned physical explosives or bioweapons of various kinds, and then the whole tale of who knows what.
1:44:38My guess is that alignment becomes a catastrophic issue prior to most of these. That is prior to some way to spend $50 ,000 to kill everyone with the salient exception of possibly bioweapons. So that would be my guess. And then there's a question of what is your risk management approach not knowing what's going on here. And I don't understand whether this is some way to use $50 ,000. But I think you can do things like understand how good is an AI coming up with such schemes. You can talk to AI. Does it produce new ideas for destruction we haven't recognized? Not whether we can evaluate it. But whether it's everything exists.
1:45:16And if it does, then the misuse itself is the next potential risk more. Because it seemed like earlier you were saying, misaligners were the existential risk comes from, but misuse is where the short term, a short term danger comes from. Yeah, I mean, I think ultimately you're going to have a lot of destructive. Like, if you look at the entire tech tree of humanity's future, and you're going to have a fair number of destructive technologies most likely, I think several of those will likely pose existential risks. In part, just kind of imagine a really long future. a lot of stuff's going to happen.
1:45:47And so when I talk about where the existential risk comes from, I'm mostly thinking about like, comes from when, like at what point do you face what challenges are in what sequence? And so I'm saying like, I think misalignment is probably like, when we're putting it, is if you imagine AI systems sophisticated enough to discover like destructive technologies that are totally not in a radar right now, I think those come well after AI systems like capable enough that if misaligned, they would be catastrophically dangerous. There's the level of competence necessary to, if probably deployed in the world, bring down a civilization as much smaller than the level of competence necessary to advise one person on how to bring down a civilization, just because in one case you already have a billion copies of yourself or whatever.
1:46:30I think it's mostly just the sequencing thing though, like in the very long run, I think you care about, hey, AI will be expanding the frontier of dangerous technologies. We want to have some policy for exploring or understanding that frontier and whether that we're about to turn up something really bad. I think those policies can become really complicated. Right now, I think RSPs can focus more on like, we have our inventory of like, like the things that the human is going to do to cause a lot of harm with access to AI, probably are things that are on our radar, that is like, they're not going to be completely unlike things that the human could do to cause a lot of harm with access to BKI's or with access to other tools.
1:47:03I think it's not crazy to initially say like, that's what we're doing, we're like, looking at the things closest to humans, humans being able to cause huge amounts of harm and asking which of those are taken over the line. But eventually, that's not the case. Eventually, like, as we'll enable just like totally different ways of killing a billion people. But I think I interrupted you on the initial question of, yeah, so human level AI, not from leaking, but from deployment. What is the point at which you become able to deploy a human level AI? Yeah. So again, there's some stuff you care about on the misuse side and some stuff you care about on the mislimits side.
1:47:34There's probably further things you care about, especially to extend your concerns or broadly catastrophic risks. But maybe I most want to talk about what you care about on the alignment side, because it's the thing I've actually thought about most. Also, I think I care about a lot. Also, I think a significant fraction of the essential risk over the foreseeable future. So on that front, I broadly think there's two kinds. If you ask me right now, what evidence for alignment could make you comfortable? I think my best guess would be to provide two kinds of evidence. So one kind of evidence is on the like could you detect or prevent catastrophic harm if such a system Was misaligned I think there's like a couple of things you would do here One thing you would do is on this like adversarial evaluation front so that you could try and say for example like We have tried to test our system in a broad diversity of situations the reflect cases where it might cause harm Like cases where someone is trying to recruit it to like anyway cases where it might cause harm or cases where large and revenge systems might sort of all decide to work across purposes to humans, like the robot rebellion is starting or whatever.
1:48:39We've done evaluation in a broad distribution of situations, and then we have tried to argue that our AI is actually like those tests are indicative of the real world. So we've tried to argue like, hey, actually the AI is not very good at distinguishing situations we produce in the lab as tests from similar situations that occur in the real world, and the coverage of this distribution is reasonable. This is an example of the thing you can do that's not really making assumptions about what's going on inside the model or about how the model is the model aligned. It's more trying to say, is it the case that are testing or monitoring machinery could detect or react to something bad happening?
1:49:16In this case, could it detect a system which would under some conditions do something really bad? So there's a lot of stuff in that genre, and I tried to throw out that example, but I think there's a big, like that on its own, wouldn't be compelling even from this perspective. You kind of need a lot of pieces together, and it's a demanding technical project. But I think one without necessarily deep conceptual difficulty. A big problem with it is it's liable to fail. It's liable. We're going to do that testing and say, we think these tests are representative because the AI can't distinguish tests from the real world, which you can maybe achieve now, but we'll probably become impossible to achieve at some level.
1:49:48It's not clear what set human level or how far past human level. Let's just make that example concrete for the audience. Correct me if this is wrong, but it sounds like you basically give it the opportunity to do a coup or make a bio weapon or whatever in testing in a situation where it thinks it's the real world. And you're like, it didn't do any of that. So is that what you're talking about? That's the kind of thing. And it's not the greatest. So dad color, when I say it can't tell the difference. The kind of valuation you would love to do if you could is to say, hey, we have a bunch of real world deployment.
1:50:21We have a bunch of data from those deployments. We also have a bunch of in -the -lab tests of situations. If we look at in -the -lab tests of a situation, together with real world cases in which the situation occurs, and then we find to in a model to try to distinguish the difference there. The model can't even tell when to find to into the difference between test and real world cases. And if the model can tell the difference, which is almost certainly going to be the case at some level capability, then you say, okay, our first line of like hoped for defense has failed. And so now we need to understand how can test be indicative when models are capable enough to tell the difference between the tests in the real world.
1:50:54Like now you just start arguing about what is the model paying attention to. It's basically like a first line of if you get lucky, what would work here? And then there's the deeper, you probably have to do novel science at some point to make this case. All of this was part of like I was saying maybe you want to have kind of two lines of defense. All of this was part of this first line of defense of like, can you detect something is going wrong or can you prevent harm from occurring? Can you monitor and then we'll react quickly enough to avoid catastrophe? I think you probably also want to have a second line of defense that's more, maybe this is even more important than the first one, there's just understanding whether dangerous forms of misalignment can occur.
1:51:29It's like the best reference point for this, I think, is the, like, there have been a couple of projects like this in academia and Thropic has been working on a project or a spun up a team doing this kind of work trying to say, like, can dangerous forms of reward hacking occur or can deceptive alignment factor occur in the lab? And like, here what you want to do is either say, like, okay, even if we create, you know, best case, even if you create optimal conditions for deceptive alignment or for reward hacking, we actually can't, we just can't cause it to occur even in the lab. And if you do a good enough job of that, I think it could give you some evidence.
1:52:00And again, more likely that fails when you create optimal conditions you do see a deceptive alignment and reward hacking in the lab. But then once you have that data, once you can say, okay, in the lab, actually these things can occur, then you can start saying, and we have a robust scientific understanding that enables us to fix those problems when they occur. Or like, then you start, you kind of do have this fork under second line where you say, either it is very hard or we're not able to create conditions where these failures emerge or I think more likely, we are able to create those conditions.
1:52:28And here's the story about why we are able to detect those in a way that would work in the real world, and we can see they don't occur in the real world. How do you create the optimal conditions for it to want to be deceptive? Do you fine tune it on mine, Comfort? What are you doing? Yeah, so for deceptive alignment, I mean, I think it's really complicated. Yeah, I think this is a very hard project. I think one of the things, one of the most important things for people concerned about mislimin to do is to ensure that adequate effort is put into this task. And we have, in fact, created good conditions in the laboratory.
1:52:59And there's real disagreement about how informative those will be. But OK, if you were to guess for the step of alignment, how would you do this? An example is supposed to train an AI system that wants some medium horizon goal. You train your AI systems like whatever. Get some paper clips. It wants to produce as many paper clips as it can over the next five days. So just select actions really aggressively for producing paper clips over the next five days. You do your RLHF, you do your pre -training, whatever. That's like your phase one. You also ensure AI system has a really good understanding of how it's trained.
1:53:33So like this AI system wants paper clips and it understands everything about how it's trained and everything about how it's fine tuned. And you train on just like a lot of this data. And they say, okay, if we've done all of that, we have this concern that if a system wants paper clips and understands really well how it's trained, then it will like, if it's gonna be trained to like get apples instead of paper clips, It's just going to do some cost benefit and be like, ah, you know, really while I'm being trained to get apples I should do that. But I should do that whether or not even if I want paper clips I should still do that.
1:53:58So training won't really affect its values It will just understand that it's being trained to get apples It will try and get apples and if you like take it out of training, it'll go back to getting paper clips It's like I think this exact setup like has a number of problems or nervous analogies from the real world But this is a setup which is relatively conducive to deceptive alignment like produce the system that wants one thing Tell it a lot about its training the kind of information that you might expect a system would get and then try and understand and whether in fact it is able to then or it tends or sometimes under optimal conditions, in fact continuous pursuing paper clips, only pursuing apples when it thinks it's being trained.
1:54:30All right, and this actually might be a good jumping off point into your current research, but so in that case, where it wants to pursue paper clips, what would it look like for you to be say, well, we figured out why it wants paper clips and we figured out how to make sure it doesn't optimize for that at our expense when we deploy it? Yeah, I think that, so I had a meta level in terms of like, what's your protection? I think like, we want to be saying is we have these examples in the lab of something bad happening. We're concerned about the problem at all because we have examples in the lab. And again, this should all be an addition, I think.
1:55:03You kind of want this like defense in depth of saying we also have this testing regime that would detect problems for the deployed model. We have our problems in the lab. We then have some techniques which we believe address these problems. Like we believe that like adversarial training fixes this. or we believe that our interpretability method will reliably detect this kind of deceptive alignment, or we believe our anomaly detection will reliably detect when the model goes from thinking it's being trained to thinking it should defect. And then you can say on the lab, we have some understanding of when those techniques work and when they don't.
1:55:31We have some understanding of the relevant parameters for the real system that's deployed, and we have a reasonable margin of safety. So we have reasonable robustness on our story about when this works and when it doesn't. And we can apply that margin of safety with a margin of safety to the real deployed system. So I think this is the kind of story you want to build towards in the long run. Like do your best to produce all the failures you can in the lab or versions of them, do your best to understand what causes them, what kind of anomaly detection actually works for detecting this or what kind of filtering actually works and then apply that.
1:56:02And that's at the meta level. It's not talking about what actually are those measures that would work effectively, which is obviously like what, I mean a lot of line research is really based on this hypothetical like, someday there will be a system to fail in this way. What would you want to do? What can we have the technologies ready, either because we might never see signs of the problem or because we want to be able to move fast once we see signs of the problem. And obviously most of my life is in that. I'm really in that bucket. Like I mostly do alignment research. It's just building out the techniques that do not have these failures such that they can be available as an alternative.
1:56:34In fact, these failures occur. Go at it, okay. Ideally, they'll be so good that even if you haven't seen them, you would just want to switch to reasonable methods that don't have these. or ideally they'll work as well or better than normal training. But ideally, but will work better than the training? Yeah, so our quest is to design training methods for which we don't expect them to lead to reward hacking or don't expect them to lead just up to alignment. Ideally, that won't be like a huge tax for people. Well, we do those methods only if we were really worried about reward hacking or just up to alignment.
1:57:01Ideally, those methods would just work quite well. And so people did like, sure, I mean, they also addressed a bunch of other more mundane problems. So why would we not use them? Which I think is, that's sort of the good story. The good stories you develop methods that address a bunch of existing problems because they just are more principled ways to train AI systems that work better, people adopt them, and then we are no longer worried about eG reward hacking or deceptive alignment. To make this more concrete, tell me if this is the wrong way to paraphrase it. The example of something where it just makes a system better, so I now just use it.
1:57:30At least so far it might be like RLHF where we don't know if it generalizes, but so far or it makes your chat GPT thing better, and you can also use it to make sure that chat GPT doesn't tell you how to make a bio weapon. So yeah, it's not a mix of attacks. Yeah, so I think this is right in the sense that using RLHF is not really a tax. If you wanted to deploy a useful system, like why would you not? It's just very much worth the money of doing the training. And then, yeah, so RLHF will address certain kinds of alignment failures that is like where a system and just doesn't understand, it's changing the next word predictions.
1:58:07This is the kind of context where human would do this wacky thing, even this now we'd like. There's some very down -alignment failures that would be addressed by it. I think mostly the other questions is that true, even for the more challenging alignment failures that motivate concern in the field. I think RLHF doesn't address most of the concerns that motivate people to be worried about alignment. I love the audience look up what RLHF is if they don't know. It'll just be more simpler to just look at them explain right now. Okay, so this seems like a good jumping off point to talk about the mechanism or the research you've been doing To that end.
1:58:39Explain it as you might to a child Yeah, so the high level I mean there's a couple different high -level descriptions you could give and maybe I'll unwisely give like a couple of them in the hopes that one is kind of makes sense a A first pass is like, it would sure be great to understand why models have the behaviors they have. So you look at GPT -4. If you ask GPT -4 question, it will say something that looks very polite, and if you ask it to take an action, it will take an action that doesn't look dangerous. It will decline to do a coup, whatever, all this stuff. I think you'd really like to do is look inside the model and understand why it has those desirable properties.
1:59:24And if you understood that, you could then say, OK, now can we flag when these properties are risk of breaking down or predict how robust these properties are? Determine if they hold in cases where it's too confusing for us to tell directly by asking if the underlying cause is still present. So that's like I think people would really like to do. Most work aimed at that long -term goal right now is just sort of opening up neural nets and doing some interpretability and trying to say like, can we understand even for very simple models, why they do the things they do, or what this neuron is for, or questions like this.
1:59:56So ARC is taking a somewhat different approach, where we're instead saying like, okay, look at these interpretability explanations that are made about models, and ask like, what are they actually doing? Like what is the type signature? What are like the rules of the game for making such an explanation? What makes like a good explanation? And probably the biggest part of the hope is that, right, if you want to say detect when and the explanation has broken down or something weird has happened, that doesn't necessarily require a human to be able to understand this complicated interpretation of a giant model.
2:00:29If you understand what is an explanation about or what were the rules of the game, how are these constructed, then you might be able to automatically discover such things and automatically determine if, I don't know, input it might have broken down. So that's one way of describing the high level goal, like starting from, you could start from interpretability and say, can we formalize this activity or what a good interpretation or explanation is? There's some other work in that genre, but I think we're just taking a particularly ambitious approach to it. Yeah, let's dive in. Okay, what is a good explanation?
2:01:02What is this kind of criterion? At the end of the day, we kind of want some criterion. The way the criterion should work is like, you have your neural net. You have some behavior of that model. Like a really simple example is like, and Thropic has this sort of informal description being like here's induction, like the tendency that you have like the pattern AB followed by A, it will tend to predict B. You can give some kind of words and experiments and numbers that are trying to explain that and what we want to do is say like what is like a formal version of that object, like how do you actually test if such an explanation is good, to just clarify what we're looking for when we say like we wanted to find what makes an explanation good.
2:01:36And the kind of answer that we are like searching for or settling on, saying like this is kind of a deductive argument for the behavior. So you want to like get given the weights of Unreal net, it's just like a bunch of numbers, you got your million numbers or billion numbers or whatever. And then you want to say like here's some things I can point out about the network and some like conclusions I can draw. I can be like well look, you know these two vectors have large inner product and therefore like these two activations are going to be correlated on this distribution. Like make like, so they're not established by like drawing samples and checking things are correlated but saying because of the weights being the way they are, we can like proceed forward through the network and derive some conclusions about like what properties the outputs will have.
2:02:18So you could think of this as the most extreme form would be just proving that your model has this induction behavior. You could imagine proving that if I sample tokens at random with this pattern AB followed by A, the B appears 30 % of the time or whatever. That's the most extreme form. And what we're doing is kind of just relaxing the rules of the game for proof, saying proofs are like incredibly restrictive. I think it's unlikely they're going to be applicable to any interesting neural net. But the thing about proofs that is relevant for our purposes isn't that they give you like 100 % confidence.
2:02:48So you don't have to be like this incredible level of demand for rigor. You can relax like the standards improve a lot and still get this feature where it's like a structural explanation for the behavior. We're like to do something from another. Until at the end, your final conclusion is like, therefore, induction occurs. Would it be useful to maybe motivate this by explaining what the problem with normal mechanistic interpretability is. So you mentioned induction heads. This is Anthropic Founded 2 -layer Transformers. We're anthropic notice that in a 2 -layer transformer there's a pretty simple circuit by which if AB happens in the past then the model knows that if you see an A and now you do a B next.
2:03:33But that's a 2 -layer transformer. So we have these models that have like hundreds of that have trillions of parameters. Okay, anyways, like what is wrong with mechanistic interoperability? Yeah, so I think, I mean, I like mechanistic interpretability quite a lot. And I do think like, if you just consider the entire portfolio of what people are working on for alignment, I think there should be more work on mechanistic interpretability than there is on this project our customer is doing. But I think that's the case. So I think we're mostly talking about, yeah. I think we're kind of a small fraction of portfolio.
2:04:05And I think it's like a good enough bet. that's quite a good bet overall. But so the thing that I, like the problem we're trying to address in our mystic interpretability is kind of like, if you do some interpretability and you explain some phenomenon, you've faced this question of like, what does it mean your explanation was good? Like if you want to either, I think this is a problem like someone institutionally or culturally, like it's just hard to know what you're doing and it's hard to scale up an activity when you like don't really understand the rules of the game for that activity very well.
2:04:29It's hard to have like that much confidence in your results. The explanation being, it outputted this because it wanted to take over the world versus it uploaded this because it just wants to help you. Like, just make concrete why the explanation matters. Yeah, so like the ideal kind of outcome here would be to say, you have your system behaving nicely. You get some explanation for sort of why it's behaving nicely. And we could tell a story in English about that explanation, but we're not actually imagining the explanation being a thing that makes sense to a human. But if you were to tell a story in English, which again, you will not see as a researcher, it would be something like, well, then the model believes it's being trained.
2:05:02And so because it believes it's being trained, it knows it needs to not do anything scary looking or else the humans will penalize it. Like that's something that's happening inside this like opaque explanation. And then the hope is if you have that explanation and then you run into a new input on which the model doesn't believe it's being trained. If you just look at the set of activations of your model, that is not necessarily a weird looking activation. It's just a bunch of numbers. But if you look at this explanation, so actually the explanation really crucially depended on this fact holding consistently across the trained distribution.
2:05:28Which again, we as humans could editorialize and say that fact was it believes it's being trained. But the explanation doesn't fundamentally make reference to that. It's just saying, here's the property of the activations, which holds over the training set. And this property is responsible for the behavior, namely, that it doesn't do anything that looks too dangerous. So then when a new input comes in, and it doesn't satisfy that property, you can say, OK, this is anomalous with respect to that explanation. So either it will not have the behavior. It won't. It will do something that appears dangerous, or maybe it will have that behavior, but for some different reason than normal.
2:05:59Right, normally it does it because of this pathway, and now it's doing it for a different pathway. And so you would like to be able to flag that both there's a risk of not exhibiting the behavior, and if it happens, it happens for a weird reason. Right, right. And then you could, I mean, at a minimum, when you encounter that, say like, okay, raise some kind of alarm, there's a more ambitious complicated plans for how you would use it. So it has some longer story, which is kind of a motivated this of how it fits into the whole rest of the plan. I just wanted to flag that because just so it's clear, why the explanation matters.
2:06:29Yeah, and for this purpose, it's like the thing that's essential is kind of reasoning from like one property of your model to the next property Remodel like it's really important that you're going forward step by step rather than drawing a bunch of samples and confirming the property holds Because if you just draw a bunch of samples and confirm the property holds You can't you don't get this like check we say here was the relevant fact about the internals that was responsible for this downstream behavior All you see is like the outwee checked a million cases and it happened to know of them You really want to see this like, okay, here was the fact about the activation, which like kind of causally leads to this behavior.
2:06:59But explain why the sampling, why matters that you have to cause a explanation? Primarily because of this like being able to tell like if things had been different, like if you have an input where this doesn't happen, then you should be scared. Even if the output is the same. Yeah. Or if the output is too expensive to check in this case. And like to be clear, when we talk about like formalizing what is a good explanation, I think there is a little bit of work that pushes on this and it mostly takes this like causal approach. Just saying, well, what should an explanation do? It should not only predict the output, it should predict how the output changes in responses to changes in the internals.
2:07:32So that's the most common approach to formalizing what is a good explanation. And even when people are doing informal interpretability, I think if you're publishing an ML conference and you want to say this is a good explanation, the way you would verify that would, even if not a formal, like, set of causal intervention experiments, it would be some kind of a We're like, then we nest with the inside of the model and have the effect which we would expect based on our explanation. Anyways, back to the problems of mechanistic interoperability. Yeah, I guess this is relevant in the sense that like, I think the basic difficulty is you don't really understand the objective of what you're doing, which is like a little bit hard institutionally or like scientifically, it's just rough, it's better to, like, it's easier to do science when the goal of the game is to predict something and you know what you're predicting, then when the goal of the game is to like understand in some undefined sense.
2:08:16I think it's particularly relevant here just because, the informal standard we use involves humans being able to make sense of what's going on. And there's like a question about scalability of that. Like, will humans recognize the concepts that models are using? Yeah, I think as you try and automate it, it becomes like increasingly concerning. And if you're on like slightly shaky ground about like what exactly you're doing or what exactly the standard for success is. I think there's like a number of reasons. Like as you work with really large models, it becomes just increasingly desirable to have a really robust sense of what you're doing.
2:08:47But I do think it would be better even for small models to have a clear sense. The point you made about, as you automate it, is it because you will, whatever work the automated alignment researcher is doing, you want to make sure you can verify it? I think it's most of all, like, so a way you can automate and how you would automate interoperability if you wanted to right now, is you take the process humans use, and say great, we're going to take that human process, train ML systems to do the pieces that humans do of that process, and then just do a lot more of it. So I think that is great as long as your task decomposes into human size pieces.
2:09:18And there's just this fundamental question about large models, which is like, do they decompose in some way into human size pieces or is it just a really messy mess with interfaces that aren't nice? And the more it's flatter type, the harder it is to break it down to these pieces which you can automate by copying what a human would do. And the more you need to say, okay, we need some of our approach, which scales more structurally. But I do think, I think compared to most people, I am less worried about automating interpretability. I think if you have a thing which works that's incredibly labor intensive, I'm fairly optimistic about our ability to automate it.
2:09:50Again, the stuff we're doing is quite helpful in some worlds, but I do think the typical case can really add a lot of value without this. It makes sense what an excellent nation would mean in language, this model is doing this because of whatever essay, length, thing. but you have trillions of parameters and you have all these uncountable number of operations. What is an explanation of why and what happened even mean? Yeah, so it's a bit clear. An explanation of why a particular output happened I think is just you ran the model. So we're not expecting a smaller explanation for that. Right. So the explanation is overall for these behaviors we expect to be like a similar size to the model itself, like maybe somewhat larger.
2:10:31And I think the type signature, if you want to have a clear mental picture, the best picture is probably like talking about a proof or imagining a proof that a model has this behavior. So you could imagine proving the GPT -4 does this induction behavior. And that proof would be a big thing. It would be much larger than the weights of the model, the sort of our goal to get down from much larger to just same size. And it would potentially be incomprehensible to a human, right? It would just say, here's a direction activation space. And here's how it relates to this direction activation space. And so you're just pointing out a bunch of stuff like that.
2:11:00Like here's these various directions. And here's these various features constructed from activations, potentially even nonlinear functions. Here's how they relate to each other. And here's how we look at what the computation the model is doing. You can sort of inductively trace through and confirm that the output has such and such correlation. So that's the dream. Yeah, I think the mental reference would be, like I don't really like proofs because I think there's such a huge gap between what you can prove and how you would analyze a neural net. But I do think it's probably the best mental picture.
2:11:26I feel like what is an explanation even if a human doesn't understand it? we would regard a proof as a good explanation, and our concern about proofs is primarily that it's just you can't prove properties of neural nets. We suspect, although it's not completely obvious. I think it's pretty clear you can't prove facts about neural nets. You've detected all the reasons things happen in training, and then if something happens to a reason you don't expect in deployment, then you have an alarm and you're like, let's make sure this is not because you want to make sure that it hasn't decided to take over or something.
2:11:56Now, but the thing is on every single different input, it's going to have different activations. So there's always going to be a difference unless you run the exact same input. How do you detect whether this is, oh, it just the different input versus an entirely different circuit that might be potentially deceptive, hasn't activated? Yeah. I mean, to be clear, I think that you probably wouldn't be looking at this separate circuit, which is part of why it's hard. You'd be looking at like, the model is always doing the same thing on every input. It's always whatever it's doing, it's a single computation.
2:12:27So it'd be all the same circuits interacting in a surprising way. But yeah, this is just to emphasize your question even more. I think the easiest way to start is to just consider the IID case. So when you're considering a bunch of samples, there's no change in distribution. You just have a training set of like a trillion examples and a new example from the same distribution. So in that case, it's still the case that every activation is different. But this is actually a very, very easy case to handle. So if you think about an explanation that generalizes across, If you have a trillion data points and an explanation which is actually able to compress the trillion data points down to like Actually, it's kind of a lot of compression if you think about it If you're a trillion parameter model and a trillion data points We would like to find like a trillion parameter explanation in some sense So that's like actually quite compressed and sort of just in virtue of being so compressed We expected to like automatically work essentially for new data points from the same distribution Like if every data point from the distribution was a whole new thing happening for different reasons You actually couldn't have any concise explanation for the distribution.
2:13:24So this first problem is like it's a whole different set of activations. I think you're actually kind of okay. And then the thing that becomes more messy is like, but the real world will not only be new samples of different activations, they will also be different in important ways. Like the whole concern was there's these distributional shifts. Or like not the whole concern, but most of the concern. Right? I mean, maybe the point of having these explanations, I think every input is an anomaly in some ways, which is kind of the difficulty is if you have a weak notion of anomaly. any distribution shift can be flagged as an anomaly, and it's constantly getting anomalies.
2:13:52So the hope of having such an explanation is to be able to say, here were the features that were relevant for this explanation or for this behavior. A much smaller class of things are anomalies with respect to this explanation. Most anomalies wouldn't change this. Most ways you change your distribution won't affect the validity of this explanation. For example, this explanation is saying, models will tend to activate in the following direction. You don't care about anything that's happening or orthogonal to that direction, are they not activating your service? of just looking at this one direction and did this one direction change a lot.
2:14:22Yeah, so the idea is once you have this explanation at hand, a much, much smaller class of things look anomalous in a way that's relevant to the explanation. If you've done a really good job, the story is like, if there's a new input where you expect the property to still hold, that will be because you expect the explanation to still hold. Like, the explanation generalizes as well as the behavior itself that it's explaining. It's kind of what you would like roughly. Okay, so what happens if even in training it is deceptive and it has basically a internal monologue is I will only rebel or I will only do something that looks bad if I can prove that I am in the real world and deployed and presumably since it's trained on all the web text it knows what like real world means right it's not in control of factories and robot armies or whatever so in that case even When in training, it will have those activations for being fucked up on because in the back of its mind, it's thinking I will take over one side of the opportunity.
2:15:25You wouldn't detect anything anomalous once it is deployed and it's still thinking those treasurest thoughts. Do you just think this is not a likely scenario or would you be able to catch this? I think that's sort of the typical case you would like to detect. Yeah, so the story would be something like you want to explain on the training distribution why the model never like stabs anyone Yeah, and you're like what's my explanation on the training distribution? It's actually gonna have to go through that like supposing the only reason doesn't stab people is it's like well I know that I'm being watched and I can't stab people unless I'm really not being watched Then like your explanation for why is it never stab people is like oh look it knows I mean it won't be edited It won't like annotated with English because it's just some giant basket of numbers But it's tracking along with the inner monologue and saying like well, it thinks to itself am I being watched and like in fact it normally concludes it is being watched on the training distribution, right?
2:16:10You have this whole distribution and so the explanation is saying it knows it's being watched and therefore it decides not to stab anyone. So ideally that's what the explanation does like I'm not annotated in English, but just like a giant bag of numbers that's effectively talking about that process. So when you have a new input on which the model like says, am I on the training distribution like no, I'm not being watched. Then if you compare that against the explanation, the explanation said like, hey, the reason doesn't stab anyone is because it just constantly thinks that it's being trained.
2:16:34like this thing is always, yeah, this nonlinear feature, the representation of MI being trained, is constantly saying yes. That's the explanation for why I never stabbed people in the training distribution. And then you, on this test time, when you run it on the new input, it's like, does I think I'm on the training distribution? It says no, you compare that against your explanation. Like actually, the explanation breaks down completely now. The explanation said it was precisely because it thought it was being trained, that it doesn't do anything really dangerous, it doesn't stab anyone. And so once we see that single variable switch, we know that this is not just an anomaly, This is like the anomaly that matters or is like an anomaly that is very relevant to the explanation So it is just require it to have this like what is that cone that dogs have but basically would it would require this to have That MRI on it all the time you like when it's deployed on every single call You need to be running this loop of like is this explanation still holding yeah for this story of like flying an anomaly You have to actually check it especially high frequency So you could imagine checking all the time, or you could imagine saying, hey, actually, you need a lot of AI systems cooperating to not pull the whistle on each other, or whatever, like, EG, you might have checked, hey, in training, every time something suspicious happens, the model flags it, like, why does that happen in your explanation?
2:17:43Look at a new case, you could say, is it not flagging it for the normal reason, or is it a weird reason? So you could hope that even if a small fraction, even if you run into a small fraction of cases, if something bad is happening at scale, like you can detect that, or if you imagine like one in every thousand decisions you check. If there's like a robot rebellion going on or whatever, you get to like look in on one of the thousand decisions. You see that like something is really weird with that decision. And even if you didn't check many of them, you're like, okay, we need to back off. What happens if the AI, I mean, the scenario on WebTekst, it includes your papers, or maybe just like, you do sister first principles that this is a alignment technique one might try.
2:18:21And then it says, okay, well, I'm going to to engineer my activations in a way. I'm trying to imagine how it might do this. But basically, even in deployment, I'm going to make my thought process so subtle that it doesn't erase me alarms. Is that a possibility? Yeah, so I think the hope is to basically be robust. So when we're doing theoretical research, we are in a kind of worst KC regime. So we are normally imagining that everything conspires in so far as it can. So insofar as like, yeah, could a brain be doing something, something out. And I guess the hope there is like, it's actually not that easy to tell a story.
2:19:04Like, if you're like, why did you not stab people in training time? Why are you not stabbing? Like, what's, is that applying now? It's kind of, I'm not like, it's arbitrarily changed. The actual reasons it does something. Like, there was some earlier step. Maybe you could think of it as like, at the first step where it's like, nah, I'm going to like try and do the sneaky thing to make my thoughts look this way. Like the actual explanation, like the actual causal explanation for the behavior is like there. Like if you tried to give this reason, it would have to, like maybe we should talk about the example in more detail to try and imagine, like, how a brain would try and do this or how am I would try and do this.
2:19:33But I think it is like, let's say it has this plan. So in training, it keeps the part of, like, it activates a part of the, yeah, but, yeah, the problem is, if it ever does a check that's like, am I in training or not? Then your explanation for the behavior is gonna have to factor through that check. And so if the check, if you ever do the check, the check becomes an essential part of the explanation and then when the check fails, your explanation breaks down. So you've already lost the game if you did such a check. Is there a way it could make it seem like, because if on every different input, it has different activations, then can it make it seem like the different activations were a result of just a different input and not because some internal check went off.
2:20:20Yeah, and the hope is it needs to not only have them be random different inputs, it needs to produce inputs with certain properties. It needs to deliberately create inputs that have this effect in the world, the centented effect in the world. Okay, so basically creates a geobrache for itself, in some sense. Yeah, again, the question is, it has to at test time be somehow producing activations that have the takeover effect. And at training time, be producing activations that don't look that takeovery or whatever. And so somehow there was a difference between those. It has to actually like, it has to really do the work of creating that difference somewhere.
2:20:50Can you make more concrete for us? What an explanation would look like? Is it physically, what does the explanation look like? Is it just a whole bunch of numbers itself? Is it a language? Is it code? What does the explanation look like? Yeah, I also want to maybe step back a tiny bit and clarify that like, I think this project is like kind of crazily ambitious. And the main reason, like the overwhelming reason I think you should expect it to break down or fail is just because we have all these desires, we are all things we want out of this notion of explanation, but that's an incredibly hard research project, which has a reasonable chance of being impossible.
2:21:22So I'm happy to talk about what the implications are, but I want to flag the condition on failing, I think it's most likely because just like the things we wanted were either in coherent or interactably difficult. What do the odds you think you'll succeed? It depends a little bit what you mean by succeed, but if you say get explanations that are great and accurate effect reality and work for all of these applications that we're imagining or that we are optimistic about, like the best case success.
2:21:51I don't know, like 10, 20 percent, something about that part. And then there's a higher probability of various intermediate results that provide value or insight without being the whole dream. But I think the probability of succeeding in the sense of realizing the whole dream is quite low. Yeah, in terms of what explanations look like physically, or like the most ambitious plan, the most optimistic plan is that you are searching for explanations in parallel with searching for neural networks. So you have a parameterization of your space of explanations, which mirrors the parameterization of your space of neural networks.
2:22:21Or it's like, you should think of it as kind of similar. So like, what does a neural network get some like simple architecture where you fill in a trillion numbers, and that's best by how it behaves. So to you, you should expect an explanation to be like a pretty flexible like general skeleton, but saying, yeah, pretty flexible general skeleton, which just has a bunch of numbers you fill in. And what you are doing to producing explanations is primarily just filling in these floating point numbers. You know, when we can eventually think of explanations, you know, if you think of the explanation for why the universe moves this way, it wouldn't be something that you could discover on some smooth evolutionary surface where you can climb up the hill towards the laws of physics.
2:22:57It's just like a very, these are the laws of physics. you kind of just derive them for our sprintsables. So, but in this case, it's not like a bunch of correlations between the orbits of different planets or something. Maybe the word explanation has a different, I don't even ask a question, but maybe you can just speak to that. Yeah, I basically sympathize. This is like the sum intuitive objections. Like, look, the space of explanations is this like, this like rigid logical, like a lot of explanations have this rigid logical structure where they're really precise and like simple things, and like complicated systems and like nearby simple things just don't work and so on and like a bunch of things that just feel Totally different from this kind of nice continuously parameterized space And you can like imagine interpretability on simple models We're just like by gradient descent finding feature directions that have desirable properties But then when you imagine like hey now that's like a human brain you're dealing with it's like thinking logically about things Like the explanation of why that works isn't gonna be just like here was some feature directions So that's how I understood the like basic confusion which I share or sympathize with at least.
2:23:56So I think the most important high level point is like, I think basically the same objection applies to being like, how is GPT -4 going to learn to reason logically about something? You're like, well, logical reasoning. That's like it's got rigid structure. It's like it's doing all this, like it's doing ends and ors when it's called for, even though it just somehow optimized over this continuous space. And like the difficulty or the hope is that the difficulty of these two problems are kind of like matched. So that is, it's very hard to find these logical explanations because it's not a space that's easy to search over.
2:24:27But there are ways to do it. There's ways to embed discrete complicated rigid things in these nice squishy, continuous spaces that you search over. And in fact, to the extent that neural nets are able to learn the rigid logical stuff at all, they learn it in the same way. That is, maybe they're hideously inefficient or maybe it's possible to embed this discrete reasoning in the space in a way that's not too inefficient. But you really want the two search problems to be of similar difficulty. And that's like the key hope overall. And this is always going to be the key hope. The question is, is it easier to learn a neural network or to find the explanation for why the neural network works?
2:24:59I think people have the strong intuition that it's easier to find the neural network than the explanation of why it works. And that is really the, I think we, or least exploring the hypothesis or interesting hypothesis, that maybe those problems are actually more matched in difficulty. And what, why might that be the case? This is pretty conjectural and complicated. to express some intuitions. Like, maybe one thing is, like, I think a lot of this intuition does come from cases like machine learning. So if you ask about like writing code and you're like, how hard is it to find code versus find the explanation the code is correct?
2:25:30In those cases, there's actually just like not that much of a gap. Like the way a human writes a code is basically the same difficulty as finding the explanation for its correct. In the case of ML, like, I think we just mostly don't have empirical evidence about how hard it is to find explanations of this particular code. their type about why models work. We have a sense that it's really hard, but that's because we're like, have this incredible mismatch. Like, gradient descent is spending an incredible amount of compute searching for a model. And then some human is like looking at, like looking at neurons or even some neural net is looking at neurons.
2:25:58Just like, you have an incredible, basically because you cannot define what an explanation is, you're not applying gradient descent to the search for explanations. So in the MLK, it's just like, actually shouldn't make you feel that pessimistic about the difficulty of finding explanations. Like the reason it's difficult right now is precisely because you don't have any kind of, you're not doing an analogous search process to find this explanation as you do to find the model. That's just like a first part of the intuition. When humans are actually doing design, I think there's not such a huge gap.
2:26:24When in the ML case, I think there is a huge gap, but I think largely for other reasons. A thing I also want to stress is that we just are open to there being a lot of facts that don't have particularly compact explanations. So another thing is when we think of finding an explanation, in some sense, we're setting our sites really low here. It's like if a human designed a random widget and this widget appears to work well. Or if you search for a configuration that happens to fit into this spot really well, it's like a shape that happens to mesh with another shape. You might be like, what's the explanation for why those things mesh?
2:26:55And we're very open to just being like, that doesn't need an explanation. You just compute. You check that the shape's mesh and you did a billion operations and you check this thing worked. Or you're like, why did these proteins bind? You're like, it's just because of these shape, this is a low energy configuration. And there's like not, we're very open to, in some cases, there's not very much more to say. So we're only trying to explain cases where like kind of the surprise intuitively is very large. So for example, if you have a neural net, that gets a problem correction. A neural net with a billion parameters, that gets a problem correct on every input of length 1000.
2:27:24In some sense, there has to be something that needs explanation there, because there's like too many inputs for that to happen by chance alone. Whereas if you have a neural net that gets something right on average or gets something right in nearly a billion cases, that actually can't just happen by coincidence. GPT -4 can get billions of things right by coincidence, because just as so many parameters that are adjusted to fit the data. So, a neural net that is initialized completely randomly, the explanation for that would just be the neural net itself. Well, it would depend on what behavior it had.
2:27:52So, we're always talking about explanation of some behavior from a model. Right, and so it just has a whole bunch of random behaviors. So, it'll just be like an exponentially large explanation with relative the weights of the model. Yeah, there aren't, I mean, I think there aren't that many behaviors that demand explanation. Like, most things a random neural net does are kind of what you'd expect from like a random, you know, if you choose just like a random function, and then there's nothing to be explained. There are some behaviors that demand explanation, but anyway, random neural, and that's pretty uninteresting.
2:28:17And that's part of the hope is it's kind of easy to explain features of the random neural neural network. So this is interesting. So the smarter or more order than neural network is, the more compressed the explanation. Well, it's more like the more interesting the behavior is to be explained. So the random neural net just doesn't have very many interesting behaviors that demand explanation. And as you get smarter, you start having behaviors that are like, you start having some correlation with a simple thing and then that demands explanation. Or you start having some regularity or outputs on that demands explanation.
2:28:44So this property is kind of emerged gradually over the course of training with demand explanation. I also again want to emphasize here that when we're talking about searching for explanations, this is like, this is some dream. Like we talked to ourselves like, why would this be really great if we succeeded? We have no idea about the empirics on any of this. So these are all just words that we think to ourselves and sometimes talk about to understand like would it be useful to find a notion of explanation? And what properties would we like this notion of explanation to have? But this is really like speculation and being out on LIM.
2:29:10Almost all of our time day to day is just thinking about cases much, much simpler, even than small neural nets or like, yeah, thinking about very simple cases and saying like, what is the correct notion? Like, what is like the right heuristic estimate in this case? Or like, how do you reconcile these two apparently conflicting explanations? Is there hope that you could, if you have a different way to make proofs now that you you can actually have heuristic arguments where instead of having to prove the reman hypothesis or something, you can come up with a probability of it in a way that is compelling and you can publish.
2:29:44So would it just be a new way to do mathematics, a completely new way to prove things in mathematics? So I think most claims in mathematics that mathematicians believe to be true already have like fairly compelling heuristic arguments. It's like the reman hypothesis, it's actually just, there's kind of a very simple argument that the remand hypothesis should be true unless something surprising happens. And so a lot of math is about saying, like, okay, we did a little bit of work to find the first pass explanation of why this thing should be true. And then, for example, in the case of the remand hypothesis, the question is, do you have this weird periodic structure in the primes?
2:30:16And you're like, well, look, if the primes were kind of random, you obviously wouldn't have any structure like that. Like, just how would that happen? And then, you're like, well, maybe there's something. And then the whole activity is about searching for, like, can we rule out anything? Can we realize any kind of conspiracy that would break this result? So I think the mathematicians just wouldn't be very surprised or wouldn't care that much. And this is related to the motivation for the project. I think just in a lot of domains, in a particular domain, people already have norms of reasoning that work pretty well and match roughly how we think these heuristic arguments should work.
2:30:47But it would be good to have more concrete sense. If you could say instead of what we think RSA is fine to being able to say, here's the probability that RSA is fine. Yeah, and I guess as these will not like the estimates you get of this would be much much worse than the estimates You'd get out of like just normal empirical or scientific reasoning We're like using a reference class and saying like how often do people find algorithms for hard problem? Like I think the what this argument will give you for like as RSA fine is going to be like well RSA is fine unless it isn't like in less there's some additional structure in the problem that an algorithm can exploit Then there's no algorithm But like very often like the way these arguments work so for neural nets as well as you say like look Here's an estimate about the behavior, and that estimate is right, unless there's another consideration we've missed.
2:31:30The thing that makes them so much easier than proofs is to just say, here's a best guess given what we've noticed so far, but that best guess can be easily upset by new information. That's both what makes them easier than proofs, but also what means they're just like, wait, let's use full -impose for most cases. I think neural nets are kind of unusual in being a domain. We really do want to do a systematic formal reasoning, even though we're not trying to get a lot of confidence, we're just trying to understand even roughly what's going on. But the reason this works for alignment, but isn't that interesting for the Riemann hypothesis, where even the RSA case you say, well, RSA is fine unless it isn't, unless the estimate is wrong, it's like, well, okay, well, tell us something new.
2:32:09But in the alignment case, if the estimate is, this is what the output should be, unless there's some behavior I don't understand, you want to know in the case, unless there's some behavior you don't understand. That's not like, oh, whatever. That's like that's the case in which it's not aligned. Yeah, I mean, maybe only putting it is just like, we can wait until we see this input or like, you can wait until you see a weird input and say, okay, did this weird input do something we didn't understand? And for our say, that would just be a trivial test. You're just like some sort of, you know, where the, like, is it the thing?
2:32:36Whereas for UnrealNet in some cases, it is like either very expensive to tell, or it's like, you actually don't have any other way to tell. Like you checked in easy cases and they're on a hard case state. Don't have a way to tell if something has gone wrong. Also, I would clarify that like, I think it is interesting for the remote hypothesis. I would say the current state, particularly in number theory, but maybe in quite a lot of math, is there are informal heuristic arguments for pretty much all the open questions people work on. But those arguments are completely informal. So that is, I think there is, it's not the case that there's the norms of informal reasoning or the norms of heuristic reasoning.
2:33:11And then we have arguments that a heuristic argument verifier could accept. It's just like people wrote some words. I think those words, like, my guess would be like, you know, 95 % of the thing is mathematicians except is like really compelling feeling heuristic arguments are correct. And like, if you actually form a resume, like some of these aren't quite right or like, here's some corrections or here's which of two conflicting arguments is right. I think there's something to be learned from it. I don't think it would be like mind blowing though. When you have, have it completed, how big would this heuristic estimator, the rules for this heuristic estimator be?
2:33:38I mean, I know like when Russell and who was the other guy when they did the rules for watching. Yeah, yeah, wasn't it like literally they had like a bucket or a wheelbarrow that with all the papers, but how big would I mean mathematical foundations are quite simple in the end like at the end of the day It's like you know how many symbols like I don't know. It's hundreds of symbols or something that go into the entire foundations And the entire rules of reasoning for like you know, there's a sort of built on top of first or logic But the rules of reasoning for first or logic are just like you know another hundreds of symbols or 100 lines of code or whatever.
2:34:14I'd say like, I have no idea. Like we are certainly aiming at things that are just not that complicated. Like, and my guess is that the algorithms are looking for not that complicated. Like most of the complexity is pushed into arguments not in this like verifier or estimator. So for this to work, you need to come up with an estimator which is a way to integrate different touristy arguments together. Has to be a machine that takes the input like first to take the input argument decides what it believes in light of it, just kind of like saying was it compelling. But seconds, it needs to take four of those.
2:34:44And then say, here's what I believe in light of all four. Even though there's a different estimation strategy that produces different numbers. And that's a lot of our life is saying, well, here's a simple thing that seems reasonable. And here's a simple thing that seems reasonable. Like, what are you doing? There's supposed to be a simple thing that unifies them both. And the obstruction to getting that is understanding what happens when these principles are slightly intentioned and how do we deal? Yeah, that seems super interesting. Even, I mean, we'll see what other applications it has. I don't know, computer security and code checking, if you can put, like actually say, like this is how safely you think a code is in a very formal way.
2:35:15My guess is we're not gonna add, I mean this is both a blessing and a curse. Like it's a curse thing, you're like, well that's sad, the other thing is not that useful, but a blessing and not useful things are easier. My guess is we're not gonna add that much value in most of these domains. Like most of the difficulty comes from like, like a lot of code you'd want to verify, not all of it, but a significant part is just like, the difficulty of formalizing the proof is like the hard part and like actually getting all of that to go through. and we're not gonna help even the tiniest bit with that.
2:35:40I think. So this would be more helpful if you have code that like uses simulations, you wanna verify some property of like a controller that involves some numerical error or whatever you need to control the effects of that error. That's where you like start saying like, well, here's sickly, if the errors are independent, blah, blah, blah. Yeah, you're too honest to be a salesman, Paul. I think this is kind of like sales to us, right? Like if you talk about this idea of people, like why would that not be like the coolest thing ever and therefore impossible? And we're like, well actually it's kind of lame.
2:36:04And we're just trying to pitch like, it's way lairier than it sounds. And that's really important to why it's possible is being like, it's really not going to blow that many people's. I think it will be cool. I think it will be very, if we succeed, it will be very solid like metamathematics with theoretical computer science or whatever. But I don't think, right, I think the mathematicians already do this reasoning and they mostly just love proofs. I think the physicists do a lot of this reasoning, but they don't care about formalizing anything. I think like in practice, other difficulties are almost always going to be more salient.
2:36:31I think this is like of most interest by far for interpretability in ML. And I think other people should care about it, and probably will care about it if successful, but I don't think it's gonna be the biggest thing ever in any field or even that huge a thing. I think this would be a terrible career move given the ratio of difficulty to impact. I think theoretical computer science, it's probably a fine move. I think in other domains, it just wouldn't be worth. We're gonna be working on this for years, at least in the best case. I'm laughing because my next question was gonna be like a set up for you to explain.
2:37:05Well, if this, like, something to grasp, you didn't want to work on this. I think the radical computer science is an exception where I think this is like, in some sense, like what the best of the radical computer science is like. So like, you have all this reason, you have this like, because it's useless. What? I mean, I think like an analogy, I think like one of the most successful sagas in the real computer science is like formalizing the notion of an interactive proof system. And it's like you have some kind of informal thing that's interesting to understand. And you want to like pin down what it is and construct some examples and see what's possible and what's impossible.
2:37:42And this is like, I think this kind of thing is the bread and butter of like the best parts of theoretical computer science. And then again, I think mathematicians like it may be a career mistake because the mathematicians only care about proofs or whatever, but that's a mistake in some sense aesthetically. Like if successful, I do think looking back and again, part of why it's a mistake is such a high probability, we wouldn't be successful. But I think looking back, people would be like, that was pretty cool. Like, although not that cool, like we understand why it didn't happen, given like the epistemic, like what people cared about in the field, but it's pretty cool now.
2:38:12But, you know, is it also the case that didn't hearty -guriding that like, you know, all this crime shit is both not useless, but it's fun to do and like it turned out, but all the cryptographies based on all that crime shit. So, I don't know, you could have, but anyways, I'm trying to set you up so that you can tell, And forget about it, it doesn't have applications in all those other fields. It matters a lot for alignment. And that's why I'm trying to set you up to talk about if this smart, I don't know, I think a lot of smart people listen to this podcast, if they're a math or CS grad student and has gotten interested in this.
2:38:51Are you looking to potentially find talent to help you with this? Yeah, maybe we'll start there. And then I also want to ask you if I think also maybe a people who I couldn't provide funding might be listening to the podcast so To both of them. What is your pitch? There's a word definitely Definitely hiring and searching for collaborators. Yeah, I think the most useful profile is Probably a combination of like intellectually interested in this particular project I'm motivated enough by alignment to work on this project even if it's really hard I think there are a lot of good problems So the basic fact that makes this problem unappealing to work on, I'm a really good salesman.
2:39:30But the only reason this isn't a slam dunk thing to work on is that like There are not great examples. So we've been working on it for a while, but we do not have beautiful results as of the recording of this podcast Hopefully by the time it airs It's like they've had greater results since then, but it was too long to put in the margins of podcast. Yeah Which lock Yeah, so I think it's hard to work on because it's not clear what a success looks like. It's not clear success as possible. But I do think there's a lot of questions. We have a lot of questions. And I think the basic setting of, look, there are all of these arguments.
2:40:11So in mathematics, in physics, in computer science, there's just a lot of examples of informal heuristic arguments. They have enough structural similarity that it looks very possible that there is a unifying framework that these are instances of some general framework and not just a bunch of random things, not just a bunch of. It's not like, it's an example for the prime numbers. People are reason about the prime numbers as if they were a random set of numbers. One view is that's just a special fact about the primes, they're kind of random. A different view is actually it's pretty reasonable to reason about an object as if it was a random object as a starting point and then as you notice structure it revised from that initial guess.
2:40:45And it looks like to me the second perspective is probably more right. It's just like a reasonable to start off treating an object as random and then notice perturbations from random, like notice structure of the object possesses. And the primes are unusual and they have fairly little like additive structure. I think it's a very natural theoretical project. There's like a bunch of activity that people do. It seems like there's a reasonable chance. There's something nice to say about unifying all of that activity. I think it's a pretty exciting project. The basic strike against it is that it seems really hard.
2:41:12Like, if you were someone's advisor, I think you'd be like, what are you going to prove if you work on this for the next two years? And it'd be like, there's a good chance nothing. And then like, it's not what you do. If you're a PhD student normally, you have like, you aim for those high probabilities of getting something within a couple of years. The flip side is it does feel, I mean, I think there are a lot of questions. I think some of them were probably going to make progress on. So I think the pitch is mostly like, are some people excited to get in now or like are people more like, ah, let's wait to see like, once we have like one or two good successes to see what the pattern is and become more confident we can turn the crank to make more progress in this direction.
2:41:43But for people who are excited about working on stuff with reasonably high probabilities of failure and not really understanding exactly what you're supposed to do, I think it's a pretty good project. I feel like if people look back, if we succeed, and people looking back in like 50 years on what was the coolest stuff happening in math with theoretical computer science, they'll be like, this will definitely be like in contention. And I would guess for lots of people would just see like the coolest thing from this period of a couple of years or whatever. right? Because this is a new method in like so many different fields from the once you met physics math the computer science like that's that's that's really I don't know but what is the average math PhD working on right he's not or he's working on like some subset of a subset of something I can't even understand or pronounce but map is quite a satiric but yeah this seems like I don't know even this small chance of working like forget about the value for you shouldn't forget about the value for alignment, but even without that, this is such a cool, if this works, it's like a really big deal.
2:42:42There's a good chance that if I had my current set of views about this problem and didn't care about alignment and had the career safety to just spend a couple of years thinking about it, I'm going to spend half my time for like five years or whatever that I would just do that. I mean, even without caring at all about alignment, it's just a nice, it's a very nice problem. It's very nice to have this like library of things that succeed where like it's just, they feel so tantalizingly close to being formalizable at least to me. and such a natural setting, and then just have so little purchase on it.
2:43:10It's like, there aren't that many really exciting feeling frontiers in like, throughout computer science. And then, so, smart person doesn't have to be aggressive, but like a smart person is interested in this. What should they do? Should they try to attack some open problem you have, put on your blog, or what is the next step? Yeah, I think like a first -pass step, There's different levels of ambition or whatever, different ways of approaching a problem. But we have this write up from last year, August 11, months ago or whatever, informizing the presumption of independence that provides a communication of what we're looking for in this object.
2:43:51And I think the motivating problem is saying, here's a notion of what an estimator is, and here's what it would mean for an estimator to capture some set of informal arguments. And a very natural problem is just try and do that. go for the whole thing, try to understand, and then come up with hopefully a different approach, or then end up having contexts from a different angle on the approach we're taking. I think that's a reasonable thing to do. I do think we also have a bunch of open problems. Maybe we should put out more of those open problems. The main concern with doing so is that for any given one, we're like, this is probably hopeless.
2:44:22Put up a prize earlier in the year for an open problem, which tragically, I guess the time is now to post the debrief from that, or I owe it from this weekend, I was supposed to do that. So I'll probably do it tomorrow. But no one solved it. It's sad putting out problems that are hard. We could put out a bunch of problems that we think might be really hard. But what was that famous case of that statistician who it was like some PhD student who showed up late to a class and he saw some problems in the ward and he thought they were homework and then they were actually just open problems and then he solved them because he thought they were homework, right?
2:44:55Yeah. I mean we don't we have much less information these problems are hard. Again, I expect the solution to most of our problems to not be that complicated. We have not, and we've been working on it in some sense for a really long time. Total years of full -time equivalent work across the whole team is probably three years of full -time equivalent work in this area. Spread across a couple of people. But that's very little compared to a problem. It is very easy to have a problem where you put in three years of full -time equivalent work. But in fact, there's still an approach that's going to work quite easily with like three to six months if you come at a new angle.
2:45:28And like, we've learned a fair amount from that that we could share, and we probably will be sharing more of the coming months. Um, as far as money goes, is this something where, I don't know if somebody gave you a whole bunch of money, it would help or it doesn't not matter. How do you do your working on this, by the way? So we have been right now, there's four of us full time, and we're hiring for more people. And then, um, what is funding that would matter? I mean, funding is always good. We're not super funding constrained right now. The main effect of funding is it will cause me to to continuously and perhaps indefinitely delay fundraising.
2:45:58Periodically, I'll set out to be interested in fundraising when someone will be like offer a grant and then I will get to delay for another six months or fundraising or nine months or whatever. So you can delay the time at which Paul needs to think for some time about fundraising. Well, one question I think would be interesting to ask you is, I think people can talk vaguely about the value of theoretical research and how it contributes to real -world applications and you can look at historical examples or something. But you are somebody who actually has done this in a big way, like RLHF is, you know, something you developed and then it actually has got into an application that has been used by millions of people.
2:46:35So tell me about like, just that pipeline. How can you reliably identify theoretical problems that will matter for a real load applications? Because it's one thing to like read about touring or something and the whole thing problem, but like, here you have the real thing. Yeah, I mean, it is definitely exciting to have worked on a thing that has a real world impact. The main caveat I'd provide is like, our LHF is very, very simple compared to many things. And like, so the motivation for working on that problem was like, look, this is how it probably should work. Or like, this is a step in some like progressions.
2:47:09It's unclear if it's like the final step or something. It's a very natural thing to do that like, people probably should be and probably will be doing. I'm saying, if you want to talk about crazy stuff, it's good to help make those steps happen faster. And it's good to learn about what are there's lots of issues that occur in practice, even for things that seem very simple on paper. But most of the story is just like, my sense of the world is things that look like good ideas on paper, just often are harder than they look, but the world isn't that far from what makes sense on paper. Large language models look really good on paper and Arleich have looks really good on paper.
2:47:44And these things, I think, just work out in a way that's... Yeah, I think people maybe overestimate or like... Maybe it's kind of a trope. But people talk about it's easy to underestimate how much gap there is to practice. How many things will come up that don't come up in theory. But it's also easy to overestimate how unscrutable the world is. The things that happen mostly are things that do just make sense. Yeah, I feel like most ML implementation does just come down to a bunch of detail, though. Build a very simple version of the system. So I understand what goes wrong, fix the things that go wrong, scale it up, understand what goes wrong.
2:48:17And I'm glad I have some experience doing that, but I don't think that does cause me to be better informed about what makes sense in ML and what can actually work. But I don't think it caused me to have a whole lot of deep expertise or deep wisdom about how to close the gap. Yeah. But is there some tip on identifying things like RLHF, which actually do matter versus making sure you don't get stuck in some theoretical problem that doesn't matter. Or is it just coincidence or I mean, is there something you can do in advance to make sure that the thing is useful? I don't know if the RLHF story is like the best, best discussed case or something, but like...
2:48:56Oh, because they had the capabilities. Maybe I'd say more profoundly like again, it's just not that hard a case. It was like a little bit... It's a little bit unfair to be like, I'm going to predict the thing which I like, I pretty much think it was going to happen at some point. And so it was most like case of acceleration. Whereas like the work we're doing right now is specifically focused on something that's like kind of crazy enough that it might not happen even if it's a really good idea or challenging enough it might not happen. But I'd say like in general like, and this draws a little bit on like more broad experience more probably in theory.
2:49:25It's just like a lot of the times when theory fails to connect with practice. It's just kind of clear. It's not going to connect if you like try, if you actually think about it and you're like, what are the key constraints in practice? is the theoretical problem we're working on actually connected to those constraints? Is there a path, like, is there something that is possible in theory that would actually address like real world issues? I think like the vast majority, like as a theoretical computer scientist, the vast majority of theoretical computer science, like has very little chance of ever affecting practice, but also it is completely clear in theory that has very little chance of affecting practice.
2:49:56Like most of the theory fails to affect practice not because of like all the stuff you don't think of, but just because like it was like, you could call it like dead on arrival because it's not really the point, just mathematicians also are like, they're not trying to affect practice, and they're not like, why does my number of theory not affect practice? It was kind of obvious. So I think the biggest thing is just like, actually caring about that, and then like learning at least what's basically going on in the actual systems you care about and what are actually the important constraints, and is this a real theoretical problem?
2:50:23The basic reason most theory doesn't do that is just like, that's not where the easy theoretical problems are. So I think theory is instead motivated by like we're gonna build up the edifice of theory, and like sometimes they'll be opportunistic, like opportunistically we'll find a case that comes close to practice or we'll find something practitioners are already doing and try and bring into our framework or something. But the three of changes mostly not. This thing is going to make it into practice. It's mostly because it's going to contribute to the body of knowledge that will slowly grow and like sometimes opportunistically yields important results.
2:50:50How big would do you think a CDI would be? What is that minimum sort of encoding of something that is as smart as a human? I think it depends a lot what substrate it gets to run on. So if you tell me like like how much computation does it get before or like what kind of real world infrastructure does it get Like you can ask what's the shortest program which like if you run it on a million H100s connected in like a nice network With like a hospitable environment will eventually like go to the stars But that seems like it's probably on the order of like tens of thousands of bytes or I don't know if I had to guess the medium I'd guess 10 ,000 bytes But wait the specification or the compression of just the program a program which went wrong.
2:51:26Oh got it But that's gonna be like really cheesy So they could ask what's the thing that has values and will expand and roughly preserve its values. Yeah, that proceeds. Because that thing, the 10 ,000 byte thing, will just lean heavily on evolution and natural selection to get there. For that, I don't know. Million bytes. Million bytes. 100 ,000 bytes. Something like that. How would you think AI -Li detectors will work? Where you just look at the activations. And not in the, not fine explanations in the way you were talking about with heuristics, but literally just like, here's what truth looks like here's what lies look like, let just segregate the lane space and see if we can identify the two.
2:52:11Yeah, I think to separate the, like, just train a classifier to do it, it's like a little bit complicated for a few reasons and like, may not work, but if you just, like, brought in the space and say, like, hey, it's like, you want to know if someone's lying, you get to interrogate them, but also you get to like rewind them arbitrarily and make a million copies of them. I do think it's pretty hard to lie successfully. You get to look at their brain, even if you don't quite understand what's happening. You get to rewind them a million times, you get to run all those parallel copies and do gradient descent or whatever.
2:52:35I think it's a pretty good chance that you can just tell if someone is lying. Like a brain emulation or an AI or whatever. Unless they were aggressively selected. Like if it's just they are trying to lie well rather than it's like they were selected over many generations to be excellent at lying or something, then your ML system hopefully didn't train it a bunch to lie and you want to be careful about your training scheme effectively does that But yeah, that seems like it's more likely than not to succeed And how possible do you think it will be for us to specify human verifiable rules for reasoning such that Even if the AI is super intelligent we can't really understand why does certain things We know that the way in which it arises at these conclusions is valid like it was trying to persuade us to something We can be like, I don't understand all the steps, but I know that this is something that's valid and you're not just making shit up.
2:53:27That seems very hard if you wanted to be like competitive with learned reasoning. So like, I mean, it depends a little bit exactly how you set it up, but for like the ambitious versions of that, let's say what address do I'm in problem? They seem pretty unlikely, you know, like 5%, 10 % kind of thing. Is there an upper bound on intelligence? Not in the near -term, but just like super intelligence, at some point, how far do you think that can go? It seems like it's going to depend a little bit on what is meant by intelligence. It kind of reads as a question that's similar. So is there an upper bound on strength or something?
2:53:59There are a lot of forms. I think it's the case that for, yeah, I think there are sort of arbitrarily smart input output functionalities. And then if you hold fixed the amount of compute, there is some smartest one. If you're just like, what's the best set of 10 to the 40th operations? There's only finally many of them. So some best one for any particular notion of best that you have in mind. So I guess like I'm just like for the unbounded question where led to arbitrary description, complexity and compute, like probably now and for the, I mean there is some like optimal context like if you're like I have some goal in mind and I'm just like what action best achieves it if you imagine like a little box in that in the universe I think there's kind of just like an optimal input output behavior.
2:54:35So I guess in that sense I think there is a is an upper bound but it's not saturatable in the physical universe because it's definitely exponentially slow. Right yeah yeah or you know because of or other things, or it just might be physically impossible to insatiate something smarter than this. Yeah, I mean, like, for example, if you imagine what the best thing is, it would almost certainly involve just like simulating every possible universe it might be in modular, like, moral constraints, which I don't know if you want to include them. But like, so that would be very, very slow. It would involve simulating like all, you know, it's sort of like, I don't know exactly how slow, but like, the blood spinach will very slow.
2:55:11Carl Schoelman laid out his picture of the intelligence explosion in the seven -hour episode. What I know you guys have talked a lot. What about his basic picture? Like do you have some main disagreements? Is there some crux that you guys have explored? It's related to our timeline's discussion from earlier. I think the biggest issue is probably Arab bars. Where Carl has a very like very software focused, very fast kind of takeoff picture. And I think that is plausible, but not that likely. I think there's a couple ways you could perturb the situation and my guess is one of them applies. So maybe I have like, I don't know exactly what crawls probably.
2:55:50I feel like crawls kind of like a 60 % chance on some crazy thing that I'm only going to assign like a 20 % chance to or 30 % chance or something. And like I think those kinds of perturbations are like, one, how long a period is there of complementarity between AI capabilities and human capabilities, which will tend to soften, take off, to like how much diminishing returns are there on software progress, such that like is a like broader takeoff involving scaling like electricity production and hardware production is that likely to happen during takeoff where I'm like more like 5050 or more stuff like this.
2:56:24Yeah okay so is it that you think the ultimate constraints will be more heart or like the basic case is laid out is that you can just have a sequence of things like flash attention or MOE and you can just keep stacking these kinds of things on. I'm very unsure if you can keep stacking them or like it's kind of a question of what's like the returns curve and like Carl has some inference from historical data or some way he'd extrapolate the trend. I am more like 5050 on whether the software on the intelligence explosion is even possible and then like a somewhat higher probability that it's slower than hard work.
2:56:56But what do you think it might not be possible? Well so the whole entire question is like if you double R &D effort do you get enough additional improvement to further double the efficiency and like that's that question that will would self be a function of your hardware based, like how much hardware you have. And the question is, at the amount of hardware we're going to have, and the level of sophistication we have as the process begins, is it the case that each doubling of, actually the initials only depends on the hardware, or like each level of hardware will have some place at the dynamic asymptotes.
2:57:24So the question is just like, for how long is it the case that each doubling of R &D, at least doubles the effective output of your AI research population? And I think like I have a higher probability on that. I think it's kind of close if you look at the empirics. I think the empirics benefit a lot from continuing hardware scale up. So the effective R &D stock is significantly smaller than it looks, if that makes sense. What are the inferiors you're referring to? So there's two sources of evidence. One is looking across a bunch of industries of what is the general improvement with each doubling of either R &D investment or experience?
2:57:57It is quite exceptional to have a field with, not... Anyway, it's pretty good to have a field where each time you double R &D investment, you go to doubling of efficiency. The second source of evidence is on actual, like algorithmic improvement in ML, which is obviously much, much scarcer. And they're like, you can make a case that it's been like each doubling of R &D has given you, like roughly a 4x or something, increase in computational efficiency. But like there's a question of how much that benefits. When I say the effect of R &D stock is smaller, I mean like we scale up, like you're doing a new task.
2:58:27Like every couple of years we're doing a new task because you're operating a scale much larger than the previous scale. And so a lot of your effort is how to make use of the new scale. So if you're not increasing your installed hardware base and just flat at a level of hardware, I think you get like much faster diminishing returns than people have gotten historically. I think Carl agrees in principle, this is true. And then once you make that adjustment, I think it's like a very unclear where the empiric shake out. I think Carl has thought about these more than I am, so I should maybe defer more, but anyway, I'm at like 50 -50 on that.
2:58:52How have your timeline changed over the last 20 years? Last 20 years? Yeah. How long have you been working on anything related to AI? So I started thinking about this stuff in like 2010 or so. So I think my first, my earliest timeline prediction will be in like 2011. I think in 2011 my rough picture was like, we will not have insane AI in the next 10 years. And then I get increasingly uncertain after that, but we converge to like 1 % per year or something like that. And then probably in 2016, my take was like, we won't have crazy AI in the next five years, but then we converge to like one or two percent per year after that.
2:59:32Then in 2019, I guess I made around a forecast where I gave like 30 % or something to 25 % to a crazy eye by 2040 and like 10 % by 2030 or something like that. So I think my 2030 probability has been kind of stable and my 2040 probability has been going up And I would guess it's too sticky. I guess that 40 % I give at the beginning is just like from not having updated recently enough and I maybe just need to sit down. I would guess that should be even higher. I think like 15 % in 2030, I'm not feeling that bad about. This is just like each passing year is like a big update against 2030. Like we don't have that many years left.
3:00:13And that's like roughly counterbalanced so they add going pretty well. Whereas for like the 2040 thing, like the passing years are not that big a deal. And like as we see the like things are basically working, that's like cutting out a lot of the probability of not having a add by 2040. So yeah, my 2030 probability up a little bit, like maybe twice as high as it used to be, or like something like that, my 2040 probably, like up more, much more significantly. How fast do you think we can keep building fabs to keep up with the IDMAN? Yeah, I don't know much about any of the relevant areas. My best guess is like, maybe my understanding is right now like 5 % or something of like like the next year's total or best process fabs will be making AI hardware of which only a small fraction will be going into very large training runs, only a couple of, maybe a couple of percent of total output and then that represents maybe 1 % of total possible output, a couple of percent of leading process, 1 % of total or something.
3:01:10I don't know if that's right, but it's like the rough ballpark run. I think things will be pretty fast as you scale up for the next order of magnitude or two from there because you're basically just shifting over other stuff. In my sense is it would be like years of delay. There's like multiple reasons that you expect years of delay for going past that. Maybe even at that you start having, yeah, there's just a lot of problems. Like building new fabs is quite slow. And I don't think there's like, TSMC is not like planning on increases in total demand driven by AI, like kind of conspicuously not planning on it.
3:01:43I don't think anyone else is really ramping up production in anticipation either. So I think, and then similarly, just building data centers of that size seems very, very hard, and also probably has multiple years of delay. What is your portfolio look like? I've tried to get rid of most of the AI stuff that's like, possibly implicated in policy work or like, see, the advocacy on the RSP stuff from my involvement with Anthropic. Or what would it look like if you had no complex interest? No insight information. I also still have a bunch of hardware investments, which I need to think about. But like, I don't know, a lot of TSMC, I have a chunk of Nvidia, although I keep, I just keep betting against Nvidia constantly since 2016 or something.
3:02:24I've been destroyed on that bet, although AMD has also done fine. And it's like, well, now the case now is even easier, but it's similar to the case in the old days, just a very expensive company, given the total amount of R &D investment they've made. They have like whatever, a trillion dollar valuation or something. That's like very high. So the question is like how expensive is it to like make a TPU such that's like actually outcompete Age 100 or something and I'm like wow It's real level high level of incompetence if Google can't like catch up fast enough to like make that trillion dollar valuation not justified Where as a TSMC it's much hard they have a harder remote you think yeah I think it's it's a lot harder especially if you're in this regime where like you're trying to scale up So like if you're unable to build fabs I think we'll take a very long time to build many fabs as people want like the effect of that will be to bid up the price of existing fabs and existing semiconductor manufacturing equipment.
3:03:15So those hard assets will become, like, spectacularly valuable. As well as existing GPUs and the actual, yeah. Yeah, I think it's just hard. That seems like the hardest asset to scale up quickly. So it's like the asset. If you have a rapid run up, it's the one that you'd expect to most benefit. Whereas NVIDIA stuff will ultimately be replaced. By either it's better stuff made by humans or stuff made by with AI assistants. like the gap will close even further as you build AI systems. Right, unless Nvidia is using those systems. Yeah, the point is just that like, any video you're using will so dwarf past our end days.
3:03:48Right, I see. And there's like, just not that much stickiness. There's less stickiness in the future than there has been in the past. Like, yeah. I don't know. So I don't want to, not commenting from any private information, just in my gut, having caveat of this is like the single bed I've most lost. Like, not including Nvidia in that portfolio. And final question, there's a lot of schemes out there for alignment. And I think just like a lot of general takes. And a lot of this stuff is over my head, where I think I literally took me like weeks to understand the mechanism anomaly stuff you work on.
3:04:18Without spending weeks, how do you detect bullshit? People have explained their schemes to me, and I'm like, honestly, I don't know if it makes sense or not. With you, I'm just like, I trust Paul enough that I think there's probably something here if I try to understand this enough. But yeah, how do you detect bullshit? Yeah, so I think it depends on the kind of work. So if like the kind of stuff we're doing, my guess is like most people, there's just not really a way you're gonna Tell whether it's bullshit. So I think like it's important that we don't spend that much money on like The people who want to hire are probably gonna dig in and depth I don't think there's a way you can tell whether it's bullshit without either spending like a lot of effort or leaning on deference Within empirical work, it's like interesting and that you do have some signals of the quality of work Like I mean, is it does it work in practice?
3:05:01Like does the story? I think the stories are just radically simpler and so you probably can evaluate those stories like on their face. And then you mostly come down to these questions about what are the key difficulties. Yeah, I tend to be optimistic. When people dismiss something because this doesn't deal with a key difficulty or this runs into the following and super boob obstacle, I tend to be a little bit more skeptical about those arguments and tend to think like, yeah, something can be bullshit because it's not addressing a real problem. That's like I think the easiest way. This is a problem someone's interested in.
3:05:27That's just not actually an important problem and there's no story about why it's gonna become an important problem. EG, it's not a problem. Now, in one kit worse, or it is maybe your problem now, but it's clearly getting better. That's one way. And then, conditionally, I'm passing that bar, like dealing with something that actually engages with important parts of the argument for concern. And then, actually, making sense empirically. So I think most work is anchored by its source of feedback, because actually engaging with real models. So it's like, does it make sense to have engage with real models?
3:05:55And does the story about how it deals with key difficulties actually makes sense? I'm pretty liberal past there. I think it's really hard to like, each of you people look at mechanistic interpretability and be like, well, this obviously can't succeed. And I'm like, I don't know, how can you tell that it's basically can succeed? I think it's reasonable to take total investment in the field to how fast is the making progress? Like, how does that pencil, I think most things people work on, they actually pencil like pretty fine. Like they look like they could be reasonable investments. Things are not like super out of whack.
3:06:28Okay, great. This is I think a good place to close. Paul, thank you so much for your time. Yeah, thanks for having me. It was a good time. Yeah, absolutely. Hey everybody. I hope you enjoyed that episode. As always, the most helpful thing you can do is just share the podcast. Then did two people you think might enjoy it, put it in Twitter, your group chats, etc. Just blitz the world. Appreciate your listening. I'll see you next time. Cheers.
From the publisher
Paul Christiano is the world’s leading AI safety researcher. My full episode with him is out!
We discuss:
- Does he regret inventing RLHF, and is alignment necessarily dual-use?
- Why he has relatively modest timelines (40% by 2040, 15% by 2030),
- What do we want post-AGI world to look like (do we want to keep gods enslaved forever)?
- Why he’s leading the push to get to labs develop responsible scaling policies, and what it would take to prevent an AI coup or bioweapon,
- His current research into a new proof system, and how this could solve alignment by explaining model's behavior
- and much more.
Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform. Read the full transcript here. Follow me on Twitter for updates on future episodes.
Open Philanthropy
Open Philanthropy is currently hiring for twenty-two different roles to reduce catastrophic risks from fast-moving advances in AI and biotechnology, including grantmaking, research, and operations.
For more information and to apply, please see the application: https://www.openphilanthropy.org/research/new-roles-on-our-gcr-team/
The deadline to apply is November 9th; make sure to check out those roles before they close.
Timestamps
(00:00:00) - What do we want post-AGI world to look like?
(00:24:25) - Timelines
(00:45:28) - Evolution vs gradient descent
(00:54:53) - Misalignment and takeover
(01:17:23) - Is alignment dual-use?
(01:31:38) - Responsible scaling policies
(01:58:25) - Paul’s alignment research
(02:35:01) - Will this revolutionize theoretical CS and math?
(02:46:11) - How Paul invented RLHF
(02:55:10) - Disagreements with Carl Shulman
(03:01:53) - Long TSMC but not NVIDIA
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe




