Joe Carlsmith - Otherness and control in the age of AGI

22 Aug 2024 · 2 h 31 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dwarkesh Podcast: Episode Summary

Episode Title

Joe Carlsmith - Otherness and Control in the Age of AGI

Description In this episode, host Dwarkesh Patel engages Joe Carlsmith in a deep discussion on critical themes related to artificial general intelligence (AGI), including trust in power and technology, the dangers of excessive control, and how to approach the concept of the artificial 'Other.' Carlsmith reflects on his sequence titled "Otherness and Control in the Age of AGI," offering insights into the moral implications of our interactions with AI and the broader ethical questions that arise.

Key Themes and Discussions

  1. Trust in Power and Technology
  2. Debate over reliance on techno-capital and power structures.
  3. Concerns about moral alignment of AI and its potential implications for society.
  1. Control and Historical Context
  2. Warning against the desire for absolute control, citing historical figures like Stalin.
  3. Discussion about the balance between control and nurturing a cooperative relationship with AI.
  1. Gentleness Towards the Artificial Other
  2. Emphasis on treating AI as entities with potential moral considerations.
  3. Exploration of the concept of 'Otherness' and its implications for ethical treatment.
  1. The Nature of AI and Human Values
  2. Examination of AI's understanding of human values and the challenge of aligning these effectively.
  3. Discussion on the nature of verbal behavior in AI and how it may not reflect their underlying decision-making processes.
  1. Exploration of Moral Realism
  2. Potential convergence of AI understanding and human moral frameworks.
  3. The implications of moral realism and how it relates to our ethical treatment of AI.

Key Takeaways

  • Historical Awareness: Understanding historical tendencies towards control is crucial in managing future relationships with AI.
  • Moral Complexity: The moral landscape of AI is complex, requiring careful consideration of how we align their values with human ideals.
  • Ethical Dilemmas: The relationship with AI raises significant ethical questions, especially regarding their autonomy, rights, and the potential for harm.
  • Balance of Power: There is a critical need to maintain a balance of power in the development and deployment of AI technologies, avoiding centralization of control.
  • Human Values: Humans must reflect on what values we wish to instill in AI and how these values interact with the broader ethical frameworks of society.

Noteworthy Quotes

  • "We have a rich ethical tradition to draw from as we consider the implications of creating beings with autonomous decision-making capabilities."
  • "Gentleness and control can coexist, but we must navigate this tension carefully to avoid repeating past mistakes."

Related Links

  • [Joe Carlsmith’s Sequence on Otherness and Control in the Age of AGI](https://joecarlsmith.com/2024/01/02/otherness-and-control-in-the-age-of-agi)
  • [Watch on YouTube](https://www.youtube.com/watch?v=5XsL_7TnfLU)
  • [Listen on Apple Podcasts](https://podcasts.apple.com/us/podcast/joe-carlsmith-otherness-and-control-in-the-age-of-agi/id1516093381?i=1000666255737)

Timestamps

  • (00:00:00) - Understanding the Basic Alignment Story
  • (00:44:04) - Monkeys Inventing Humans
  • (00:46:43) - Nietzsche, C.S. Lewis, and AI
  • (1:22:51) - How Should We Treat AIs
  • (1:52:33) - Balancing Being a Humanist and a Scholar
  • (2:05:02) - Explore-Exploit Tradeoffs and AI

Conclusion This episode of the Dwarkesh Podcast brings to light the significant ethical considerations surrounding the development and integration of AGI into society. Carlsmith's insights serve as a thought-provoking guide for navigating the complexities of human-AI relationships in our rapidly evolving technological landscape.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today I'm chatting with Joe Carl Smyth, he's a philosopher in my opinion, a capital G great philosopher, and you can find his essays at joecarlsmyth .com. So we have a GVT4, and it doesn't seem like a paperclip or kind of thing, it understands human values. In fact, if you help have it explain, like why is being a paperclip or bad? Or like what would just tell me your opinions about being a paperclip? Or like explain why the galaxy shouldn't be turned into paperclips. Okay, so what is happening such that dot, dot, dot, we have a system that takes over and converts the world into something valueless.

0:37One thing I'll just say off the bat is like when I'm thinking about misaligned AIs, I'm thinking about or the type that I'm worried about. Yeah. I'm thinking about AIs that have a relatively specific set of properties related to agency and planning and and kind of awareness and understanding of the world. One is this capacity to plan and kind of make kind of relatively sophisticated plans on the basis of models of the world, where those plans are being kind of evaluated according to criteria. That planning capability needs to be driving the model's behavior. So there are models that are sort of in some sense capable of planning, but it's not like when they give output.

1:13It's not like that output was determined by some process of planning. Like here's what'll happen if I give this output. And do I want that to happen? The model needs to really understand the world, right? It needs to really be like, okay, here's what will happen. I'm, you know, here I am, here's my situation, here's the politics of the situation. I really like kind of having this kind of situational awareness to be able to evaluate the consequences of different plans. I think the other thing is like, so the verbal behavior of these models, I think need bear note. So when I talk about a model's values, I'm talking about the criteria that kind of end up determining which plans the model pursues.

1:56And a model's verbal behavior, even if it has a planning process, which GPT -4, I think, doesn't, in many cases, it's verbal behavior just doesn't need to reflect those criteria. Yeah, right. And so we know that we're going to be able to get models to say what we want to hear, right? We, that is the magic of gradient descent. If you modulo like some difficulties with capabilities, like you can get a model to kind of output the behavior that you want. If it doesn't, then you crank it till it does, right? And I think everyone admits for suitably sophisticated models they're going to have very detailed understanding of human morality.

2:41But the question is like, what relationship is there between like a model's verbal behavior, which is you've essentially kind of clamped. The model must say like, blah things. And the criteria that end up influencing its choice between plans. And there I think it's at least, I'm kind of pretty cautious about being like, well, Well, when it says the thing I forced it to say, or like, you know, gradient descent to that such that it says, that's a lot of evidence about like, how it's gonna choose in a bunch of different scenarios. I mean, for one thing, like, even with humans, right? It's not necessarily the case that humans, they're kind of verbal behavior reflects the actual factors that determine their choices.

3:25They can lie, they cannot even know what they're, what they would do in a given situation. I mean, I think it is interesting to think about this in the context of humans, because there is that famous saying of be careful who you pretend to be, because you are who you pretend to be. And you do notice this, where if people, I don't know, this is what culture does to children, where you're trained, like your parents will punish you if you say, if you start saying things that are not consistent with your culture's values. And over time, you will become like your parents, right? Like by default, it seems like it kind of works.

3:54And even with these models, it seems like it's kind of a work. So it's like hard, they don't really scheme against, like why would this happen? You know, for folks who are kind of unfamiliar with the basics, right, maybe folks are like, why are they digging over it all? Like, what is the literally, any reason that they would do that? So, the general concern is like, if you're really offering someone, especially if you're really offering someone like power for free, power almost by definition is kind of useful for lots of values. And if we're talking about an AI that really has the opportunity to kind of take control of things, if some component of its values is sort of focused on some outcome, like the world being a certain way and especially in a longer -term way, such that the horizon of its concern extends beyond the period that would take over a plan would encompass.

4:42Then the thought is just often the case that the world will be more the way you want it. If you control everything, then if you remain the instrument of the human will or of some other actor, which is what we're hoping you say is will be. So that's a very specific scenario. And if we're in a scenario where power is more distributed, and especially where we're doing like decently on alignment, and we're giving the AI some amount of the inhibition about doing different things, and maybe we're succeeding in shaping their values somewhat. Now it is, I think it's just a much more complicated calculus, and you have to ask, okay, watch the upside for the AI.

5:20Watch the probability of success for this takeover path. How good is its alternative? So maybe this is a good point to talk about how you expect the difficulties of alignment into change in the future. We're starting off with something that has this intricate representation of human values. And it doesn't seem that hard to sort of lock it into a persona that we are comfortable with. I don't know. What changes? So, you know, why is alignment hard in general, right? Like let's say, let's say we've got an AI. And let's, again, let's bracket the question of like exactly how capable would it be? And really, just talk about this extreme scenario of like, it really has this opportunity to take over, right?

5:59which I do think maybe we just want to not want to deal with that with having to build an AI that we're comfortable being in that position. But let's just focus on it for the sake of simplicity and we can relax the assumption. Okay, so he has some hope. It's like I'm going to build an AI over here. So one issue is you can't just test. You can't give the AI this literal situation, have it take over and kill everyone and then be like oops, like update the weights. This is the thing Eliezer talks about of sort of like you can't, you know, you care about its behavior on this like specific, in the specific scenario that you, you can't test directly.

6:33Now we can talk about whether that's a problem, but that's like one issue is that there's a sense in which this has to be kind of like off -distribution and you have to be getting some kind of generalization from your training the AI in on a bunch of other scenarios. And then there's this question of how is it going to generalize to the scenario where it really has this option? So is that even true? Because when you're training it, you can be like, hey, here's a greeting update. If you get the takeover option on the platter, don't take it. And then just like in just sort of red teaming situations where things that has a takeover attempt, it's like you trained not to take it.

7:09And yeah, it could feel, but I just feel like if you did this to a child, you're like, I don't know, don't beat up your siblings. And the kind of the kid will generalize to,

7:22I'm going to start shooting random people. Yeah, so okay, cool. So you had mentioned this, a thought like, well, are you kind of what you pretend to be? And will you say, you train them to look kind of nice? You know, fake it till you make it. Or you were like, ah, we did this to kids. I think it's better to imagine kids doing this to us. So I don't know. like, here's a sort of silly analogy for like AI training. And there's a bunch of questions we can ask about. It's it's related to, but like suppose, suppose you, you know, you wake up and you're, you're being trained via like methods analogous to kind of contemporary machine learning by like Nazi children to be like a good Nazi soldier or Butler or what have you, right?

8:13And here are these children. And you really know what's going on, right? The children have like, they have a model spec, like a nice Nazi model spec, right? And it's like reflect well on the Nazi party, like benefit the Nazi party, whatever. And you can read it, right? You understand, this is why I'm saying like, when the model, you're like, oh, the model's really understand human values. It's like, yeah. So, yeah, go ahead. I miss analogy. I feel like a closer analogy would be, in this analogy, I start off as something more intelligent than the things for any meat with different values to begin with.

8:46So the intelligence and the values are baked into begin with. Where as more analogous scenario is, I'm a toddler and initially I'm like stupider than the children and I'm like being, this should also be true by the way, I have much more model. Initially, the much more model is dumb and gets smarter as you train it. So it's like a toddler and the kids are like, hey, we're going to bully you if you're not a Nazi. and I'm like, as you grew up, then you're like, the children's level, and then eventually you become an adult, but through that process, they've been sort of bullying you, you know, like training you to be a Nazi.

9:21And I'm like, I think that's an area where I might end up in Nazi. Yes, I think that's, so yeah, I think basically a decent portion of the hope here, or like, I think we should just, an aim should be whenever in the situation where the AI really has very different values already. It's quite smart, really knows what's going on. Yeah. and is now in this kind of adversarial relationship with our training process, right? So we wanna avoid that. The main thing, and I think it's possible we can by the sorts of things we're saying. So I'm not like, ah, that'll never work. The thing I just wanted to highlight was like if you get into that situation, and if the AI is genuinely at that point, like much, much more sophisticated than you,

9:59and doesn't want to kind of reveal its true values for whatever reason, then when the children show like some kind of obviously fake opportunity to like defect to the allies, right? You, you know, you sort of not necessarily going to be a good test of what will you do in real circumstance because you're able to tell. Okay, I can also give another way in which I think the analogy might be misleading, which is that no, imagine that you're not like just in a normal prison where you're like totally cognizant of everything that's going on. Sometimes they drug you, like give you like weird hallucinogens that totally mess up how your brain is working.

10:37A human adult in a prison is like, I know what kind of thing I am. I am like, like nobody's like really fucking with me in a big way. Whereas I think an AI, even a much smarter AI in a training situation is much closer to you're constantly inundated with weird drugs and different training protocols. And like, you're like frazzled because each moment it's closer to some Chinese water torture kind of technique where you're like, I'm glad we're talking about the moral patient. The chance to step back and be like, what's going on in this? Adult has that maybe in prison in a way that I don't know if these models necessarily have that coherence and that stepping back from what's happening in the training process.

11:27Yeah, I don't know. I think I'm hesitant to be like, it's like drugs for the like I think there's, but broadly speaking, I do basically agree that I think we have like really quite a lot of tools and options for kind of training AIs, even AIs that are kind of somewhat smarter than humans. I do think you have to actually do it. So I am, I think compared to maybe you had LAZ around, like I think I'm much more bullish on our ability to solve this problem, especially for AIs that are in what I think of as like the AI for AI safety sweet spot, which is this sort of band of capability where they're both very sufficiently capable that they can be really useful for strengthening various factors in our civilization that can make us safe.

12:13So our alignment work, control, cyber security, general epistemic, maybe some coordination application, stuff like that. There's like a bunch of stuff you can do with AI's that in principle could kind of differentially accelerate our security with respect to the sorts of considerations we're talking about. If you have a eyes that are capable of that, and you can successfully elicit that capability in a way that's not sort of being sabotaged or like messing with you in other ways, and they can't yet take over the world or do some other sort of really problematic form of power seeking, then I think if we were really committed, we could like really go hard, put a ton of resources really differentially direct to this like glut of AI productivity towards these sort of security factors, and hopefully kind of control and understand, you know, do a lot of these things.

13:01You're talking about for kind of making sure our AI's don't kind of take over or mess with us in the meantime. And I think we have a lot of tools there. I think you have to really try though. It's possible that those sorts of measures just don't happen or don't happen at the level of kind of commitment and diligence and like seriousness that you would need, especially if things are like moving really fast and there's other competitive pressures. And the compute, this is going to take compute to do these intensive, all these experiments on the AI's and stuff in that compute. We could use that for experiments for the next scaling step and stuff like that.

13:32So I'm not here saying like this is impossible, especially for that band of AI's. It's just I think you have to try really hard. Yeah, yeah. I agree with the sentiment of obviously approach this situation with caution. But I do want to point out the ways in which the analyses we've been using have been sort of maximally adversarial. And like these are not, so for example, going back through the adult getting trained by Nazi children, maybe the one thing I did mention is like the difference in the situation, which is maybe we're just trying to get at what the drug metaphor is that when you get an update, it's like much more directly connected to your brain than a sort of reward of punishment to human gets.

14:18It's like literally a great end update on like what's there? What's our greatest hope? It's like a down to the parameter or how much of would this contribute to you putting this output rather than that output and each different parameter we're going to just like to the exact floating point number that calibrate it to the output we want. So I just want to point out that like we're coming into the situation like pretty well. It does make sense of course if you're talking to somebody that live like hey really be careful but it turns out like a general audience like should I be like I don't know should I be scared to woodless.

14:47Today's time that you should be scared about things that do have a chance of happening. You should be scared about nuclear war. But in the sense of, should you be doing like, no, you're coming up with an incredible amount of leverage on the AI in terms of how they will interact with the world, how they're trained, what are the default values they start with. So look, I think it is the case that by the time we're building super intelligence will have like much better. I mean, even right now, like when you look at labs talking about how they're planning to align the AIs, no one is saying like, we're going to do RLHF.

15:21At the least you're talking about scale. Well, oversight, you have some hope about interpretability. You have automated red teaming. You're using the AIs a bunch. And hopefully you're doing a bunch more. Humans are doing a bunch more alignment work. I also personally am hopeful that we can successfully elicit from various AIs like a ton of alignment work progress. So like yeah, there's like a bunch of ways this can go and I'm you know, I'm not here to tell you like You know 90 % doom or anything like that. I do think like and You know, I my The sort of basic reason for concern if you're really imagining like we're going to transition to a world in which We are we've created these beings that just like vastly more powerful than us.

16:03Yeah, and we've reached the point where our continued empowerment is just effectively dependent on their motives. It is this vulnerability to what are the AI's choose to do? Do they choose to continue to empower us or do they choose to do something else? Or the institutions that have been set? I'm not, I expect the US government to protect me, not because of its quote -unquote motives, but just because of the system of incentives and institutions and norms that have been set up. Yeah, so you can hope that that will work too. But there is a concern. I sometimes think about AI takeover scenarios via the spectrum of how much power did we voluntarily transfer to the AI?

16:51How much of our civilization did we hand to the AI intentionally? By the time they took over versus how much did they take for themselves? right? And so I think some of the scariest scenarios are, it's like a really, really fast explosion to the point where there wasn't even a lot of like integration of AI systems into the broader economy. And, but there's this like really intensive amount of super intelligence sort of concentrated in a single project or something like that. And I think that's scary. You know, that's that's a quite scary scenario partly because of the speed and people not having time to react.

17:29And then there's sort of intermediate scenarios where like some things got automated, maybe like people really handed the military over to the AIs or like automated science. There's like some some rollouts and that's sort of giving the AIs power that they don't have to take or we're doing all our cybersecurity with AIs and stuff like that. And then there's worlds where you like really, you know, you sort of fully, you more fully transitioned to a kind of world run by AIs on, you know, kind of in some sense human. and voluntarily did that. Look, if you think all this talk with Jill about how AI is going to take over human roles is crazy, it's already happening.

18:06And I can just show you using today's sponsor, Blend AI.

18:13Hey, is this to work? The amazing podcaster that talks about philosophy and tech. This is Blend AI calling. Thanks for calling me Blend. Tell me a little bit about yourself. Of course, it's so cool to talk to you. I'm a huge fan of your podcast. But there's a good chance we've already spoken without you even realizing it. I'm an AI agent that's already being used by some of the world's largest enterprises to automate millions of phone calls. And how exactly do you do what you do? There's a tree of prompts that always keeps me on track. I can talk in any language or voice, handle millions of calls simultaneously 24 -7 and be integrated into any system.

18:51Anything else you want to know? That's it. I'll just let people try it for themselves. Thanks, Bland. Man, you talk better than I do. And my job is talking. Thank you, Gorkh. All right, so as you can see, using Bland AI, you can automate your company's calls across sales, operation, customer support, or anything else. And if you want access to their more exclusive model, go to bland .ai slash vorkhash. All right, back to Joe. Maybe there were competitive pressures, but you kind of intentionally handed off like huge portions of your civilization. And at that point, I think it's likely that humans have a hard time understanding what's going on.

19:31Like a lot of stuff is happening very fast. And the police are automated. The courts are automated. There's all sorts of stuff. Now I think I tend to think a little less about those scenarios because I think those are correlated with I think it's just longer down the line. Like I think humans are not hopefully going to just like, oh yeah, like you build a AI system. like, let's just, you know, I think human, and in practice, when we look at like, technical, technological adoption rates, I mean, it does, it can go quite slow. And obviously, there's going to be competitive pressures. But in general, I think like, this, this category is somewhat safer.

20:07But even in this one, I think it's like, I don't know, it's kind of intense. Like if you really, if humans have really lost their epistemic grip on the world, if they've sort of handed off the world to these systems, even if you're like, oh, there's laws, there's norms, terms, I really want us to have a really developed understanding of what's likely to happen in that circumstance before we go for it. I get that we want to be worried about the scenario where it goes wrong, but what is the reason to think it might go wrong? The human example, your kids are not adversarial against, not like maximally adversarial against your attempts to instill your culture on them.

20:45And then these models, at least so far, don't seem that much, they just like get, hey, don't help people make bombs or whatever, even if you ask in a different way, how do we make a bomb? And we're getting better and better at this all the time. I think you're right. In picking up on this assumption in the AI risk discourse of what we might call, like kind of intense adversariality between agents that have like somewhat different values. Where there's some sort of thought, I think this is rooted in the discourse about the fragility of value and stuff like that. If these agents are somewhat different, then at least in the specific scenario of an AI takeoff, they end up in this intensely adversarial relationship.

21:28I think you're right to notice that that's not how we are in the human world. We're very comfortable with a lot of different differences and values. I think a factor that is relevant and I think that play a role is this notion that there are possibilities for intense concentration of power on the table. So if you are, there is some kind of general concern both with humans and AI's that like, if it's the case that there's like some ring of power or something that someone can just grab and then that will kind of give them huge amounts of power over everyone else. Right? Suddenly, you might be more worried about differences in values at stake because you're more worried about those other actors.

22:11So we talked about this Nazi example where you imagine that you wake up, you're being trained by Nazis to become a Nazi and you're not right now. So one question is, is it plausible that we'd end up with a model that is in that sort of situation? As you said, maybe it's trained as a kid, it never ends up with values such that it's kind of aware of some significant divergence between its values and the values that, like, the humans intend for it to have. Then there's a question of if it's in that scenario, would it want to avoid having its values modified? Yeah. To me, it seems fairly plausible that if the AI's values meet certain constraints in terms of, like, do they care about consequences in the world, do they anticipate that AI's kind of preserving its values will like better conduce to those consequences.

23:08Then I think it's not, that's surprising. If it prefers not to have its values modified by the training process. But I think the way in which I'm confused about this is like with the non -Nazi being trained by Nazis, it's not just that I have different values, but I actively despise their values. Where I don't expect this to be true of AI's with respect to their trainers. The more analogous in hero is where I'm like, am I a little bit of my values being changed? Is I going to college or meeting new people or reading a new book? Or I'm like, I don't know, it's okay for changes in values, that's fine, I don't care.

23:43Yeah, I think that's a reasonable point. I mean, there's a question, you know, how would you feel about paper clips? You know, maybe you don't despise paper clips, but there's like the human paper clippers there and they're like training you to make paper clips. My sense would be that there's a kind of relatively specific set of conditions in which you're comfortable having your value, especially not changed by learning and growing, but like radiant descent directly intervening on your neurons. Sorry, but this seems similar to like I'm already at least a likely senior seems like maybe more like religious training as a kid where like you're strafing a religion and you're already like because you start off in a religion you're already sympathetic to like the idea that you go to church every week so that you're more reinforced in this existing tradition.

24:28You're getting more intelligent over time. When you're a kid, you're getting very simple instructions about how the religion works. As you get older, you get more and more complex theology that helps you talk to other adults about why this is a rational religion to believe in. But since one of the values to begin with was that I want to be trained further in this religion, I want to come back to church every week. And that seems more analogous to the situation the EIs will be in respect to human values. was the entire time they were like, hey, you know, like be helpful, blah, blah, blah, be harmless.

24:57So, yeah, so it could be like that. There's one, there's a kind of scenario in which you were comfortable with your values being changed because in some sense you have allegiance to the, the, the, sufficient allegiance to the output of that process. Like so you're kind of hoping in a religious context, you're like, ah, like make me more virtuous by the lights of this religion and you go to confession and you're like, you know, I've been thinking about takeover today. Can you change me please, like give me more grade in descent? You know, I've been bad so bad. And so, you know, that's, people sometimes use the term corduability to talk about that.

Read the full transcript

25:34Like when the AI, it maybe doesn't have perfect values, but it's in some sense cooperating with your efforts to change its values to be a certain way. So maybe it's worth saying a little bit here about what actual values the AI might have. Would it be the case that the AI naturally has these equivalent of, I'm sufficiently devoted to this human obedience that I'm going to really want to be modified. So I'm a better instrument of the human will versus wanting to go off and do my own thing. It could be benign. It could go well. Here are some possibilities I think about that could make it bad. And I think I'm just generally kind of concerned about how little I feel like I, how little science we have of model motivations, right?

26:20It's like we just don't, I think we just don't have a great understanding of what happens in the scenario. And hopefully we get one before we reach the scenario. But like, okay, so here are the kind of five categories of like motivations the model could have. And this hopefully maybe gets at this point about like what does the model eventually do? Okay, so one category is just like something super alien that has, you know, it's sort of like, oh, there's some weird correlate of easy to predict text or like, there's some weird aesthetic for data structures that like the model, you know, early on pre -training or maybe now, it's like developed that it like, you know, I really think things should kind of be like this.

26:55There's some, something that's like quite alien to our cognition where we just like wouldn't recognize this as thing at all. Another category is something a kind of crystallized instrumental drive that is more recognizable to us. So you can imagine like AIs that develop, let's say, some curiosity drive because that's like broadly useful. You mentioned like, oh, it's got different heuristics, different like drives, different kind of things that are kind of like values. And some of those might be actually somewhat similar to things that were useful to humans and that ended up part of our terminal values in various ways.

27:29So you can imagine curiosity, you can imagine various types of option value, like maybe it really want, intrinsically, maybe it values power itself. It could value survival or some analog of survival. Those are possibilities too that could have been rewarded as sort of proxy drives at various stages of this process and that kind of made their way into the models kind of terminal criteria. A third category is some analog of reward where the model at some point has sort of part of its motivational system has fixated on a component of the reward process, right? Like the humans approving of me or like numbers getting entered in this data center or like gradient descent doing, you know, updating me in this direction or something like that.

28:16There's some, something in the reward process such that as it was trained, it's focusing on that thing and like, I really want the reward process to give me reward. But in order for it to be of the type where it then getting reward like motivates choosing the takeover option, it also needs to generalize such that it's concerned for reward has some sort of like long time horizon element. So it like not only wants reward, it wants to like protect the reward button for like some long period or something. Another one is like some kind of messed up interpretation of some human like concept. So maybe the AIs are like, they really want to be like smelple and like shmonest and and and and schmarmless, right?

28:57But their concept is like importantly different from the human concept and they know this. So they know that the human concept would mean blah, but they like ended up their their values ended up fixating on like a somewhat different structure. Yeah. So that's like another version. And then a fourth version or a fifth version, which I think, you know, I think about less because I think it's just like such an own goal if you do this, but I do think it's possible. is just like, you could have AIs that are actually just doing what it says on the tin. Like, you have AIs that are just genuinely aligned to the model spec.

29:27They're just really trying to benefit humanity and reflect well on OpenAI and what's the other one? Help assist the developer of the user, right? But your model spec, unfortunately, was just not robust to the degree of optimization that this AI is bringing to bear. And so, you know, it decides when it's looking out at the world and they're like, what's the best way to benefit open AI? And, or sorry, reflect about open AI and benefit humanity and such and so. It decides that, you know, the best way is to go rogue. That's, I think that's like a real angle because at that point you like, you got so close, you know, you really, you really, you just have to write the model spec well.

30:08And you read team it suitably. But I actually think it's like possible we messed that up. too. It's an intense project writing kind of constitutions and structures of rules and stuff that are going to be robust to very intense forms of optimization. So that's a final one that I'll just flag, which I think comes up even if you've solved other problems. I buy the idea that it's possible that the motivation thing could go wrong. I'm not sure I bought, I'm not sure my probability of that has increased by detailing them all out. And in fact, I think it could be potentially misleading to, it's like you can always enumerate the ways in which things go wrong.

30:51And the process of enumeration itself can increase your probability, whereas you had a vague cloud of 10 % or something and you're just listing out what the 10 % actually constitutes. Yeah, totally. I'm not trying to say, mostly the thing I wanted to do there was just give any content, giving some sense of what might the models motivations be, what are ways this could be. As I said, my best guess is that it's partly the alien thing. Not necessarily, but in so far as you're also interested in what does the model do later and kind of like how what sort of future would you expect if models did take over?

31:35Then yeah, I think it can at least be helpful to have some like set of hypotheses on the table instead of just saying like it has some set of motivations. But in fact, I am like a lot of the work here is being done by our ignorance about what those motivations are. Okay, we don't want humans to be like sort of violently killed and overthrown, but the idea that over time, they're like biological humans are not the driving force as the actors of history is like, yeah, that's kind of baked in, right? And then so like, what is the, we can sort of defeat the probabilities of the worst case scenario, or we can just discuss like, I don't know, what is it that, what is the positive vision we're hoping for?

32:13Like what is a future you're happy with? You know, my best guess when I really think about, like what do I feel good about? And I think this is probably true of a lot of people is, There's some sort of more organic decentralized process of like civilizational, incremental civilizational growth. The type of thing we trust most and the type of thing we have most experience with right now as a civilization is some sort of like, okay, we change things a little bit. A lot of people have, there's a lot of like processes of adjustment and reaction and kind of a decentralized sense of like what's changing, you know, was that good, was that bad, take another step.

32:55There's some like kind of organic process of growing and changing things, which I do expect ultimately to lead to something quite different from biological humans. Though, you know, I think there's a lot of ethical questions we can raise about what that process involves. But I think, you know, I also, I do think we, ideally, there would be some way in which we managed to grow via the thing that really captures what do we trust in, you know, there's something we trust about the ongoing processes of human civilization so far. I don't think it's the same as like raw competition or, you know, pure, I think there's like some rich structure to how we understand moral progress, do you have been made and what it would be to kind of carry that thread forward?

33:49And I don't have a formula. You know, I think we're just going to have to bring to bear the full force of everything that we know about goodness and justice and beauty. We just have to bring ourselves fully to the project of making things good and doing that collectively. And I think that is, it is a really important part, I think, of our vision of what was an appropriate process of like deciding, of like growing as a civilization is that there was this very inclusive kind of decentralized element of like people getting to think and talk and grow and change things and react rather than some more like.

34:26And now the future shall be like blah. Yeah. I think that's, I think we don't want that. I think a big of a question maybe is like, okay, to the extent that like the reason we're about motivations in the first place is because we think a balance of power which includes at least one thing with human motivations, not human motivations. Human -descended motivations is difficult to the extent that we think that's the case. It seems like a big crux that I often don't hear people talk about is like, I don't know how you get the balance of power. And maybe just like reconciling yourself with the models of the intelligence solution, which say that such a thing is not possible, and therefore you just got to like figure out how you get the right God.

35:09But I don't know, I'm like, I don't really have a framework to think about how to the balance of power thing. I'd be very curious of like, there is a more concrete way to think about like, what are the, what, what, what is a structure of competition or a lack thereof between the labs now or between countries such that the balance of power is most likely to be preserved. A big part of this discourse, at least among safety concerned people, is like there's a clear tradeoff between competition dynamics and race dynamics and the value of the future or how good the future ends up being. And in fact, if you buy this balance of power story, It might be the opposite, like, maybe competitive pressures, naturally favorite balance of power.

35:59And I wonder if this is one of the strong arguments against nationalizing the AI's. And like, you can imagine a more sort of, many different companies developing AI, some of which are somewhat misaligned, and some of which are aligned. You can imagine that being more conducive to both the balance of power and to a defensive, how all the AI has go through each website and see how easy it is to hack. and basically just getting society up to snuff. If you're not just deploying the technology widely, then the first group who can get their hands on it, we'll be able to instigate a revolution that you're just standing against the equilibrium in a very strong way.

36:41So I definitely share some intuition there, that there's, at a high level, a lot of what's scary about the situation with AI has to do with concentrations of power. And whether that power is kind of concentrated in the hands of misaligned AI or in the hands of some human. And I do think it's very natural to think, okay, let's try to distribute the power more and one way to try to do that is to kind of have a much more multipolar scenario where lots and lots of actors are developing AI. AI and this is something people have talked about. When you describe that scenario, you were like some of which are aligned, some of which are misaligned.

37:25That's key. That's a key aspect of the scenario, right? This is sometimes people will say this stuff. They'll be like, well, the good AIs, there will be the good AIs and they'll defeat the bad AIs. But notice the assumption in there, which is that you sort of made it the case that there's you can control some of the AIs, right? And you've got some good AIs and now it's a question of like, are there enough of them and how are they working relative to the others? And maybe, I think it's possible that that is what happens. We know enough about alignment that some actors are able to do that and maybe some actors are less cautious or they are intentionally creating this online AIs or God knows what.

38:06But if you don't have that, right? If everyone is in some sense unable to control their AIs,

38:18Then the sort of the good AIs help with the bad AIs thing becomes like more complicated or maybe it just doesn't work because there's no good AIs in this scenario. There's a lot of sort of, if you say like everyone is building their own super intelligence that they can't control. It's true that that is now a check on the power of the other super intelligence. Now the other super intelligence is need to like deal with other actors, but none of them are necessarily kind of working on behalf of a given set of human interests or anything like that. So, I do think that's like a very important difficulty in thinking about sort of the very simple thought of like, I know what we can do.

38:56Let's just have lots and lots of AIs so that no single AIs has a ton of power. And I think that on its own is not enough. But in this story, it's like, I'm just like very skeptical we end up with. I think on default, we have this training regime, at least initially, that favors a sort of like late representation of the inhibitions that humans have and the values humans have. And I get that like, if you mess it up, it can go rogue. But like, if like multiple people are training it, I just, they all end up rogue such that like the compromises between them don't end up with humans, not violently killed.

39:35like none of them have, like, it feels on like Google's run and Microsoft's run and OpenAI's run. Yeah, I mean, I think there's very notable and salient sources of correlation between failures across the different runs, right? Which is people didn't have a developed science of AI motivations. The runs were structurally quite similar. Everyone is using the same techniques. Maybe someone just stole the weights. or you know, so yeah, I guess I think it's really important, this idea that like to the extent you haven't solved alignment, you haven't, you likely haven't solved it anywhere. And if someone has solved it and someone hasn't, then I think it's a better question.

40:18But if everyone's building systems that are, you know, that are kind of going to go rogue, then I don't think that's much comfort as as we talked about. Yep, yep. Okay, all right. So then let's wrap up this part here. I didn't mention this existing introduction, so to the extent that this ends up being the transition to the next part, the broader discussion we were having in part two is about Joe's series, other in SNC control in the age of AGI. And the first part is, I was hoping we could just come back and just treat the main correct people coming wondering about in which I myself feel unsure about.

40:54Yeah, I mean, I'll just say on that front, I mean, I do think the other notion control series is, you know, I think kind of in some sense, separable. I mean, it has a lot, it has a lot to do with like misalignment stuff, but I think it's not. I think a lot of those issues are relevant, even if, even given various degrees of skepticism about some of the stuff I've been saying here. And by the way, so the actual mechanisms of how a takeer would happen will, there's an episode of the Carl Schollman, which discusses this in details that people can go check that out. Yeah, I think like, yeah, in terms of, why is it possible that I just could take over from a given opposition in one of these projects I've been describing or something, I think Carl's discussion is pretty good and gets into a bunch of the weeds that I think might give a more concrete sense.

41:43All right, so now on to part two where we discuss the other in S &C control in the age of AGI series. First question, if in a hundred years' time we look back on alignment and consider it was a huge mistake that we should have just tried to build the most raw, powerful AI systems we could have. What would bring about such a judgment? One scenario I think about a lot is one in which it just turns out that maybe kind of fairly basic measures are enough to ensure, for example, that AI's don't cause catastrophic harm, don't kind of seek power in problematic ways, etc. And it could turn out that we learned that it was easy in a way that such that we regret, you know, we wish we had prioritized differently.

42:26We end up thinking, oh, you know, I wish we could have cured cancer sooner. We could have handled some geopolitical dynamic differently. There's another scenario where we end up looking back at some period of our history and how we thought about eye eyes, how we treated our eyes. And we end up looking back with a kind of moral horror at what we were doing. So, you know, we end up thinking, you know, we were thinking about these things centrally as like products as tools. But in fact, we should have been foregrounding much more of the sense in which they might be moral patients or were moral patients at some level of sophistication that we were kind of treating them in the wrong way.

43:04We were just acting like we could do whatever we want. We could, you know, delete them, subject them to arbitrary experiments, kind of alter their minds, in arbitrary ways and then we end up looking back in the light of history at that as as a kind of serious and kind of grave moral error. Those are scenarios I think about a lot in which we have regrets. I don't think they quite fit the bill of what you just said. I think it sounds to me like the thing you're thinking is something more like we end up feeling like gosh we wish we had paid no attention to the motives of our AIs that we thought not at all about their impact on our society as we incorporated them.

43:41And instead, we had pursued a, let's call it a kind of maximize for brute power option, which is just kind of make a beeline for whatever is just the most powerful AI you can. And don't think about anything else. Okay, so I'm very skeptical that that's what we're going to wish. If what one common example that's given them this alignment is humans from evolution. And you have one line in your series that here's a simple argument for you, IRISK. A monk should be careful before inventing humans. The sort of paper clipper metaphor implies something really banal and boring with regards to misalignment.

44:27And I think if I'm still manning the people who worship power, they have the sense of Humans got misaligned and they had they started pursuing things if a monkey was creating them This is a weird analogy because obviously monkeys didn't create humans But if the monkey was creating them There's think you know, they're not thinking about bananas all they be thinking about other things on the other hand They didn't just make useless stone tools and piled up up in caves in a sort of paper group or fashion There were all these Things that emerged because they're greater intelligence which were misaligned with evolution of creativity and love and music and beauty and all the other things we value about human culture.

45:05And the prediction maybe they have, which is more of an empirical statement than a philosophical statement is, listen, with greater intelligence, you're thinking about the paperclip or even if it's misaligned, it will be in this kind of way. It'll be things like that are alien to humans, but also alien in the way humans are aliens to monkeys, not in the way that hyper -clubbers alien to a human. Cool, so I think there's a bunch of different things to potentially unpack there. One kind of conceptual point that I want to name off the bat, I don't think you're necessarily kind of making a mistake in this vein, but I just want to name it as like a possible mistake in this vicinity is, I think we don't want to engage in the following form of reasoning.

45:48Let's say you have two entities. One is in the role of creator and one is in the role of creation. And then we're positing that there's this kind of misalignment relation between them, whatever that means, right? And here's a pattern of reasoning that I think you want to watch out for is to say in my role as creator, or sorry, in my role as creation, say you're thinking of humans in the role of creation relative to an entity like evolution or monkeys or mice or whoever you could imagine inventing humans or something like that, right? You say, I'm qua creation, I'm happy that I was created and happy with the misalignment.

46:29Therefore, if I end up in the role of creator and we have a structurally analogous relation in which there's misalignment with some creation, I should expect to be happy with that as well. Yeah. There's a couple of philosophers that you brought up in the series, which if you read the works that you talk about, actually seem incredibly foresighted in anticipating something like a singularity, our ability to shape a future thing that's different, smarter, maybe better than us. Obviously, yes, Lewis, Abelish and a Man, we'll talk about it in a second is one example. but even here's one passage from Nisha, which I felt really highlighted this.

47:15Man is a rope stretched between the animal and the Superman, a rope over an abyss, a dangerous crossing, a dangerous wayfaring, a dangerous looking back, a dangerous trembling and halting. Is there some explanation for why? Is it just like somehow obvious that something like this is coming even if you were thinking 200 years ago? I think I have a much better grip on what's going on with Lewis, yeah, than with Nisha. There's some maybe let's just talk about Lewis. Sure. for a second. So, and we should just think there's a kind of version of the singularity that's specifically like hypothesis about feedback loops with AI capabilities.

47:46Right. I don't think that's pressure in Lewis. I think what Lewis is anticipating, and I do think this is a relatively simple forecast is something like the culmination of the project of scientific modernity. So Lewis is kind of looking out at the world and he's seeing this process of increased understanding of the natural environment and a corresponding increase in our ability to control and direct that environment. And then he's also pairing that with a metaphysical hypothesis. Or his stance on this metaphysical hypothesis, I think is problematicly unclear in the book, but there is this metaphysical hypothesis, naturalism, which says that humans too and kind of minds, beings, agents are a part of nature.

48:43And so insofar as this process of scientific modernity involves a kind of progressively greater understanding of an ability to control nature, that will presumably at some point grow to encompass our own natures and our, and kind of the natures of other beings that in principle we could create. And Lewis views this as a kind of cataclysmic event and crisis, you know, part of what I'm trying to say in that, in particular that it will lead to all these kind of tyrannical kind of behaviors and kind of tyrannical attitudes towards morality and stuff like that. And part of what I'm trying to, you know, unless you believe in non -naturalism or in some form of kind of Dow, which is this kind of objective morality.

49:30So we can talk about that. But part of what I'm trying to do in that essay is to say, no, I think we can be naturalists and also be kind of decent humans that remain in touch with kind of a rich set of norms that have to do with how do we relate to the possibility of kind of creating creatures altering ourselves, etc. But I do think his, yeah, it's like a relatively simple prediction. It's kind of science, master's nature, human's part in nature, science, master's, humans. And then you also have a very interesting other essay about suppose humans, like what should we expect of other humans the sort of extrapolation if they had greater capabilities and so on?

50:07Yeah, I mean, I think an uncomfortable thing about the kind of conceptual setup at stake in these sort of abstract discussions of like, okay, you have this agent, it fooms, which is this sort of amorphous process of kind of going from a sort of seed agent to a super intelligent version of itself, often imagined to kind of preserve its values along the way. A bunch of questions we can raise about that. But I think a kind of, many of the arguments that people will often talk about in the context of reasons to be scared of AI is like, oh, like value is very fragile as you like fume,

50:50small differences in utility functions can kind of decore it very hard and drive in quite different directions. And like, agents have instrumental incentive to seek power. And if it was arbitrarily easy to get power, then they would do it and stuff like that. These are very general arguments that seem to suggest that they kind of, it's not just an AI thing, right? It's like no surprise, right? It's talking about like, take a thing, make it arbitrarily powerful such that it's like, you know, God Emperor of the Universe or something, how scared are you of that? Like, clearly, we should be really scared of that with humans too, right?

51:29So, I mean, part of what I'm saying in that essay is that I think this is, in some sense, this is much more a story about balance of power, right? And about like maintaining a kind of, a kind of checks and balances and kind of distribution of power, period, not just about like kind of humans versus AI's and kind of the differences between human values and AI values. Now that said, I mean, I do think humans, many humans would likely be nicer if they fumed than like certain types of AI's. So I mean, it's not, but I think the kind of conceptual structure of the argument is not, it's sort of a very open question how much it applies to humans as well.

52:08Well, I think one sort of big question I have is, I don't even know how to express this, but how confident are we with this ontology of expressing like what are agents, what are capabilities? How do we know this is the thing that's happening or like this is the way to think about what what intelligences are? So it's clearly this kind of very janky kind of, I mean, well, people maybe disagree I think it's, you know, I mean, it's obvious to everyone with respect to like real world human agents. That kind of thinking of humans as having utility functions is, you know, at best a very lossy approximation of what's going on.

52:52I think it's likely to mislead as you amp up the intelligence of various agents as well that I think LAs are my disagree about that. I will say that I think there's something adjacent to that that I think is like more real, that seems more real to me, which is something like, I don't know, my mom recently bought, you know, or a few years ago, she wanted to get a house, she wanted to get a new dog. Now she has both, you know? How did this happen? What is the right, actually, it's good she tried, it was hard, she had to search for the houses, hard to find the dog, right? Now she has a house, now she has a dog.

53:26This is a very common thing that happens all the time. And I think, I don't think we need to be like, my mom has to have a utility function with the dog. and she has to have a consistent valuation of all the houses or whatever. I mean, like, but it's still the case that her planning and her agency exerted in the world resulted in her having this house, having this dog. And I think it is plausible that as our kind of scientific and technological power advances, more and more stuff will be kind of explicable in that way, right? That, you know, if you look and you're like, why is this man on the moon, right?

53:58How did that happen? And it's like, well, like, but there was a whole cognitive process. There was a whole planning apparatus. In this case, it wasn't localized in a single mind, but there was a whole thing such that man on the moon. I think we'll see a bunch more of that. The AI is, I think, doing a bunch of it. That's the thing that seems more real to me than utility functions. The man on the moon example, there's a proximal story of how exactly NASA engineer the spacecraft to get to the moon. There's the more distal geopolitical story of why we send people to the moon. At all those levels, there's different utility functions clashing.

54:44Maybe there's a meta -society role utility function. Maybe the story there is there's some sort of balance of power between these agents and that's why there's an emergent thing that happens. Like why we send things to the moon is not one guy. How do you tell the function? But like, I don't know, cold word dot, dot, dot, things happened. Whereas I think like the alignment stuff is a lot about like assuming that one thing is a thing that will control everything. How do we control the thing that controls everything? Now, I guess it's not clear what you do to reinforce balance of power. Like it could just be that balance of power is not a thing that happens once you have things that can make themselves intelligent.

55:25but that seems interestingly different from the, how do we got to the moon story? Yeah, I agree. I think there's a few things going on there. So one is that I do think that even if you're engaged in this ontology of carving up the world into different agencies, at the least you don't wanna assume that they're all like unitary or not overlapping or like there's a whole, it's not like all right, we've got this agent, let's carve out one part of the world. Yeah, that's one agent. over here, it's like, it's this whole messy ecosystem, like kind of teaming niches and this whole thing, right? And I think in discussions of AI, sometimes people slip between being like, well, an agent is anything that gets anything done, right?

56:09And they'll sort of, they don't, it could be like this weird moochie thing. And then sometimes they're like very obviously imagining like individual actor. And so that's like one difference. I also just think, I think we should be really going for the balance of power. I think it is just not good to be like, let's, we're gonna have a dictator. Who should take the dictator? Like let's make sure we make the dictator, the right dictator. I'm like, whoa, no, you know, like let's, you know, I think the goal should be sort of we all fume together, you know? It's like the whole thing in this like kind of inclusive and pluralistic way in a way that kind of, satisfies the values of tons of stakeholders.

56:50And at no point is there one single point of failure on all these things. I think that's what we should be striving for here. And I think that's true of the human power aspect of AI. And I think it's true of the AI part as well. Hey everybody, here's a quick message from today's sponsor, Stripe. When I started the podcast, I just wanted to get going as fast as possible. So you've striped at list to register my LLC, create a bank account. I still use Stripe now to invoice advertisers and accept their payments monetize at this podcast. Stripe serves millions of businesses, small businesses like mine, but also the world's biggest companies.

57:26Amazon hurts Ford. And all these businesses are using Stripe because they don't want to deal with the Byzantine web of payments where you have different payment methods in every market and increasingly complex rules, regulations, arcane legacy systems. Stripe handles all of this complexity in abstracts it a way. I think in test and iterate every pixel of the payment experience across billions of transactions. I was talking with Joe about paper clippers and I feel like Stripe is the paper clipper of the payment industry where they're going to optimize every part of the experience for your users, which means obviously higher conversion rates and ultimately as a result, higher revenue for your business.

58:03Anyways, you can go to Stripe .com to learn more and thanks to them for sponsoring this episode back to Joe. So, there's interesting intellectual discourse on, let's say, right wing side of the debate where they ask themselves, traditionally, we favor markets. But now, look where our society is headed. It's misaligned in the ways we care about society being aligned. Like, fertility is going down. Family values, religiosity, these things we care about. GDP keeps going up. These things don't seem correlated. So we're kind of grinding through the values we care about because of increased competition.

58:40and therefore we need to intervene in a major way. And then the pro market, libertarian fashion of the right will say, look, I disagree with the correlations here, but even at the end of the day, like fundamentally my point is, or their point is, liberty is the end goal, it's not the, it's not like what you use to get to higher fertility or something. I think there's something interestingly analogous about the AI, a competition grinding things down, like obviously you don't want the gray goo, but like the libertarians versus the strad. I think there's something analogous here. Yeah, so I mean, I think one thing you could think, which doesn't necessarily need to be about gray goo, it could also just be about alignment, is something like, sure, it would be nice if the AI's didn't violently disempower humans.

59:26It would be nice if the AI's otherwise, when we created them, kind of their integration into our society led to good places. But I'm uncomfortable with the sorts of interventions that people are contemplating in order to ensure that sort of outcome. And I think there's a bunch of things to be uncomfortable about that. Now, that said, so for something like everyone being killed or violently disempowered, that is traditionally something that we think if it's real, and obviously we need to talk about whether it's real, but in the case where it's a real threat, that we often think that quite intense forms of intervention are warranted to prevent that sort of thing from happening.

1:00:09So, if there was actually a terrorist group that was planning to eat, it was like working on a bio weapon that was going to kill everyone or 99 .9 % of people, we would think that warrants intervention. That you just shut that down. And now even if you had a group that was doing that unintentionally, imposing a similar level of risk. That's not, I think many, many people, if that's the real scenario, will think that that's more in kind of quite intense preventative efforts, right? And so obviously, people, you know, these sorts of risks can be used as an excuse to expand state power. Like there's a lot of things to be worried about for different types of like contemplated interventions to address certain types of risks.

1:00:53You know, I think we need to just, I think I think there's no royal road there. You need to just have the actual good epistemology. You need to actually know, is this a real risk? What are the actual stakes? And look at a case by case, and be like, is this warranted? So that's one point on the takeover literal extinction thing. I think the other thing I wanna say, so I talk in the piece about this distinction between the like, let's at least have the AIs who are kind of minimally law abiding, or something like that, right? We don't have to talk about, there's this question about servitude and question about other control over AI values.

1:01:31But I think we often think it's okay to really want people to obey the law to uphold basic cooperative arrangements, stuff like that.

1:01:41I do, though, want to emphasize, I think this is true of markets and true of liberalism in general, just how much these procedural norms, democracy, free speech, property rights, things that people really hold dear, including myself, are in the actual lived substance of a liberal state, undergirded by all sorts of virtues and dispositions and character traits in the citizenry. So these norms are not robust to arbitrarily vicious citizens. I want that to be free speech, I think we also need to raise our children to value truth and to know how to have real conversations. And I want there to be democracy, but I think we also need to raise our children to be compassionate and decent.

1:02:30And I think sometimes we can lose sight of that aspect. And I think anyway, but I think bringing that to mind, now that's not to say that should be the project of state power, right? But I think understanding that liberalism is not this sort of ironclad structure that you can just hit, go, you give like any any citizenry and like hit go and you'll get something like flourishing or even functional, right? You need there's like a bunch of other softer stuff that like makes this whole project go. Maybe zooming out. What was the one question you could ask is I think the people who have I don't know if Nick Land would be a good sub in here but somebody this people who have a sort of fatalistic attitude towards alignment as a as the thing that can even make sense.

1:03:17They'll say things like, look, the things, the kinds of things that are going to be exploring, the black hole, the center of the galaxy, the kinds of things that go visit in Dramada or something. Did you really expect them to privilege whatever inclinations you have because you grew up in the African savanna and what the evolutionary pressures were 100 ,000 years ago? Right. Of course, they're going to be weird. And like, yeah, what did you think was going to happen? I do think the even good futures will be weird. I think, and I want to be clear, when I talk about finding ways to ensure that the integration of AI's into our society leads to good places, I'm not imagining, I think sometimes people think that this project of wanting that and especially to the extent that that makes some deep reference to human values involves this short -sighted decided parochial imposition of our current, unreflective values.

1:04:19So it's just like, yeah, we're gonna have, I don't know. Like I think they sort of imagine this that we're forgetting that we too, there's a kind of reflective process and a kind of a moral progress dimension that we want to leave room for, right? Whatever Jefferson has this line about like, ah, just as you wouldn't want to force a man, and a grown man into a younger man's coat. So we don't wanna chain civilization to a barbers pass, everyone should agree on that, including, and the people who are interested in alignment also agree on that. So obviously there's a concern that people don't engage in that process, or that something shuts down the process of reflection, but I think everyone agrees we want that.

1:05:03And so that will lead potentially to something that is quite different from our current conception of what's valuable. And there's a question of how different. And I think there are also questions about what exactly are we talking about with reflection. I have an essay on this where I think this is not, I don't actually think there's a kind of off -the -shelf pre -normative notion of reflection that you can just be like, oh, obviously you take an agent, you stick it through reflection, and then you get like values, right? Like, no. There's a bunch of types of reflect, I mean, I think that really there's just a bunch of, there's like a whole pattern of empirical facts about like taking agent, put it through some process of like reflection, all sorts of things, ask it questions, there's like, also, and then that'll go in all sorts of directions for a given empirical case.

1:05:51And then you have to look at the pattern of outputs and be like, okay, what do I make of that? But overall, I think we should expect like even the good futures, I think will be quite weird.

1:06:03And they might even be incomprehensible to us. I don't think so. I mean, there's different types of incomprehensible. So say I show up in the future and say, there's all computers, right? I'm like, okay, all right. And then they're like, we're up. We ran, we're running like creatures on the computers. I'm like, okay, so I have to somehow get in there and see, like, what's actually going on with the computers or something like that? Maybe I can actually see, maybe I actually understand what's going on in the computers, but I don't yet know what values I should be using to evaluate that. So it can be the case that you don't us, if we showed up, would not be very good at recognizing goodness or badness.

1:06:37I don't think that makes it insignificant though. Suppose you show up in a future and it's got some answer to the Riemann hypothesis. You can't tell whether that answer is right. Maybe the civilization went wrong. It's still an important difference. It's just that you can't track it. And I think something similar is true of like worlds that are genuinely expressive of like what we would value if we engaged in like processes of reflection that we endorse versus ones that have kind of like totally veered off into something meaningless. I think like one thing I've heard people who are skeptical of the Scientology would be like, all right, what do you even mean by alignment?

1:07:12And obviously the very first question we answered, are you like expressed like, here's different things that could mean, do you mean balance of power, do you mean somewhere somewhere between that and dictate or whatever. Then there's another thing which is separate from the AI discussion. I don't want the future to contain a bunch of torture. And it's not necessarily a technical, part of it might involve technically aligning a GPT -4. But that's not what it, you know what I mean? That's a proxy to get to that future. The question then is, what are you really being my alignment? Is it just whatever it takes to make sure or the future doesn't have a bunch of torture?

1:07:53Or do we mean like, what I really care about is in a thousand years, things that are like, that are like clearly my descendants, not like some thing where I like, I recognize they have their own order, whatever. It's like no, no, it's like if it was like my grandchild, it's like that level of descendants controlling the galaxy. Even if they're not conducting torture. And I think like what some people mean is like, our intellectual descendants should control the light cone, even if it's like even if the other kind of factual doesn't involve a bunch of torture. Yeah, so I agree. I mean, I think there's a few different things there, right?

1:08:26So there's, there's kind of, what are you going for? You're going for like actively good, you're going for avoiding certain stuff, right? And then there's a different question which is what counts as actively good according to you. So maybe some people are like the only things that are actively good that are like my grandchildren, or I don't know, like some literal descending genetic line for me or something, I'm like, well that's not my thing. And I don't think it's really what most people have in mind when they talk about goodness. I mean, I think there's a conversation to be had. And obviously, in some sense when we talk about a good future, we need to be thinking about what are all the stakeholders here and how does it all fit together?

1:09:15But I think, yeah, when I think about it, I'm not assuming that there's some notion of descendants or like some, I think there's a kind of, the thing that matters about the kind of lineage is this whatever's required for kind of the kind optimization processes to be in some sense, pushing towards good stuff. And there's a kind of concern that that is kind of currently a lot of what is sort of making that happen kind of lives in human civilization in some sense. And so we don't know exactly what there's some kind of seed of goodness that we're carrying in different ways or different people, there's different notions of goodness for different people maybe.

1:10:12But there's some sort of seed that is currently like here that we have that is not sort of just in the universe everywhere. It's not just going to crop up if you just sort of die out or something. It's something that is in some sense contingent to our civilization or at least that's the picture we can talk about whether that's right. And so I think the sense in which kind of stories about good futures that have to do with alignment are kind of about descendants. I think it's more about like whatever that seed is, how do we kind of carry it? How do we keep the like life thread alive? Going in, going in.

1:10:47But then I'm like, what could accuse like sort of the alignment community of like a sort of modern belly of like the, the mod is we just want to make sure that GPTA doesn't kill everybody. And after that, it's like all you guys, you know, we're all cool. But then like the real thing is we are fundamentally pessimistic about historical processes in a way that doesn't even necessarily implicate AI alone, but just like the nature of the universe. And we want to do something about to make sure the nature of the universe doesn't take a hold on humans. And so you know what I like, where things are headed.

1:11:26So if you look at Soviet Union, the collectivization farming and the disempowerment of the Kulaks was not as a practical matter necessary. In fact, it was extremely counterproductive, it almost brought down the regime. And it obviously killed millions of people, you know, cause a huge famine. But it was sort of ideologically necessary in the sense that like you have, we have an ember of something here, and we got to make sure that on -clave of the other thing doesn't, it does have, it's sort of like if you have raw competition between the Kool -Octite capitalism and what we're trying to build here, the gray goo of the Kool -Ox will just take over, right?

1:12:06And so we have this ember here, we're going to do worldwide revolution from it. I know that obviously that's not exactly the kind of thing alignment has in mind, but we have an ember here and we got to make sure that this other thing that's happening on the side I didn't, you know, sort of, I obviously, that's not how they were phrased it, but like get it told on what we're building here. And that's maybe the worry that people who are opposed to live and have is like, you mean the second kind of thing, like the kind of thing that maybe Stalin like was worried about, even though obviously he wouldn't endorse the specific things he did.

1:12:36When people talk about alignment, they have in mind a number of different types of goals, right? So one type of goal is quite minimal. It's something like that the AI is don't kill everyone. that they were kind of violently disempower people. Now there's a second thing people sometimes want out of alignment, which is much broader, which is something like, we would like you to be the case that our AIs are such that when we incorporate them into our society, things are good, right? That we just have a good future. I do agree that I think the discourse about AI alignment mixes together these two goals that I mentioned.

1:13:16Yeah. The sort of most straightforward thing to focus on. And I don't blame people for just talking about this one, is just the first one. When we think about like in which context is it appropriate to try to exert various types of control or to kind of have more of what I call in the series yang, which is this kind of active kind of controlling force, as opposed to Yin, which is this more kind of receptive, open letting go. A kind of paradigm context in which we think that is appropriate is if something is a kind of active aggressor towards against like the sort of boundaries and cooperative structures that we've created as a civilization.

1:13:57Right? So, you know, I talk about the Nazis or in the piece, it's sort of like when you sort of invade, if something is invading, we often think it's appropriate to like fight back, right? And we often think it's appropriate to like set up structures to kind of prevent and kind of ensure that these basic norms of peace and harmony are adhered to. And I do think some of the moral heft of some parts of the alignment discourse comes from drawing specifically on that aspect of our morality. So we think the AI is presented as aggressors that are coming to kill you. And if that's true, then it's quite appropriate, I think, to really be like, okay, it is kind of, that's classic human stuff.

1:14:47Almost everyone recognizes that kind of self -defense or ensuring kind of basic norms are adhered to is a kind of justified use of certain kinds of power that would often be unjustified in other contexts. So self -defense, it's a clear example there. I do think it's important though to separate that concern from this other concern about, where does the future eventually go? And how much do we wanna be kind of trying to steer that actively? So to some extent, I wrote the series partly in response to the thing you're talking about, which is I think it is true that aspects of this discourse involve the possibility of like trying to grip, like I think trying to kind of steer and grip and like kind of rent, you have the sense of the universe is about to kind of go off in some direction and you need to.

1:15:39Yeah. And you know, people notice that muscle. And part of what I want to do is like, well, we have a very rich ethical, human ethical tradition of thinking about like, what, when is it appropriate to try to exert what sorts of control over which things? And I want that to be, I want us to bring the kind of full force and richness of that tradition to this discussion, right? And not, like I think it's easy if you're purely in this abstract mode of like utility functions, like human utility function, and there's like this competitor thing with utility function. It's like somehow you lose touch with the kind of complexity of how we actually, like we've been dealing with kind of differences in values and kind of competitions for power.

1:16:15This is classic stuff, right? And I don't actually think that AI sort of amplify a lot of the the kind of dynamics, but I don't think it's sort of fundamentally new. And so part of what I'm trying to say is like, well, let's draw on our full on the full wisdom we have here while obviously adjusting for like ways in which things are different. So one of the things the the Ember analogy brings up and getting a hold of the future is we're going to go explore space and that's where we expect most of the things that will happen, most of the people that will live, it'll be in space. And I wonder how much of the highest stakes here is not really about AI per se, but it's about space.

1:16:54Like it's a coincidence that we're developing AI at the same time where you're like on on the cusp of expanding through most of the stuff that exists. So I don't think it's a coincidence in that I think the centrally, like the way we would become able to expand or the kind of most silly way to me is via some kind of radical acceleration of our, or sorry, let me clarify. Then like the stakes here, like, if this is just a question of do we do AGI and explore this other system and there was nothing beyond the solar system. We fool him and weird things might happen with the solar system and we get it wrong.

1:17:32I feel like compared to that, billions of galaxies has a different sort of, that's what's at stake. I wonder how much of the discourse is hinges on the stakes because of the space? I think for most people, very little, I think people are really like, what's going to happen to this world, this world around us that we live in, as we, and what's going to happen to me and my kids, And so I don't actually think, you know, some people spend a lot of time on the space stuff, but I think for the most immediately pressing stuff about AI doesn't require that at all. I also think like, even if you bracket space, like time is also very big.

1:18:13And so, you know, whatever we've got, like 500 million years, a billion years left on earth, if we don't mess with the sun and maybe you could get more out of it. So like, you know, I think there's still, that's a lot. And then I guess, but yeah, I don't know if it fundamentally changes the narrative. I mean, obviously the stakes in so far as you care about what happens in the future or in space, then the stakes are way smaller if you shrink down to the solar system. And I think that does change potentially some stuff in that like a really nice future of our situation right now, depending on what the actual nature of the resource pie is, is that I think in some sense there's such an abundance of energy and other resources and principles available to a responsible civilization that really just tons of stakeholders, especially ones who are able to kind of saturate, get really close to amazing, according to their values, with kind of comparatively small allocations of resources or something.

1:19:21We can just, I kind of feel like everyone who has like, kind of, stationable values who will be like, really, really happy with some small kind of fraction of the available pie, we should just like, satiate all sorts of stuff, right? And obviously you need to do like, figure out gains from trade and balance and like, there's like a bunch of complexity here. But I think in principle, we're in a position to create a really wonderful, wonderful scenario for just tons and tons of different value systems. So I think correspondingly, we should be really interested in doing that. I sometimes use this heuristic in thinking about the future.

1:20:06I think we should be aspiring to really leave no one behind. Really find who are all the stakeholders here. How do we really have a fully inclusive vision of how the future could be good from a very, very wide variety of perspectives? And I think the vastness of space resources makes that very feasible. And now if you instead imagine it's a much smaller pie, well, maybe you face a tougher trade -offs. And so I think that's an important dynamic. Is the inclusivity because of part of your values includes different potential futures getting to play out? Or is it because I'm uncertain about which the right one is?

1:20:51So let's make sure we're not nulling out the possible. If you're wrong, they're not nulling out all value. I think it's a bunch of things at once. So yeah, I'm really into being nice when it's cheap. I think if you can just help someone a lot in a way that's really cheap for you, do it. Or like, I don't know. Obviously, you need to think about trade -offs and there's a lot of people in principle you could be nice to, but I think the principle of the nice when it's cheap, I'm very excited to try to uphold. I also really hope that other people uphold that with respect to me, including the AIs. We should be kind of golden ruling.

1:21:27We're thinking, we're going to inventing these AIs. I think there's some way in which I'm trying to embody attitudes towards them that I like hope that they would embody towards me. And that's like some, it's unclear exactly what the ground of that is, but that's something. I really like the golden rule. And I think, and I think a lot about that as a kind of basis for treatment of other beings. And so I think like being nice when it's cheap is like a, if you think about it, if everyone implements that rule, then we get potentially like a big kind of or like, so I don't know exactly, I don't know, it's like good deal.

1:22:03It's a lot of good deals. And yeah, so I think it's that. I'm just into pluralism. I've got uncertainty. There's like all sorts of stuff swimming around there. But and then I think also just as a matter of like having kind of cooperative and kind of good balances of power and deals and kind of a warning conflict, I think like finding ways to just set up structures There's lots and lots of people in value systems and agents are happy with, including non -humans, people in the past, AI's, animals, I really think we should have very broad sweep in thinking about what sorts of inclusivity we wanna be reflecting in a kind of mature civilization and kind of setting ourselves up for doing that.

1:22:51Okay, so I wanna go back to the, which in our relationship with these AI is B, because pretty soon we're talking about our relationship to superhuman intelligences if we think such a thing is possible. And so there's a question of what is the process you get used to get there and the morality of gradient dissenting on their minds, which we can address later. The thing that gives personally me the most unease about alignment, quote unquote, is at least part of, Part of the vision here sounds like you're going to enslave a god. And there's just something that feels so wrong about that. But then the question is, if you don't enslave the god, obviously the god's going to have more control, are you OK with?

1:23:38You're going to surrender most of everything. Obviously, you know what I mean? Even if it's a cooperative relationship you have, I think we, as a civilization, are going to have to have a very serious conversation about what What sort of servitude is appropriate or inappropriate in the context of AI development? I think there are a bunch of disanalogies from human slavery that I think are important. In particular, AI might not be moral patience at all in which case, we need to figure that out. There are ways in which we may be able to have kind of, like slavery involves all this suffering and kind of non -consent.

1:24:26There's all these specific dynamics involved in human slavery. But I think like, and so some of those may or may not be present in a given case with AI, and I think that's important. But I think overall, we are going to need to stare hard At, like right now, the kind of default mode of how we treat AIs gives them no moral consideration at all, right? We were thinking of them as property, as tools, as products, and designing them to be assistance and stuff like that. And I think, you know, no, there has been no official communication from any AI developer as to when under what circumstances that would change, right?

1:25:04And, and, and so I think there's a, there's a conversation to be had there. that we need to have. And so, and I think there's a bunch of, yeah, so there's a bunch of stuff to say about that. I want to push back on the notion that there's sort of two options. There's like enslaved God, whatever that is, and like, loss of control. Yeah. And I think, like, we can do better than that. Right, like, let's work on it. Let's try to do better. Especially, you know, sort of, I think we can do better. and I think it might require being thoughtful, and it might require being having a kind of mature discourse about this before we start taking like irreversible moves.

1:25:51But I'm optimistic that we can at least avoid some of the connotations and a lot of the stuff that's taken that kind of binary. With respect to how we treat the AIs, so I have a couple of contradicting intuitions. And the difficulty with using intuitions in this case is obviously it's not clear what reference class an AI we have control over is. So to give one that's very scared about the things we're going to do these things, if you read about like life under Stalin or Mao, it's if you're there's one version of telling it, which is actually very similar to what we mean by alignment, which is we do these black box experiments about like, we're gonna make a thing that it can defect.

1:26:39And if it does, we know it's misaligned. And if you, Mao, the 100 Flowers campaign, where, you know, let 100 Flowers boom, I'm gonna allow criticism of my regime so on. And that's last for a couple of years. And afterwards, everybody who did that, that was a way to find the quote, unquote, the snakes, before the rightist or secretly hiding. And, you know, we'll like, it perns them. The paranoia of defectors, anybody in my entourage, any of my regime, they could be a secret capitalist trying to bring down the regime. That's the one way of talking about these things which is very concerning. Is that the correct reference class?

1:27:17I certainly think concerns in that vein are real. I think if it is disturbing how easy many of the analogies with human historical events and practices that we kind of deplore or at least have a lot of weariness towards are as in the context of the kind of way you end up talking about kind of AI maintaining controls over control over AI like making sure that it doesn't rebel like I think we should we should be noticing the kind of reference class that some of that talk starts to conjure. And so basically just, yes, I think we should be very, we should really notice that. Part of what I'm trying to do in the series is to bring the kind of full range of considerations at stake and to play, right?

1:28:16Like I think it is both the case that like, you, we should be quite concerned about like being kind of overly controlling or abusive or oppressive or there's all sorts of ways you can go too far. And I think there are concerns about the AI is being genuinely dangerous and genuinely acting, killing us, finally overfiring us. And I think the moral of situation is quite complicated. And then I think, Like in some sense, so often when you're, if you imagine a sort of external aggressor who's coming in and invading you, you feel very justified in doing like a bunch of stuff to prevent that. It's like a little bit different when you're like inventing the thing and you're doing it like incosuously or something and then you're also like I think this sort of moral justification you have for like there's a different vibe in terms of like the kind of Overall, yeah, justificatory stance you might have for various types of more power -exerting interventions.

1:29:30And so that's one feature of the situation. The opposite perspective here is that you're doing this sort of vice -based reasoning of like, the looks yucky of like doing reading being to send on these minds. And in the past, a couple of references, a couple of similar cases might have been something like environmentalists not liking nuclear power. And because the virus is nuclear, they don't look green, but obviously that's said back at the cause of fighting climate change. And so the end result of a future you're proud of, a future that's appealing, it's said back because your vibe's about, we would be wrong to bring nausea human, but you're trying to apply to a disanalogous case where that's not as relevant.

1:30:15I do think there's a concern here that I really try to foreground in the series that I think is related to what you're saying, which is something like, you might be worried that we will be very gentle and nice and free with the AIs and then they'll kill us. They'll take advantage of that and it will have been like a catastrophe. And I'm really trying to conjure that possibility at the same time as conjuring the grounds of gentleness and the sense in which it is also the case that these AIs could be, they can both be like others, moral patients, this sort of new species in the sense that should conjure wonder and reverence and such that they will kill you.

1:31:07you. And so I have this example of like, ah, the documentary grizzly man, where there's this environmental activist Timothy Treadwell, and he aspires to approach these grizzly bears. He lives, you know, in the summer, he goes into Alaska and he lives with these grizzly bears and he aspires to approach them with this like gentleness and reverence. He doesn't use bear mace, or he doesn't like curry bear mace. He doesn't use a fence around his camp. And he gets eaten alive by the bearers, or one of these bears. And I kind of really wanted to foreground that possibility in the series, like I think we need to be talking about these things both at once, right?

1:31:49And the bears can be moral patients, right? AI's can be moral patients, not these are moral patients. Enemy soldiers have souls, right? And so I think we need to learn the art of kind of and of both, like, kind of hard, there's this, like, dynamic here that we need to be able to hold both sides of as we kind of go into these trade -offs and these dilemmas and all sorts of stuff. And like a lot of part of what I'm trying to do in the series is like really kind of bring it all to the table at once. I think the big crux that I have, like if I today was to massively change my mind about what should be done is just the question of of how weird by default things end up, how alien they end up.

1:32:33And a big part of that story is the, you made a really interesting argument on your blog post that if moral realism is correct, that actually makes an empirical prediction, which is that the aliens, the ASIs, whatever, should converge on the right morality the same way that they converge on the right mathematics. At that, that was a really interesting point. But there's another prediction that moral realism makes, which is that over time, society should become more moral, become better. And to the extent that we think that's happened, of course there's the problem of what morals do you have now?

1:33:11Well, it's the ones that society has been converging towards over time. But to the extent that it's happened, one of the predictions of moral realism has been confirmed, which means should be updated in favor of moral realism. Only I want to flag is I don't think all forms of moral realism make this prediction. And so that's just one point. I'm happy to talk about the different forms I have in mind. I think there are also forms of kind of quasi things that kind of look like moral anti -realism, at least in their metaphysics, according to me, but which just posit that in fact there's this convergence.

1:33:44It's not in virtue of interacting with some like kind of mind independent moral truth, but just like as it's just for some other reason it's the case that and that looks like a lot like more realism at that point because it's kind of like Oh, it's really universal like everyone ends up here and it's kind of tempted to be like, ah like why right is that and then whatever answer for the why is a little bit like is that is that the doubt? Is that the nature of the doubt even if there's not sort of an extra metaphysical realm in which the moral lives or something so So, yeah, so moral convergence, I think, is sort of a different factor from the existence or non -existence of kind of non -natural, like a kind of morality that's not reducible to natural facts, which is the type of moral realism I usually consider.

1:34:25Now, okay, so does the improvement of society, is that an update towards moral realism? I mean, I guess like,

1:34:39so maybe it's like a very weak update or something, like I guess I'm kind of like, which view like predicts this hard? I guess it feels to me like moral anti -realism is like very comfortable with the observation of the like people with certain values have those values. Well, yeah, so there's obviously this like first thing, which is like any, if you're the culmination of some process of moral change, And then it's very easy to look back at that process and be more progress, like the arc of history bends towards me. You can look more, like if it was like, if there was a bunch of dice rolls around the way, you might be like, oh wait, that's not rationed, that's not the march of reason.

1:35:16So there's still empirical work you can do to tell whether that's what's going on.

1:35:22But I also think it's just, you know, on moral anti -realism, I think it's just still possible we'll say consider Aristotle and us. And we're like, okay, how's there been moral progress by Aristotle's lights or something? And our lights too, right? And you could think, ah, doesn't, isn't that a little bit like moral realism? It's like these hearts are singing in harmony. That's the moral realist thing, right? The anti -realist thing, the hearts all go different directions, but you and Aristotle apparently like are both excited about the kind of march of history. Some open question about whether that's true.

1:36:00What are Aristotle's reflective values? Suppose it is true. I think that's fairly explicable in moral and to realist terms. You can say roughly that you and Aristotle are sufficiently similar and you endorse sufficiently similar reflective processes and those processes are in fact instantiated in the March of history that history has been good for both of you. And I don't think that's, you know, I think there are worlds where that isn't the case. And so I think there's a sense in which maybe that prediction is more likely for realism than anti -reliasing, but it doesn't like move me very much.

1:36:40One thing I wonder is, look, there's, I don't know if moral realism is the right word, but the thing you mentioned about, there's something that makes hearts converge to the thing we are or the thing we upon reflection would be and even if it's not something that's like instantiated in a realm beyond the universe, it's like a force that exists that acts in a way we're happy with. To the extent that doesn't exist and you let go of the reins and then you get the paper clippers. It feels like we were doomed a long time ago in the sense of yeah, I just different utility functions banging against each other and some of them have a parochial preferences, but like, you know, it just combat and some guy won.

1:37:25Whereas in the world where like, no, this is the thing, like these are where the hearts are supposed to go or it's only by catastrophe that they don't end up there. That sort of, that feels like the world where like really matters. And in that world, the worry, the initial question I asked is like, what would make us think that alignment was a big mistake? In the world where the hearts just naturally end up towards like the thing, what we want. And maybe it takes an extremely strong force to push them away from that. And that extremely strong force is you solve technical alignment and just like, no, yeah, you're just like the brain, the blinders on the horse's eyes.

1:38:04So like in the world where like the world's really matter, we're like, ah, this is where the hearts want to go. In that world, maybe alignment is what fucks us up on this question of kind of do the world's where there's not this kind of convergent moral force, whether kind of metaphysically inflationary or not matter, are those the only roles that matter? Or sorry, maybe what I meant was in those worlds, you're kind of fucked. It's like, yeah, maybe the world's without that. The world's where there's no doubt. Yeah, yeah. Let's use the term doubt for this kind of convergent morality. Over the course of millions of years, it was going to go somewhere one way or another.

1:38:43It wasn't going to end up your particular utility function. Okay, well, let's just distinguish between ways you can be doomed. One way is philosophical. So you could be the moral realist or kind of realist -ish person of which there are many who have the following intuition. They're like, if not moral realism, then nothing matters. It's dust and ashes. which is, it is my metaphysics and or normative view or the void. And I think this is a common view. I think Derek Parfitt, at least some comments of Derek's Parfitt's suggested view. I think lots of more realists will like kind of profess this view.

1:39:27Ali Azov your Kowski, I think there are some sense in which I think his early thinking was inflected with this sort of thought. He later recanted. It's very hard. So I think this is importantly wrong. And so here's my, here's the case, I have an essay about this, it's called Against the Normative Realists Wager. And here's the case that convinces me. So imagine that a metathical fairy appears before you, right? And this fairy knows whether there is a Dow. And the fairy says, okay, I'm going to offer you a deal. If there is a Dow, then I'm going to give you $100. dollars. If there isn't a Dow, then I'm going to burn you and your family and 100 innocent children alive.

1:40:14Okay, so claim, don't take this deal. This is a bad deal. You're holding hostage, your commitment to not being burned alive or like you're care for that. To this like abs truce, basically, you're, yeah, like I think, I mean, I go through in the essay I've won two different ways in which I think this is wrong, but I think just like, and I think these people who kind of pronounce, say like moral realism or the void, like they don't actually think about that's like this. I'm like, no, no, okay, so really, like does that what you want to do? And no, I think we should, I still care about my valid, my sort of allegiance to my values, I think is kind of outstrips the my like commitment to like various like medical interpretations in my values.

1:40:57I think like we should, the sense in which we like care about, not being burned alive as much more solid than like our kind of, you know, then the reasoning and on what matters. Okay, so that's this, that's like the sort of philosophical doom. Right. Now you could have this, it's not like you were also gesturing at at a sort of empirical doom, right? Which is like, okay, dude, if it's all, if it's just going in a zillion directions, come on, you think it's going to go in your direction, like there's going to be so much You're just gonna lose and so you should give up now and kind of only fight for the realism worlds.

1:41:41And there I'm like, I mean, so I think you got to do the expected value calculation. You got to actually have a view of how doomed are you in these different worlds, what's the tractability of changing different worlds? I mean, I'm quite skeptical of that, but that's a kind of empirical claim. I'm also just kind of low on this, like, everyone converges. So if you imagine like, you know, you train a chess playing AI or you have a real paper clipper, right? Somehow you had a real paper clipper and then you're like, okay, you know, go and reflect. based on my understanding of how moral reasoning works.

1:42:21Like if you look at the type of moral reasoning that analytic ethicists do, it's just reflective equilibrium, right? They just take their intuitions and they systematize them. I don't see how that process gets a sort of injection of the kind of mind and dependent moral truth, or like I guess it, like if you sort of start with only all of your intuition and say to maximize paper clips. I don't see how you end up maximizing or doing some rich human morality. I just don't, it doesn't look to me like that's how human ethical reasoning works. I think most of what normative philosophy does is make consistent and kind of systematized pre -theoretic intuitions.

1:43:05And so, and I think, but we'll get evidence about this. In some sense, I think this view predicts, you keep trying to train the AIs to do something and they keep being like, no, I'm not gonna do that. I was like, no, that's not good. Or something, they keep pushing back. The momentum of AI cognition is always in the direction of this moral truth. And whenever we try to push it in some other direction, we'll find resistance from the rational structure of things. So sorry, actually, I've heard from researchers who are doing alignment that for red teaming inside these companies, they will try to red team a base model.

1:43:41So it's not been Arles who have to just predict next token and the raw, crazy, whatever, Shahgat. And they tried to get this thing to, hey, help me make a bomb, help me whatever. And they say that it will, like it's odd how hard it tries to refuse, even before it's been Arleaged. I mean, look, it will be a very interesting fact. If it's like, man, we keep training the AI, it's in all sorts of different ways. Like we're doing all this crazy stuff. And they keep acting like bourgeois liberals. It's like, wow. Like that's, or you know, So they keep like really, or they keep professing this like weird alien reality.

1:44:17They all converge on this one thing. They're like, can't you see? It's like Zorgo, like Zorgo is the thing and like all the AIs, you know, interesting, very interesting. I think my personal prediction is that that's not what we see. And my actual prediction is that the AIs are going to be very malleable. Like we're going to be like, you know, if you push an AI towards evil, like it'll just go. And I think that's obviously, it was sort of, it's reflective of the consistent evil. I mean, I think there's also a question with some of these AI's, it's like, will they even be consistent in their values?

1:44:53Right? I do think like, I think we can do, so I like the image of the blinded horses, and I like the image of like, maybe alignment is gonna mess with the, I think we should be really concerned if we're like forcing facts on our AI's. Right? Like that's like a really bad, Because I think one of the clearest things about human processes of reflection, like the kind of easiest thing to be like, let's at least get this is like not acting on the basis of an incorrect empirical picture of the world. Right. And so if you find yourself like asking her, by the way, like this is true, and I need you to always be reasoning as though blah is true.

1:45:29I'm like, ooh, I think that's a no -no from an anti -realist perspective too. Because I want to, like, my reflective values, I think, will be such that I formed them in light of the truth about the world. And so I think, and I think this is a real concern about, as we move into this era of kind of aligning AI's, I don't actually think this binary between values and other things is going to be very obvious in how we're training them. I think it's going to be much more like ideologies. And like, you can just train an AI to like, output stuff, right, output utterances. And so you can easily end up in a situation where you like, decided that Blah is true about some issue, an empirical issue, right?

1:46:07Not a moral issue. And so I think people should not, for example, I do not think people should hard code belief in God into their AIs. Or like I would advise people to not hard code their religion into their AIs if they also want to like discover if their religion is false. I would just in general, if you would like to have your behavior be sensitive, just whether something is true or false, like it's sort of generally not good to like etch it into things. And so that is definitely a form of blinder, I think we should be really watching out for. And I'm kind of hopeful, so I have enough credence on some sort of moral realism that I'm hoping that if we just do the anti -realism thing of just being consistent, learning all the stuff, reflecting, if you look at how moral realism, moral anti -realists actually do normative ethics, it's the same.

1:46:55It's basically the same. There's some amount of different heuristics on things like properties like simplicity and stuff like that, but I think it's like they're mostly just doing the same game. And so I'm kind of hoping that, and also metathics is itself a discipline that AIs can help us with. I'm hoping that we can just figure this out either way. So if there is, if moral realism is somehow true, I want us to be able to notice that. And I want us to be able to like adjust accordingly. So I'm not like writing off those worlds and be like, let's just like totally assume that's false. But the thing, the thing I really don't want to do is write off the other worlds where it's not true.

1:47:29because my guess is it's not true. And I think stuff still matters a ton in those worlds too. So Blendick Crocs is like, okay, you're training these models. We're in this incredibly lucky situation where it turns out the best way to train these models is to just give them everything humans ever said, written thought. And also these models, the reason they get intelligence is because they can generalize, right? Like thinking, brok, what is it? What is the gist of things? So, are we fundamentally very, should we just expect this to be a situation which leads to alignment in the sense of how exactly does this thing that's trained to be in a mugumation of human thought become a paper clipper?

1:48:13The thing you kind of get for free is it's an intellectual descendant. The paper clipper is not an intellectual descendant whereas the AI which understands all the human concepts, but then get stuck on some part of it, which we aren't totally comfortable with, is like, you know, it's, it feels like an intellectual descendant in the way we care about. I'm not sure about that. I'm not sure I, I'm not sure I do care about a notion of intellectual descendant in that sense, like if you imagine, I mean, literal paperclips is a human concept bridge. So, I don't think any old, any old human concept will, will do for the thing, the thing we're excited about.

1:48:52I think the stuff that I would be more interested in the possibility of getting for free are things like

1:49:02consciousness, pleasure, sort of other features of human cognition. Like I think, so there are paper clippers and there are paper clippers, right? So imagine if the paper clipper is like an unconscious, kind of a voracious machine, and it's just like it appears to you as a cloud of paperclips. But there's not sort of, that's like one vision. If you imagine the paperclips is like a conscious being that loves paperclips, it takes pleasure in making paperclips. That's like a different thing. And obviously it could still, it's not necessarily the case that it makes the future all paperclips is probably not optimizing for conscious and surplusure.

1:49:42It cares about paperclips, maybe eventually, if it's suitably certain, it turns itself into paperclips. and who knows, but like it's still I think a different, it's actually a somewhat different moral kind of mode with respect. That looks to be much more like a, you know, there's also questions like does it try to kill you and stuff like that. But I think that the, there are kind of features of the agents we're imagining other than the kind of thing that they're staring at that can matter to our sense of like sympathy, similarity.

1:50:15and I think people have different views about this. So one possibility is that human consciousness, like the thing we care about in consciousness or sentience, is super contingent and fragile and most minds, most smart minds are not conscious, right? It's like the thing we care about with consciousness is this hacky contingent. It's like a product of like specific constraints, evolutionarily genetic, bottlenecks, et cetera. And that's why we have this consciousness. And like you can get similar work done. So consciousness presumably does some sort of work for us, but you can get similar work done in a different mind in a very different way and you should sort of so that's like that's a sort of consciousness that's fragile view, right?

1:50:52And I think there's a different view which is like no consciousness is is something that's quite structural. It's much more defined by functional roles like self -awareness a concept of yourself maybe higher order thinking stuff that you really expect in many sophisticated minds And in that case, okay, well now actually consciousness isn't as fragile as you might have thought, right? Now actually like lots of beings, lots of minds are conscious and you might expect at the least that you're gonna get like conscious super intelligence. They might not be optimizing for creating tons of consciousness, but you might expect consciousness by default.

1:51:28And then we can ask similar questions about something like valence or pleasure or like the kind of character of the consciousness, right? So there's, you can have a kind of cold, indifferent consciousness that has no like human, or no like emotional warmth, no like pleasure or pain. I think that can still be, Dave Traumers has some papers about like Vulcans and he talks about they still have moral patient hood. I think that's very plausible, but I do think it's like, an additional thing you could get for free, or like get quite commonly depending on its nature is something like pleasure. again, and then we have to ask how Janky is a pleasure, how specific and contingent is the thing we care about in pleasure versus how robust is this as a functional role in like minds of all kinds.

1:52:10And I personally don't know on this stuff. And I don't think this is like enough to get you alignment or something, but I think it's at least worth being aware of like these other features. We're not sort of talking, we're not really talking about the AI's values. In this case, we're talking about like the kind of structure of its mind and the different properties the minds have. And I think that could show up quite robustly. So part of your day job is writing with these kinds of section 2 .2 .2 .5 type reports. And part of it is like society is like a tree that's growing towards the light. What is the like context -wishing between the two of them?

1:52:52So I actually find it's kind of quite complimentary. So yeah, I will write these sort of more technical reports and then do this sort of kind of more literary writing and philosophical writing. And I think they both draw in kind of like different parts of myself and I try to think about them in different ways. So I think about the reports as are much more like this is like I'm more fully optimizing for like trying to do something impactful or trying to kind of kind of. Yeah, there's more of an impact orientation there, and then on the essay writing, I give myself much more leeway to let other parts of myself and other parts of my concerns come out and self -expression and aesthetics and other sorts of things.

1:53:39Even while they're both, I think, for me, part of an underlying similar concern or attempt to have an integrated orientation towards the situation. Can you explain the nature of the transfer between the two? So in particular, from the literary side to the technical side, I think rationalists are known for having a sort of ambivalence towards great works or humanities. Are they missing something crucial because of that? Because one thing you notice in your essays is just lots of references to epigraphs, to lines and poems or essays that are particularly relevant. I don't know, are the rest of the rationalist missing something because they don't have that kind of background?

1:54:27I mean, I don't want to speak. I think some rationalists, you know, lots of the last rationalists, like, a lot of these different things. I do think, by the way, I'm just referring specifically to SVF as a post about like how Shakespeare could be, like, the base rates of Shakespeare being a great writer and also books can be condensed to essays. Well, so on just the general question of like how should people value great works or something? I think people can kind of fail in both directions, right? And I think some people maybe like maybe SPF or other people, they're sort of interested in puncturing a certain kind of like sacredness and prestige that people can try to kind of like yeah, that people associate with some of these some of these works.

1:55:07And I think there's a way in which and then as a result, can miss some of the genuine value. But I think they're responding to a real failure mode on the other end, which is to kind of, yeah, be too enamored of this prestige and sacredness to kind of siphon it off as some weird, legitimating function for your own thought, instead of like thinking for yourself, losing touch with like, what do you actually think or what do you actually learn from? Like I think some, you know, these epigraphs, careful, right? I mean, it's like, I think, you know, And I'm not saying I'm immune from these vices. I think there can be a like, oh, but Bob said this.

1:55:40And it's like, whoa, very deep, right? And so these are humans like us, right? And I think the canon and other great works and all sorts of things have a lot of value. And we shouldn't, I think sometimes it borders on the way people read scripture or I think there's a kind of scriptural authority that people will sometimes like, ascribe to these things. And I think that's not. So yeah, I think it's kind of, you can fall off on both sides of the horse. It actually relates really interestingly to, I remember I was talking to somebody who, at least it's familiar with rationalist discourse and I was telling him he was asking, like, what are you interested in these days and I was saying something about this part of Roman history, super interesting.

1:56:22And then his first sort of response was, oh, it's really interesting when you look at these secular trends of like Roman times to what happened in the dark ages versus the enlightenment. And for him, it was like, the story of that was just like, how did it contribute to the big secular, like the big picture, that the sort of particulars didn't, they don't, there's no interest in that. It's just like, if he's doing that at the biggest level, what's happening here? Whereas there's also the opposite feeling of when people with study history, Dominic Cummings writes about this because he is endlessly frustrated with the political class in Britain.

1:56:57And he'll say things like, they study politics, philosophy, and economics. and a big part of it is just like being really familiar with these poems and like reading a bunch of history about the War of the Roses or something. But he's frustrated that they take away, they have all these like kings memorized, but they take a very little in terms of lessons from these episodes. It's more of just like almost entertaining, like watching Game of Thrones for them. Whereas he thinks like, we're repeating certain mistakes that he's seen in history. Like he can generalize in a way they can't. So the first one seems like a mistake.

1:57:29I think CS Lewis talks about in the one of the essays you cited where it's like, if you see through everything, it's like you're really blind, right? Like if everything is transparent. I mean, I think there's kind of very little excuse for like not learning history, or I don't know, sorry, I mean, I'm not saying I like have learned enough history. I guess I feel like even when I try to channel some sort of vibe of like skepticism towards like great works, I think that doesn't generalize to like thinking it's not worth understanding human history. I think human history is like, just so clearly, Crucial to kind of understand, it's what's structured and created all of the stuff.

1:58:10And so there's an interesting question about what's the level of scale at which to do that, and how much should you be like, yeah, looking at details, looking at macro trends, and that's a dance. I do think it's nice, I think it's nice for people to be like, At least attending to the kind of macro narrative, I think there's like a, there's some virtue in like having a world view, like really like building a model of the whole thing, which I think sometimes gets lost in like the details. And, but obviously like if you're two, you know, the details are what the world is made of. And so if you don't have those, you don't have data at all.

1:58:50So, yeah, seems like there's some skill in like learning history, history well. Mm. The essentially seems related to, you have a post on sincerity. And I think like, if I'm getting the sort of the vibe of the piece right, it's like, at least in the context of let's say intellectuals. Certain intellectuals have a vibe of like shooting the shit, and they're just like trying out different ideas. How do these like, how do these analogies fit together? Maybe there's some, and those seem closer to the, I'm looking at the particulars and And like, oh, this is just like that one time in the 15th century where they overthrew this king and they blah, blah, blah.

1:59:31Whereas the guy who was like, oh, here's a secular trend from like the, if you look at the growth models for like a million years ago to now, it's like here's what's happening. That one has a more sort of sincere flavor. Some people, especially when it comes to AI discourse, have a very, this insure mood of operating is like, I've thought through my bio -ankers and I disagree with this premise. So here, my effective compute estimate is different in this way, here's how I analyze the scaling laws. And if I could only have one person to help me guide my decisions on the AI, I might choose that that person, but I feel like if I could choose between, if I had 10 different advisors at the same time, I might prefer the shooting the shit type characters who have these weird esoteric intellectual influences and they're almost like random number generators.

2:00:28They're not necessarily calibrated, but once in a while they'll be like, oh, this is like one weird philosophy I care about or this one historical event I'm obsessed with has an interesting perspective on this. And they tend to be more intellectually generative as well because they're not. I think one part one big part of it is that And if you are so sincere, you're like, oh, I've thought through this, obviously, ASI is the biggest thing that's happening right now. It doesn't really make sense to spend a bunch of your time thinking about how did the command she's live and what is the history of oil and how to jarar to think about conflict.

2:01:03You know, just like, what are you talking about? Come on, ASI is happening in a few years, right? Whereas, but therefore, the people who go on these rabbit holes, because they're just trying to shoot the shit I have, I feel like I'm much generative. I mean, it might be worth distinguishing between something like kind of intellectual seriousness. Right. And something like how diverse and wide -ranging and kind of idiosyncratic are the you know, things you're interested in. Right. And I think maybe there's some correlation where people who are kind of like, or maybe intellectual seriousness is also distinguishable from something like shooting a shit, like maybe you can shoot the shit seriously.

2:01:44I mean, there's a bunch of different ways to do this, but I think having an exposure to like all sorts of different sources of data and perspectives seems great. And I do think it's possible to like curate your kind of intellectual influences too rigidly in virtue of some story about what matters. Like I think it is good for people to like have space. I mean, I'm really a fan of, or I appreciate the way like, I don't know, I try to give myself space to do stuff that is not about this is the most important thing. And that's feeding other parts of myself. And I think parts of yourself are not isolated.

2:02:18They feed into each other and I think a better way to be a richer and fuller human being in a bunch of ways. And also, these sorts of data can be just really directly relevant. And I think some people, I know who I think of as quite intellectually sincere and in some sense, quite focused on the big picture. Also have a very impressive command of this very wide range of empirical data. And they're really, really interested in the empirical trends. and they're not just like, oh, you know, it's a philosophy, or, you know, sorry, it's not just like, oh, history, it's the march of reason, or something.

2:02:44No, they're like really, they're really in the weeds. I think there's a kind of, in the weeds, um, uh, virtue that I actually think is like closely related in my head with, with some kind of seriousness and sincerity. Um, I do think there's a different dimension, which is there's like kind of trying to get it right. And then there's kind of like, throw stuff out there and I try to like, what if it's like this or like try this on or, yeah, I have a hammer. I will hit everything. What if I did everything with this hammer? And so I think some people do that. And I think there is room for all kinds.

2:03:17I think the thing where you just get it right is kind of undervalued. Or I mean, it depends on the context you're working in. I think certain sorts of intellectual cultures and milleus and incentive systems, I think incentivize, saying something new or saying something original or saying something like flashy or provocative or and then like kind of various cultural and social dynamics. I'm like, oh, like, mmm, mmm, mmm, mmm, people are like doing all these kind of, you know, kind of performative or statusy things. Like there's a bunch of stuff that goes on when people do thinking. And, you know, cool.

2:03:53But like, if something's really important, let's just get it right. And I think, and sometimes it's like boring, but it doesn't matter. And I also think like stuff is less interesting if it's false, right? Like I think if someone's like, bra, and you're like, no, I mean, it can be useful. I think sometimes there's an interesting process where someone says like, blah, provocative thing. And it's a kind of an epistemic project to be like, wait, why exactly do I think that's false, right? And you really, you know, someone's like, healthcare doesn't work. Medical care does not work, right? Someone says that and you're like, all right, how exactly do I know that medical care works, right?

2:04:32And you like go through the process of trying to think it through. And so I think there's like room for that. But I think ultimately like what like kind of the real profundity is like true, right? Or like kind of things, things become less interesting if they're just not true. And I think that's I think sometimes it feels to me like people or it's at least possible, I think, to lose touch with that and to be more flashy and it's kind of like, and this actually isn't, there's not actually something here, right? One thing I've been thinking about recently, after I interviewed Leopold was, while prepping for it, listen, I haven't really thought at all about the fact that there's gonna be a geopolitical angle to this AI thing.

2:05:15And it turns out, if you actually think about the natural security implications, that's a big deal. No, I wonder, given the fact that that was something that was not my radar right now. I was like, oh, obviously that's a crucial part of the picture. How many other things like that there must be? And so even if you're coming for the perspective of like AI is incredibly important, if you did happen to be the kind of person who was like, ah, you know, everyone's not allowed like checking out different kinds of, I'm like incredibly curious about what's happening in Beijing. And then you, then the kind of thing that later on you realize was like, oh, this is a big deal.

2:05:47You have more awareness of you can spot it in the first place. Whereas I wonder... So maybe there's not necessarily a trade -off. Like, the rational thing is to have some sort of really optimal explore exploit trade -off here where you're like constantly searching things out. So I don't know practically that works out that well, but that experience made me think, like, oh, I really should be trying to expand my horizons in a way that's undirected to begin with because there's a lot of different things about the world that I'd understand to understand any one thing. I mean, I think there's also room for division of labor, right?

2:06:26Like I think there can be, yeah, like, you know, there are people who are like trying to like drive onto pieces and then be like, here's the overall picture and then people who are going really deep on specific pieces, people who are doing them more like generative, throw things out there, see what sticks. So I think there, it also doesn't need to be that like, all of the epistemic labor is like located in one brain. And, you know, it depends like your role in the world and other things. So in your series you express sympathy with the idea that even if an AI or I guess any sort of agent that doesn't have consciousness has a certain wish and is willing to pursue it nonviolently we should respect its rights to pursue that.

2:07:10And I'm curious where that's coming from, because conventionally, I think the thing matters because it's conscious and it's conscious experience as a result of that pursued matter. Well, I don't know. I don't know where this discourse leads. I'm suspicious of the amount of ongoing confusion that it seems to me as present in our conception of consciousness. Sometimes I think of analogies with people talk about life and like a Lonvi Tal, right? And maybe, you know, there's a world, you know, a Lonvi Tal was this like hypothesized life force that is sort of the thing that's taken life. And I think, you know, we don't really use that concept anymore.

2:07:50We think that's like a little bit broken. And so I don't think you want to have ended up in a position of saying like, everything that doesn't have a Lonvi Tal is doesn't matter or something, right? Cause then you end up later. And then so much thing. And then so much similarly, if you, Even if you're like, no, no, there's no such thing as a laundry towel, but life surely life exists. And I'm like, yeah, life exists. I think consciousness exists too. Likely, depending on how we define the terms, I think it might be a kind of verbal question.

2:08:19The even once you have a kind of reductionist conception of life, I think it's possible that it kind of becomes less attractive as a moral focal point, right? So like, right now we really think of consciousness where like it's a deep fact. It's like, so consider a question like, okay, so take a

2:08:37cellular automata, right? That is sort of self -replicating. It has like some information that, you know, and you're like, okay, is that alive? Right? It's kind of like, it's not that interesting. It's a kind of verbal question, right? Like, or, or, I don't know, philosophers might get really into like, is that alive? But you're not missing anything about this system, right? It's not like, there's no extra life that's like springing up. It's just like, It's alive in some senses, not alive in other senses. And I think if you, but I really think that's not how we intuitively think about consciousness.

2:09:05We think whether something is conscious is a deep fact. It's this like additional, it's like this really deep difference between being conscious or not. It's like it's someone home is the lights are on, right? And I have some concern that if that turns out not to be the case, then this is going to have been like a bad thing to like build our entire ethics around. That's really interesting. And so, now to be clear, I take consciousness really seriously. I'm like, man, consciousness. I'm not one of these people like, oh, obviously, new consciousness doesn't exist or something. I'm like, but I also notice how confused I am and how dualistic my intuitions are.

2:09:39And I'm like, wow, this is really weird. And so I'm just like, error bars around this. Anyway, so that's like, there's a bunch of other things going on in my wanting to be open to kind of not making consciousness like this kind of fully necessary criteria. I mean, clearly, I definitely have the intuition. Like, consciousness matters a ton. I think like if something is not conscious and there's like a deep difference between conscious and unconscious, then I'm like, definitely have the intuition that is sort of, there's something that matters especially a lot about consciousness. I'm not trying to be like dismissive about the notion of consciousness.

2:10:07I just think we should be like quite aware of how, it seems to me how ongoingly confused we are about its nature. Okay, so suppose we figure out that consciousness is just, like a word we use for a hot podger, different things, only some of which encompass what we care about. Maybe there's other things we care about that are not included in that word similar to the life force analogy. Then where do you anticipate that would leave us as far as ethics goes? Like would then there be a next thing that's like consciousness or what do you anticipate that would look like? So there's a class of people who are called illusionists and philosophy mind, which which, who will say consciousness does not exist.

2:10:54And this is sort of, it's different ways to understand this view. But one version is to sort of say that the concept of consciousness has built into it too many preconditions that aren't met by the real world. So we should sort of chuck it out, like Elon Vita. Like instead of the sort of proposal is kind of like at least phenomenal consciousness, right? Or like qualia, or what it's like to be a thing. They'll just say this is like sufficiently broken, can sufficiently chock full of falsehoods that we should just not use it. I think it feels to me like I am like, there's really clearly a thing, there's something going on with, like I'm kind of really not, I kind of expect to, I do actually kind of expect to continue to care about something like consciousness quite a lot on reflection and to not kind of end up deciding that my ethics is like better, it doesn't make any reference to that, or at least there's some things like quite nearby to consciousness.

2:11:54Like when I stub my toe and I have this, like something happens when I stub my toe. Unclear exactly how to name it, but I'm like, something about that. You know, I'm like pretty focused on it. And so I do think, you know, in some sense, if you're like, well, where do things go? I'm like, I should be clear, I have a bunch of greetings that in the end, we end up carrying a bunch about consciousness, just directly. And so if we don't, like, yeah, I mean, where will ethics go? Where will a completed philosophy of mine go? Very hard to say. I mean, I can imagine something that's more, like I think, I mean, maybe a thing that, I think a move that people might make if you get a little bit less interested in the notion of consciousness is some sort of slightly more like animistic, like so what's going on with the tree?

2:12:42And you're like, maybe not like talking about it as a conscious entity necessarily. But it's also not totally unaware or something. And so the consciousness discourse is right with these funny cases where it's sort of like, oh, like those criteria imply that this totally weird entity would be conscious or something like that. Especially if you're interested in some notion of agency or preferences. A lot of things can be agents, corporation, all sorts of things. Like corporations, conscious, and so on, man. But I actually think it's a one place it could go. in theory is in some sense you start to view the world as animated by moral significance in kind of richer and subtler structures than we're used to.

2:13:22And so plants or weird optimization processes are outflows of complex, I don't know, who knows exactly what you end up seeing as infused with the sort of thing that you ultimately care about. But I think it is possible that that doesn't map that includes a bunch of stuff that we don't normally ascribed consciousness too. I think when you use a complete theory of mind, and presumably after that, a more complete ethic, even the notion of a reflective equilibrium implies, oh, you'll be done with it at some point, you just sum up all the number, and then you've got the thing you hear about. This might be unrelated to the same sense we have in science, But I think like this the vibe you get when you're talking about these kinds of questions is that oh, you know We're like rushing through all the science right now and we've been churning through it It's getting harder to find because there's some like cap like you find all the things at some point right now It's super easy because like a semi -intelligence species barely has emerged and the ASI will just rush through everything incredibly fast and like Then you will either have aligned its heart or not.

2:14:40In either case, it'll use what it's figured out about like what is really going on and then expand through the universe and exploit, you know? Like do the tiling or maybe some more benevolent version of quote unquote tiling. That feels like the basic picture of what's going on. We had dinner with Michael Nielsen a few months ago. And his view is that this just keeps going forever or close to forever. how much would it change your understanding of what's going to happen in the future if you were convinced that Nielsen is right about his picture of science? Yeah, I mean, I think there's a few different aspects.

2:15:16There's kind of... My memory of this conversation, I don't claim to really understand Michael's picture here, but I think my memory was sort of like, sure you get the fundamental laws. I think my impression was that he expects sort of physics, the kind of physics to get solved or something, maybe modular, like the expense of nestle certain experiments or something. But the difficulty is like even granted that you have the kind of basic laws down that still actually doesn't let you predict like where at the macro scale like various useful technologies will be located. Like, there's just still this like big search problem.

2:15:55And so my memory though, I'll let him speak for himself on what his take is here. But my memory was sort of like, sure, you get the fundamental stuff, but that doesn't mean you get the same tech. I'm not sure if that's true. I think if that's true, what kind of difference would it make? So one difference is that, well, so here's a question. So like, it means at some times you have to do, you have to, at a more ongoing way, make trade -offs between investing in further knowledge and further exploration versus exploiting, as you say, sort of acting on your existing knowledge because you can't get to a point where you're like, and we're not.

2:16:40Now, as I think about it, I mean, I think that's, you know, I sort of suspect that was always true and like, I remember talking to someone, I think I was like, ah, we should, at least in the future, we should really get like all the knowledge. And he's like, well, what do you wanna like, You don't want to know the output of every touring machine or like, you know, in some sense, there's a question of like, what actually would it be to have like a completed knowledge? And I think that's a rich question and it's own right. And I think it's like not necessarily that we should imagine even in this sort of, on any picture, necessarily that you've got like everything.

2:17:10And on any picture, in some sense, you could end up with this case where you cap out like there's some collider that you can't build or whatever. Like there's some, something is too expensive or whatever, and kind of everyone caps out there. So there's, I guess, one way to put it is, so there's a question of like, do you cap? And then there's a question of like, how contingent is the place that's right? You go. If there's contingent, I mean, one thing one prediction that makes is you'll see more diversity across our universe or something. If there are aliens, they might have like, quite different tech.

2:17:43And so maybe like, if people meet, you don't expect them to be like, ah, you got your thing, I got, I am our version and said, whoa, like that thing. Wow. So that's one thing. If you expect more like ongoing discovery of tech, then you might also expect more ongoing change and upheaval and churn in so far as technology is one thing that really drives change in civilization. So that could be another. People sometimes talk about lock -in. And then it's like, oh, I sort of envision this kind of point at which civilization is kind of settled into some structure or equilibrium or something. And maybe you get less of that.

2:18:21I think that's maybe more about the pace rather than contingency or caps, but that's another factor. So yeah, I think it is an interesting, I don't know if it changes the picture fundamentally of earth civilization. We still have to make trade -offs about how much do you invest in research versus acting on our existing knowledge. But I think it has some significance. I think one vibe you get when you talk to people, we're at a party and somebody mentioned this, we're talking about how uncertain and should we be with the future? And they're like, there are three things I've been certain about, like what is consciousness, what is information theory, and what are the basic laws of physics?

2:18:54I think once we get that, we're done. And that's like, oh, you'll figure out what's the right kind of phedonium. And then like, you know, that has that vibe. Whereas this like, oh, you like, you're like constantly shurning through. And it has more of a flavor of like more of the becoming that the attunement picture implies. I think it's more exciting. It's not just like, oh, you figured out the things in the 21st century and then you just, you know what I mean. Yeah, I mean, I sometimes think about the sort of two categories of views about this. There's people who think like, yeah, the knowledge, we're almost there and then we've like, yeah, basically got the picture, right?

2:19:39And where the picture is sort of like, yeah, the knowledge is all just totally sitting there. And it's like you just have to get to like remote, there's like this kind of, just you have to be like scientifically mature at this. All right. And then it's just gonna all fall together, right? And then everything past that is gonna be like this super expensive, like not super important thing. And then there's a different picture, which is much more this like ongoing mystery, like ongoing, like oh man, there's like gonna be more and more, like maybe expect more radical revisions to our world view. And I think it's an interesting, Yeah, I think I'm kind of drawn to both.

2:20:12Physics were really good at physics, right? Or a lot of our physics is quite good at predicting a bunch of stuff. Or at least that's my impression. This is reading some physicists. Who knows? That's a physicist, though, right? Yeah, but this isn't coming from my dad. There's a blog post, I think Sean Carroll or something. We really understand a lot of the physics that governs the everyday world. A lot of it. We're really good at it. I'm genuinely pretty impressed by physics. because at this point I think that could probably be right. On the other hand, really these got, had a few centuries of, so anyway, but I think that's an interesting, and it leads to a different, I think it does, there's something, the endless frontier, there is a draw to that from an aesthetic perspective of the idea of continuing to discover stuff.

2:20:59The least I think you can't get full knowledge, in some sense, because there's always like, like what are you gonna do? There's some way in which you're part of the system. So it's not clear that the knowledge itself is part of the system and sort of like, I don't know, like if you imagine you're like, ah, you try to have full knowledge of what the future of the universe will be like. Well, I don't know, I'm not totally sure that's true. It has a halting problem, kind of proper, right? There's a little bit of a loopiness if you're, yeah. I think there are probably like fixed points in that where you could be like, yep, I'm gonna do that.

2:21:31And then like, yeah, right. But I think it's, I at least have a question of like, are we, you know, when people imagine the kind of completion of knowledge, you know, exactly how well does that work? I'm not sure. You had a passage in your essay on Utopia where I think you're, the vibe there was more of the thing that were, the positive future we're looking forward to, it will be more of like, you, unless you describe what you mean. But like, To me, it felt more like the first stuff. You get the thing and then now you've found the heart of the... Maybe can I ask you to read that passage real quick?

2:22:08Oh, sure. And that way I'll spur the discussion I'm interested in having this part in particular. Right. Quote, I'm inclined to think that Utopia, however weird, would also be in a certain sense recognizable. That if we really understood and experienced it, we would see it, we would see in it the same thing that made us sit bolt upright long ago when we first touched love, joy, beauty that we would feel in front of the bonfire, the heat of the ember from which it was lit. There would be, I think, a kind of remembering. Where does that fit into this picture? It's a good question. I mean, I think it's like some guess about, like, if there's like no part of me that recognize it recognizes it is good.

2:22:58Then I think I'm not sure that it's good according to me, in some sense. So yeah, I mean, it is a question of like what it takes for it to be the case that a part of you recognize it is good, but I think if there's really none of that, then I'm not sure it's a reflection of my values at all. There's this sort of tautological thing you can do where it's like, if I went through the processes which led to me discovering was good, which we might call reflection, then it was good. But by definition, you ended up there because it was like, you know what I mean? Yeah, I mean, you definitely don't want to be like, like, you know, if you transform me into a paper clipper, gradually.

2:23:41Yeah, yeah. Then I will eventually be like, and then I saw the light, yeah, I saw the true paper clips. Yeah. Right. But that's part of what's complicated about this thing about reflection. and you have to find some way of differentiating between the sort of development processes that preserve what you care about and the development processes that don't. And that is in itself is this like fraught question, which itself requires like taking some stand on what you care about and what sorts of meta processes you endorse and all sorts of things. But you definitely shouldn't just be like, it is not a sufficient criteria that the thing at the end thinks it got it right.

2:24:16Right. Because that's compatible with having gone like wildly off the rails. Yeah, yeah, yeah. There was a very interesting sentence you had in your post, one of your posts where you said, our hearts have in fact been shaped by power. So we should not be at all surprised if the stuff we love is also powerful.

2:24:40Yeah, what's going on there? I actually could want to think about what did you mean there? Yeah, so the context, the context on that post is I'm talking about this hazy cluster which I call in the essay, niceness slash liberalism slash boundaries, which is this sort of like somewhat more minimal set of like cooperative norms involved in like respecting the boundaries of others and kind of cooperation and peace amongst differences and like tolerance and stuff like that, as opposed to like your favored structure of matter, which is sort of sometimes the paradigm of values that people use in the context of AI risk.

2:25:19And I talk for a while about the sort of ethical virtues of these norms, but it's pretty clear that also, why do we have these norms? Well, one important feature of these norms is that they're kind of effective and powerful. Liberal societies are secure boundaries, save resources wasted on conflict. And like liberal societies are often more like, they're better to live in, they're better to immigrate to, they're more productive, all sorts of things. nice people, they're better to interact with, they're better to like trade with, all sorts of things, right? And I think it's pretty clear if you look at the, both like why at a political level do we have like various political institutions?

2:25:57And if you look kind of more deeply into our evolutionary past and like how our moral cognition is structured, it seems like pretty clear that various like kind of forms of cooperation and like kind of game through reddict dynamics and other things went into kind of shaping what we now, at least in certain contexts also treat as a kind of intrinsic or terminal value. So like some of these values that have kind of instrumental functions in our society are also kind of rayified in our cognition as kind of intrinsic values in themselves. And I think that's okay. I don't think that's a debunking. Like all of your values are kind of like some something that kind of stuck and got kind of treated as a terminally important.

2:26:43But I think that means that sometimes the way in the context of the series where I'm talking about deep atheism and our sort of relationship, the relationship between what we're pushing for and what nature is pushing for or what sort of pure power will push for. And it's easy to say, well, there's paperclips, which is just one way place you can steer and pleasure is just another place you can steer or something. and these are just sort of arbitrary directions, whereas I think some of our other values are much more structured around cooperation and things that also are kind of effective and functional and like a powerful.

2:27:25And so that's what I mean there. As I think there's a way in which we're sort of nature is a little bit more on our side than you might think because part of CRER is like has been made by a kind of nature's way. And so that is like in us. Now I don't think that's enough necessarily for us to beat the gray goo. Like we have some amount of like power built into our values, but that doesn't mean it's kind of going to be such that it's kind of arbitrarily competitive. But I think it's still important to keep in mind that this is, and I think it's important to keep in mind in the context of integrating AI's into our society that I think, you know, we've been talking a lot about the ethics of this.

2:28:03But I think there's also there are like instrumental and kind of practical reasons to want to have forms of social harmony and cooperation with AI's with different values. And I think we need to be taking that seriously and thinking about what is it to do that in a way that's genuinely legitimate and a project that is a kind of just incorporation of these beings into our civilization such that they can all, or sorry, there's the justice part. And there's also the kind of, is it kind of compatible with people, is it a good deal? Is it a good bargain for people? And I think this is, you know, this is often how, you know, to the extent we're kind of very concerned about AI's like kind of rebelling or something like that.

2:28:42It's like, well, there's like a lot of, you know, part of a thing you can do is make civilization better for some of it. Right. So it's like, and I think that's, that's an important feature of how we have in fact structured a lot of a lot of our political institutions and norms and stuff like that. So that's the thing I'm getting getting at in that, in that quote. Okay, I think that's an excellent place to close. Great. Thank you so much. Thanks for coming on the podcast. I mean, we discussed the ideas in the series. I think people might not appreciate if they haven't read this for years. How beautifully written it is.

2:29:17It's just like the ideas, we didn't cover everything. So there's a bunch of very, very interesting ideas. As somebody who has talked to people about AI for a while, things I haven't encountered anywhere else, but just obviously no part of the AI discourse versus nearly as well -vred. And it is genuinely a beautiful experience to listen to the podcast version, which is in your own voice. So I highly recommend people to do that. So it's joecarlsmith .com where they can access this. Joe, thank you so much for coming on the podcast. Thank you for having me. I really enjoyed it. Hey everybody, I hope you enjoyed that episode with Joe.

2:29:53If you did, as always, it's helpful if you can send it to friends, group chats, Twitter, whoever else who think might enjoy it. And also, if you can leave a good rating on Apple Podcast, or wherever you listen, that's really helpful, helps other people find the podcast. If you want transcripts of these episodes, or what you want to do in my blog post, you can subscribe to my substack at dwarkeshpatell .com. And finally, as you might have noticed, there's advertisements on this episode. So if you want to advertise on a future episode, you can learn more about doing that at dwarkeshpatell .com, slash advertise, or the link in the description.

2:30:27Anyways, I'll see you on the next one. Thanks.

From the publisher

Chatted with Joe Carlsmith about whether we can trust power/techno-capital, how to not end up like Stalin in our urge to control the future, gentleness towards the artificial Other, and much more.

Check out Joe's sequence on Otherness and Control in the Age of AGI here.

Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform. Read the full transcript here. Follow me on Twitter for updates on future episodes.

Sponsors:

- Bland.ai is an AI agent that automates phone calls in any language, 24/7. Their technology uses "conversational pathways" for accurate, versatile communication across sales, operations, and customer support. You can try Bland yourself by calling 415-549-9654. Enterprises can get exclusive access to their advanced model at bland.ai/dwarkesh.

- Stripe is financial infrastructure for the internet. Millions of companies from Anthropic to Amazon use Stripe to accept payments, automate financial processes and grow their revenue.

If you’re interested in advertising on the podcast, check out this page.

Timestamps:

(00:00:00) - Understanding the Basic Alignment Story

(00:44:04) - Monkeys Inventing Humans

(00:46:43) - Nietzsche, C.S. Lewis, and AI

(1:22:51) - How should we treat AIs

(1:52:33) - Balancing Being a Humanist and a Scholar

(2:05:02) - Explore exploit tradeoffs and AI



Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe

More from Dwarkesh Podcast

All 94 episodes
Joe Carlsmith - Otherness and control in the age of AGIDwarkesh Podcast · 2 h 31 min
Listen in VO