In short
Podcast Episode Summary: Win-Win with Liv Boeree - Episode #42
Episode Title
Anthony Aguirre and Malo Bourgon - Why Superintelligent AI Should Not Be Built
Episode Description In this episode, host Liv Boeree engages in a critical conversation with Anthony Aguirre and Malo Bourgon regarding the implications of developing superintelligent AI. They discuss the risks associated with AI development, the concentration of power, and potential geopolitical ramifications. The conversation emphasizes alignment, governance, and innovative solutions to ensure technology serves humanity positively.
Key Themes
- Risks of Superintelligent AI: The primary concern raised is the potential for AI systems to surpass human intelligence, leading to uncontrollable scenarios that could jeopardize humanity's future.
- Power Concentration: Aguirre and Bourgon highlight the dangers of AI power being centralized within a few entities, posing threats to democracy and societal structures.
- Market Self-Regulation: The notion that the market can self-regulate AI risks is challenged; the speakers argue for more proactive measures in governance and regulation.
- Win-Win Solutions: They explore methods to create beneficial frameworks for AI development that align with human values and societal goals.
Key Takeaways
- Different Categories of Risks:
- Existential risks resulting from uncontrollable superintelligence.
- Societal and economic disruptions due to job displacement caused by AI.
- Power imbalances created by concentration of AI capabilities.
- Why the Market Can't Self-Regulate AI:
- The complexity and potential consequences of AI systems exceed the market's ability to manage risks effectively.
- Individual companies may prioritize profit over societal welfare, leading to harmful outcomes.
- Geopolitical Implications:
- The race for AI development may escalate tensions between nations, potentially leading to arms races comparable to nuclear proliferation.
- Proposed Solutions:
- Compute Governance: Implement regulations around powerful AI hardware and ensure transparency in AI development.
- Open-Source Safety Measures: Encourage collaborative efforts among nations to create safety protocols for AI deployment.
- International Treaties: Establish legal frameworks to manage the global implications of AI technology.
Chapter Breakdown
- 0:00 - Intro
- 4:08 - Different Categories of Risks: Overview of the various risks AI development poses.
- 10:51 - Why Can't the Market Manage AI Risks Sufficiently?: Discussion on the inadequacies of market self-regulation.
- 15:33 - The Dangers of Power Concentration: Examination of how AI could centralize power and threaten democracy.
- 22:52 - What Solution Space Exists?: Exploration of win-win scenarios and potential solutions.
- 32:55 - Should Compute be Regulated?: Arguments for and against regulating computational power in AI.
- 38:38 - Evolutionary Nexus: Discussion on whether humanity is at a pivotal point regarding AI.
- 42:07 - Decentralization vs. Centralization: Evaluating the risks and benefits of both approaches in AI development.
- 46:46 - Nuclear Power & AI: Drawing parallels between AI development and nuclear power regulation.
- 51:06 - Wrong Stories Attract the Wrong People: The importance of narrative in shaping societal attitudes towards AI.
- 57:39 - Coherent Extrapolated Volition: Discussion on aligning AI with human values.
- 1:05:59 - Can Win-Win be Embedded into AI?: Exploring the feasibility of integrating win-win principles into AI systems.
Conclusion The episode serves as a crucial discourse on the ethical considerations surrounding AI development. Aguirre and Bourgon's insights emphasize the importance of proactive governance, public awareness, and collaborative efforts to ensure AI technology benefits humanity rather than poses existential risks. The conversations encourage listeners to think critically about the future of technology and its implications for society.
Links and Resources
- [Keep the Future Human Paper](https://keepthefuturehuman.ai)
- [Future of Life Institute](https://futureoflife.org)
- [MIRI](https://intelligence.org)
- [Superintelligence Book by Nick Bostrom](https://www.amazon.com/Superintelligence-Dangers-Strategies-Nick-Bostrom/dp/0198739834)
Credits
- Hosted by: Liv Boeree
- Produced by: Luca de Vico
---
This summary captures the essence of the episode and provides a structured outline of the discussions and key points made by the guests, making it easier for readers to grasp the complex issues surrounding superintelligent AI.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00If you give every person on earth like a magical genie that is superpowered that they can just command and will do whatever they want. You know what happens when superheroes fight? You get like New York City creamed every time, right? We're talking about systems that even setting aside the risks are going to make all cognitive labor economically redundant if we succeed at making them. And that will create a change in the world that will make the Industrial Revolution look like small beans, but instead of happening over the course of 1 to 200 years, might happen over the course of 5 to 10 years.
0:28Hello, friends, and welcome to this very special Win Win episode. This is a recording of the second ever Win Win IRL event that we did here in Austin a few months ago on the theme of AI. I'm speaking to Anthony Aguirre and Malo Borgon, who spearhead two of the organizations who pioneered AI alignment research. These guys have been working on the problem of how could we actually control or at least steer something that is super intelligent to humanity. They've been working on this for nearly 15 years. And I wanted to talk to them because some of their viewpoints are quite controversial, certainly in technology circles.
1:07And in fact, Antony has just written a very interesting paper called Keep the Future Human, which makes the argument that we should not be trying to build it at all. I wanted to really sort of examine their ideas and understand why they believe that this technology could be so dangerous. It was also a cool opportunity to allow the guests in the audience, many of whom are regular win-win fans, and many of whom you might actually recognise, to also ask their questions too. So yeah, let me know if you like this format, because we want to do more events like this. So maybe drop a comment if you would want to attend one of these in the future.
1:41On that note, let's dig in.
1:45So yeah, maybe to get us started, can you just sort of explain both what it is, What is the primary area of work that you're both focused on? And sort of what are your core arguments or theses right now? Great. Yeah, I'm Molo. I run the Machine Intelligence Research Institute. You know, there's a bunch of ways that I could describe it. I don't love the term doomer, but a short explanation is we're one of the OG doomers. I don't think the term doomer is quite correct because most of the people who get labeled as doomers are often the people who imagine that if we get things right with this tech, that it'll be, you know, radically, transformatively positive for the future of humanity.
2:18And that's the thing that we're ultimately focused on. So yeah, Miri's been around for about 25 years now. I've been working there since 2012, so quite a while. Historically, a lot of Miri's work has actually been on technical work. You know, we did some of the early work laying the foundations for the field, which people now call the field of AI alignment. And more recently, we've transitioned away from doing that work, mostly because although we felt that we're making some slow incremental progress, it was kind of very high level insight-bound research and the march of capabilities was progressing much faster, it seemed, than our ability and the field's ability to come up with some of the important solutions and understanding of how we'd actually be able to build the tech in a more safe, understandable, transparent way.
3:02And so now most of our focus is on communicating. After the chat to be team moment, the world is paying attention. Part of me is like, oh my God, I've been training for this for the last 12 years. And so, yeah, we're spending our time trying to talk to the world about it, doing some more policy engagement and doing some governance work to try and figure out what the world can do to actually avoid the risks and get the benefits? So my background is theoretical physics, and this is helpful because I take abstract arguments seriously. And I think this is one of the issues with AGI and superintelligence and talking about those things.
3:34We don't have them yet, and yet we have to take, we have to somehow reason about them because if we wait until we have them, it may be well too late to do anything important about it. So I've switched from theoretical physics slowly over the past 10 years to thinking about AI and both its risks and benefits, now head the Future of Life Institute. And we have that dual mission, not just for AI, but other transformative technologies of what do we do now to make that transformation that is coming at one level or another as positive as possible and avoid the large-scale risks while doing so. So can you take us through what these different categories of risks are?
4:12Because it seems like some of them almost like, they're almost like lying at different ends of the spectrum. And there's many different categories, some are much larger than others. But so can you just try and sort of lay out the landscape for us? I think the basic thing that I worry about, which, you know, as I'm talking about it to people for the first time, I might just kind of lay out as the what I call the common sense argument for the large scale catastrophic extinction risk level things, which is, you know, dating back to the start of the field of AI, the goal was to make computers that can do the thinky thing that we do.
4:42And the thinky thing that we do has allowed us to be the ones who are steering the future of this planet. That's the thing that causes us to be able to put, you know, rockets out into the sky and send things to the moon and beyond. And that thinky thing is extremely powerful. And even back when the field was founded, folks like I.J. Goode and Alan Turing were already looking at, you know, what if we succeed here and we make something that is much, much smarter than us? Controlling that might be very difficult. Steering that in positive directions if its interests aren't aligned with ours is likely to end in catastrophe.
5:12And so, I mean, there's a bunch of details to get into on that particular risk, but I kind of just like to lead with that, where I think there's kind of this default case that I hear from some people, which is, well, unless you can tell me a very specific story about how this goes wrong, then this all just sounds like crazy nonsense to me. And I'm like, from this pitch, I don't think that we should all just agree that obviously the risk, you know, everyone should just automatically agree that there are huge risks here that are kind of the default risks, but it seems like at least it should be a foundation for being like there is an important thing here to consider that seems real difficult to potentially navigate.
5:45And, you know, based on the current paradigm of the systems we build, and we're in a place where we have very little understanding of how these models actually work. You know, we're not coding them. We're almost growing them. And they result in these trillions of floating point numbers or hundreds of billions that get multiplied together, making systems that are built that way to be vastly more intelligent than us is a thing that I think probably does not go well by default unless we figure out a way to understand them and steer them in directions that are aligned with human interests. What are some other types of risks aside of the sort of classic losing control of, you know, a more powerful being?
6:21Yeah, I think there are underappreciated risks even... Well, yeah, so, I mean, we think of it as, like, job one, curtail the intelligence runaway that ends in doom for everybody. But then there's like lots that is left because even if you don't have the runaway, you still have extremely capable systems that proliferate lots of powerful capabilities to lots of people. For the most part, that's good. And also it is dangerous. And so we have to have a society that is resilient far more than our current one is to things like bio risk or proliferation of cyber capabilities or like auto blackmail and hacking or, you know, filing a billion lawsuits at the press of a button.
7:03Like there's just many things that have used to be hard and require a whole lot of human expertise and will become very easy. And for the things we want, that's great. And for the things that we don't want, that's really bad. And so when you think of so you can you can make a long list. I think at the top of the list for most humans will be that even if we don't have a runaway to superintelligence, if we build AGI at a very competent level, it's going to largely replace them in their work by default, especially if it happens very quickly. So lots of people will be at minimum displaced and depending on how far and how competent those systems get, largely replaced.
7:41And then the information, like if most of the information on the internet or everywhere else, you know, is not just chosen for us by AI systems with some misaligned optimization, but also written by AI systems that are under control of some corporation with some agenda, that's also not great. if like all of our democratic institutions are basically, you know, at the mercy of our understanding of the world, that is at the mercy of the AI systems that we're interacting with and the news ecosystem that we're interacting with, that's not great. So if most of the, you know, information and power and wealth is running through AI systems run by one or two or three companies and maybe one or two governments, that's a level of power concentration that seems dangerously not good for most people so there are lots of those i think the but they a lot of them stem from like i think a lot of those are in principle manageable though like if you let me put it this way i think there are people who believe that ai is not going to be that big of a deal like it's going to be another technology like the internet or maybe like even crypto or social media, but it's going to be another technology that's going to be a tool for human use.
9:01We're going to get used to it. It's going to cause some job displacement, but it's going to make more jobs than it destroys. And ultimately, it's going to just boost the economy and let us do all kinds of stuff that we couldn't do before. And it's not going to be that super dangerous. Those people on our current trajectory are wrong. There's another set of people who think that AI is going to get stronger. It's going to have not just intelligence and generality, but autonomy. It's going to combine all of those and things that can be slotted in one for one for humans in their work. It's going to be able to be slotted in one for one for AI researchers, which means that the production of new AI research is going to accelerate and they're going to get smarter.
9:40And it's going to quickly lead to superintelligence, which we cannot control. and that's going to be possibly wonderful, but most likely very bad for most of humanity. Those people, I believe, currently are correct. The problem is we would like those first people to be right and the second people to be wrong. So I think our goal is to make the first people right, to make AI like a great empowering tool that is just another technology that lets us do all kinds of great things that we couldn't do before, but isn't a replacement for us individually, isn't a replacement for humans as a species, is just another in a long line of great tools that we've invented.
10:16That would be my optimal outcome. And I do think that is possible, but we would have to really, really decide to do that. Why can we not just trust the market to manage these types of risks? You know, like, generally speaking, I think most people would agree that it seems like the market tends to produce things that we want, right? it's yes there are externalities sometimes but the market has produced all these cool things that we have in this room and our lives are pretty great compared to how they were you know 100 plus years ago so why is this case different yeah i think again it it is insofar as ai is another technology that we're developing i think insofar as it is a like full replacement for people that is not going to be controllable, that does not apply.
11:05So I think there's a fundamental difference between AGI and superintelligence versus the sort of AI tools that we have now and that we could have if we go down that path. What is that difference? It's really how much of the full human capability we build in. So the way I've been thinking about it is not like AGI is better not thought of as artificial general intelligence, but autonomous general intelligence, because the key sort of the trifecta of capabilities that humans have is we're like very intelligent at doing particular tasks, you know, and we have AI systems that are superhuman at particular tasks like at Go or folding proteins and like even identifying images and whatever.
11:49We have superhuman narrow AI, but humans are also very general. And now we have general AI that is as general as humans in a lot of ways and even more so. No human speaks 100 languages and has 10 whatever they are, bachelor's or master's degrees or whatever, or 100. So they're more general in a lot of ways. What they're not very good at is autonomy. Humans are really good at autonomy. We get stuff done, not always so effectively, but we are very autonomous systems. And that trifecta is what allows, that combination is what allows us to really be in control of the world. I think if we allow systems to have all of those properties at human or above level, then there's just not, like, that is where the danger zone lies in terms of their capability to, you know, self-improve and sort of run away from us.
12:38Like, that's where their uncontrollability lies. And that's where the ability to, like, fully replace people rather than just be tools lies. So I think that that's the key difference. Like, do we forcibly make AI a tool like previous technologies, or do we allow it to become more like a second species that's more like a biological thing that reproduces, that improves itself, that is agential, that isn't really controllable? We have a good analogy for all of those things. Those are biological systems. And if we really go along those lines, but they're also super smart and super general, that's not going to be good news for us.
13:13But we can make them as tools, but we will have to decide to do that. Why do you think it's a given that something won't naturally... I think the most common pushback I hear to this, and I can sympathize with it, is that, well, if something is super intelligent, then it'll figure out, Like, why would it see us as a thing to compete against? It's going to unlock types of abundance that we could never have imagined. So why would we assume that it would see us as some kind of actual threat in any way that would affect it? I mean, it doesn't have to. You know, I think we should be careful with analogies, but we can look at ourselves as a species and the other species on the planet.
13:55And often people conceptualize some of these big loss of control risk scenarios as, you know, the AIs are evil and, you know, there's going to be some sort of Terminator situation where they try and take us out for some sort of malevolent purpose. But, you know, I don't think that's how we've oriented to most of the species on the planet. But some 10 ,000 plus of them are extinct, not because we went out of our way to try and make them extinct, because we had things that we were trying to do in the world. They happened to be in the way. And so we were very good at accomplishing our goals, at radically transforming the world to accomplish the tasks and to do the things that we wanted to do.
14:31And they were collateral damage. And so when I'm thinking about kind of this big picture loss of control risk, it really comes down to the question of if we bring these very powerful superintelligence systems in the world, if they don't care about the things we care about, if they don't care about our flourishing, then if there's kind of any misalignment with that, then we should be worried about the consequences of them pursuing the things that they end up pursuing.
14:57on the other end of the sort of loss of control you have the idea of excessive centralization of power i mean we touched on it briefly you know i've heard some people say that oh well you know if we just sort of open source and allow everyone to sort of you know it's allow a race to agi but also have open sourcing then that's going to be sufficient to make sure make sure the power does not concentrate that said anthony in your in your recent essay which by the way everyone should read it's it's called keep the people keep the future human and one of the lines you one of the claims you make in there is that actually agi sort of by default will tend to absorb or seek power more than distribute it so it would also as well as you've got the sort of the the risk of sort of like techno anarchy with all like more empowered bad actors and so on but and the possibility of runaway ai you also have this it's also amplifying the the risk of centralized control so can you that's not quite clear to me why yeah well i think like super like agi or super intelligent systems are really dangerous if they're under control or if they're not under control if they're dangerous if they're under control because people will have control of them and people will do bad things and like the the ability of you know give if you give every person on earth like a magical genie that is super powered that they can just command and will do whatever they want.
16:13Like, is that going to be a good world? I don't know. It's going to be like pretty, they're going to send their genies out to like make a lot of money and defeat their enemies and all these things. And you know what happens when superheroes fight, you get like New York City creamed every time. Right. So, so I think like having people duking it out with the super intelligences that are under their control is quite dangerous. Having super intelligences that aren't under control, also super dangerous, right? Because those, as Mala said, if, But unless you've done an amazing job at having those things fully and deeply understand exactly what humanity wants at any given time and whatever that could possibly mean and give it to you, then the huge amount of power that they're wielding is going to end up with all sorts of things happening that humans don't want.
16:57Now, in terms of like going a little bit below superintelligence, as we start to build more and more capable AI systems, there are sort of a couple of different paths. They like we could lose control to them because they decide, well, I would really rather do that other thing rather than the thing you're asking. Or we could give them control because they're just going to be better at doing those things than we are. So like if you have a system like suppose I have my I want to run a company, you know, I'm a CEO. So all the other companies have their AIs and their CEOs and their AIs are really smart and their CEOs are giving them advice.
17:36and I'm my CEO and like I can take advice from my powerful AI system or I can just make up my own mind. So what am I going to do? It's smarter than me. All my competitors are taking that advice. So I'm going to have to take advice from my AI advisor as well. And like it's more effective than me. It's competing with them. It's got to make some decision, like it's got to make, give some advice really quickly. And I'm going to have to decide really quickly because all the other people are getting advice and deciding really quickly. So I'm not really going to have that much time to like think through and have the AI totally explain everything to me so that my puny human brain can understand why it's a good idea to short this stock and buy a whole bunch of widgets from these guys.
18:15So I'm just going to start pressing like, yeah, like let's do that. Yes. You've given me good advice. Yeah. You're making me a ton of money. Yes, yes, yes. So at some point I'm just going to realize like, am I really running the company? Like, yeah, I'm formally in power, but everything I do is told to me by this smarter than me AI system. So that's the sense in like so at that point the ai really is in charge like formally i am i'm formally the ceo and like technically i could turn off the ai but will i like no because my company's going to go bankrupt my shareholders are going to be like why did you turn off your c your ai dude you're fired so the ai is really in charge of the company and so that's the sense in which like if it's more capable than us as we're in a competitive environment and have to make decisions we're going to be either having that control taken by the ai systems or give it to them willingly and Either way, we lose it.
19:03I mean, what I'd add to that is you started by bringing up kind of the open question. I do want to acknowledge that I feel like I have some fairly large sympathy for the people who are, you know, very supportive about continuing to open source models. I think, you know, currently with the, you know, models that we have now and in many worlds, depending on the capabilities of models, there's enormous value to diffusing that capability throughout the world. There's a ton of people who could be doing good for the world, designing better products, incorporating them into all kinds of research that if they can have the models themselves, they can use them with the data that they have.
19:38Like there's an enormous benefit there. There's also an enormous benefit from the kind of research and safety side to the extent to which there are more people who can get their hands on models that are close to the frontier. They can be contributing to understanding those models. You know, I think that the challenge of actually doing interpretability work, for example, to actually really get an understanding is hard. But I'd certainly love for more people who thought they had new ideas to try to be able to try that. The thing that I think is, you know, an unfortunate, difficult challenge that we're going to have to face that is going to be very uncomfortable is that even though there are those benefits, as models become more capable, even if we're not, you know, going well past AGI to superintelligence, they will also be able to do very powerful things in the world.
20:19They will be able to help malicious actors, you know, do more damage, you know, CBRN threats, you know, chemical, biological, radiological, nuclear, cyber, that kind of thing. They could also be just very societally disruptive. You know, one of the things I feel like I spend a lot of my time when I'm talking to folks about this is, you know, I'll talk about the loss of control stuff. But I feel like there's kind of a missing mood when even among people who are kind of starting to take these ideas seriously of when I hear them talking about it, like, well, you know, if we do it right, we'll get better drugs and we'll have, you know, cleaner forms of energy and we'll have to deal with the risks.
20:51And we can talk about the offense, defense balance of cyber or whatever and get all those details. But I feel like there's there's a missing mood of like we're talking about systems that even setting aside the risks are going to make all cognitive labor economically redundant if we succeed at making them. And that will create a change in the world that will be, you know, make the Industrial Revolution look like small beans, but instead of happening over the course of one to 200 years, might happen over the course of five to 10 years. And we're not kind of responding in a way. There's kind of not this orientation towards what might be about to happen.
21:23It's not for sure. I'm not sure that this will happen in five years or two years. It could be 20. But I think it's coming and there's kind of this missing grappling with how radical things might change, how quickly, even setting aside the risks. And I just try and impress that upon people and think about like if we just diffuse this in an open way as quickly as possible, there are an enormous amount of ways in which that could go wrong. And we have to be careful and think about as capabilities increase what we're actually doing here. So I'd love to hear you guys explore what kind of solution space there exists, because, you know, like the name of the show is Win Win.
22:01You know, so I always try to find like, what are the win win solutions where we can, even though I agree that like right now, to me, it seems like the near term risks are vastly outweighing the potential near term benefits. But many people might disagree. And that's fine. And please say so when we open this up. But what are some possible solutions are there that, as best as you can tell, allow us to minimize the worst risks but also maximize the benefits, of which there are obviously many? Yeah, I'd say a lot of it is sort of a motivational question of what do you build an AI system for? Like I'm a big believer in when there's a problem that you have that you would like to be able to solve and humans can't solve it, let's build a new tool that can help us solve that problem.
22:48So if we can't solve protein folding, let's make an AI system that can solve protein folding. Like everyone agrees this is great. Like there could be some problems with like proliferating the ability to fold proteins and figure out like protein interactions. That's definitely something we have to keep an eye on. But I sort of have confidence that like as we that that as a species, we've dealt with like fairly powerful tools before. If we're just a little bit smart about it, we can deal with like steadily more powerful tools of this type. The similarly, like in some of our social interactions or with our news, like right now, AI is making all of our you know, we've got the whole information gathering and aggregating and understanding system of society.
23:30That's totally fucked. Right. Like we've got AI algorithms that are choosing what we what we are like given on news feeds according to an optimization algorithm that is not built for us, like for societal benefit or better understanding of things or anything else, but like for corporate profit through engagement. And like now we have lots of the content actually being written by AI with the same incentive structure. Like how much of the news we consume is written by AI? Nobody knows. It's like literally an unknowable number. Amusingly, I Googled how much of current AI, how much of current like news stories are written by AI.
24:10And I got every time I do this, I get the same article from LinkedIn written by AI that says 10%. There's no citations. There's like no evidence for this. It's just like 10 % from like, so this is where we are. Like we don't know. And we're not going to know. We're like, we're never going to know unless we do something quite different. Now, can we do something different? Absolutely. like ai could absolutely be used to build a beautiful epistemic system that we could all trust like where we look at a news article and we have an ai system that will tell us like oh here's the or like let's just probe into the origin of all these different claims that are in this thing here are the things that everybody agrees with here are the things that are like controversial the things that everyone agrees with this is why they agree with them and you could like delve down all the way back to like data taken in a scientific experiment at a place in time or a photograph that is like cryptographically registered to the blockchain and like is an actual photograph taken by this camera in this place.
25:02We could do all of that. And we could have our loyal fiduciary responsibilities to us, AI assistants, like check all those articles for us if we wanted to, to say like, yeah, these are legit. And we could have new sources that have those capabilities that we trust and other ones that don't have those capabilities. We're like, why would I ever read that? We could do all these things. Like we're not building those sorts of tools. We're instead like just making them all crappier by dumping AI into things where it's not actually helping. So I think like the critical difference, I think, for me is can we identify the things we actually want to do and then build the AI systems to do it rather than build the general AI capability and then just jam it in wherever it makes things like more convenient or follows local incentives?
25:45Because that's what we're doing right now. And I think we can hugely win if we do that first thing. Yeah, I think I agree with some of that picture. I think the challenge is having the world have that orientation seems difficult to me when there isn't a great appreciation of the risks and what the alternative might look like. I think we do have examples in history where humanity, you know, society has looked at a certain thing and decided we don't want to go there. So human germline engineering. It's not like there's this national surveillance, you know, or international deep surveillance program in order to stop anyone from doing this.
Read the full transcript
26:21But, you know, the scientific community kind of agreed that doing this, if any one actor unilaterally did it, that this would kind of be fundamentally changing what it meant to be human in a way that would be like a radically not cool thing to do unless somehow humanity decided that this is a path that we wanted to go down. And this formed a taboo that is like very well enforced. As soon as anyone does this type of work, you know, funders won't support it. like, you know, even in China, for example, there was an example of someone who tried to do this and they were immediately kind of shunned from their institution.
26:53And you're shunned. Yeah. And this is, but this is the case where there is a clear, understandable thing that the scientific community rallied around doing. And if we got to a place with AI where there is a big enough appreciation for the risks, I think there's a bunch of challenges we actually have to probably deal with in terms of people still not trying to do that, the large scale risks that I'm worry about and push systems in that direction. But it certainly would be easier if we could kind of have that shared understanding. And I think that's kind of still what I'm pointing out with like the missing mood is that not building these smarter than human systems that pose this risk, certainly until we're not, until we have a much deeper understanding of how we would do so safely and can proceed down that path, in some sense is very easy.
27:34I think that given the current incentive landscape, you know, there are many actors that seem like they're trying to goad China into a race to build a decisively strategic technology. And I'm just like, why are we doing this? China doesn't seem like they're trying to race us. Why are we trying to goad them into this race? So I think, in principle, the solution is easy. In practice, I think, coming to enough of a shared understanding that we could actually not do the things that are dangerous and set up the kind of what I think would be required, the verification enforcement mechanisms to actually supervise this kind of work and make sure we were working on the pro-social, easier to understand, more tool-like AI systems is the path we should go down.
28:11But I think that right now, given the state we're in, it looks extremely challenging. That was your optimistic answer? Yeah. Well, my optimistic answer is kind of like, I don't know, man, the nuclear weapons thing seemed like a big deal. In the Cold War, everyone thought that maybe we were going to die in a nuclear holocaust. And we got to the point where at some point, Gorbachev and Reagan looked at each other and went, what the fuck are we doing here? Like, maybe we should actually just take a step back because this is kind of crazy. And I don't know how we get to that point in this world, but this is a place where I think we need to get to, where as capabilities increase, I think this will be more salient to people.
28:47But I think it, I find it hard to imagine how we get to that state without more of this realization happening with our senior decision makers across the general populace, et cetera. But that's why it's easy, because you can have that moment. But I think the nuclear arms reductions is really instructive because I think that largely why that really started to happen was the realization of nuclear winter, which meant that there isn't just mutually assured destruction, but self-assured destruction. Like even if the U.S. sends all its missiles to Russia and Russia just sits there and does nothing, basically everyone in the U.S.
29:19still starves because we have nuclear winter. And so this fundamentally changes the dynamics of the situation. And I think similarly with AGI and superintelligence, the thing that we have been really trying to emphasize is not that it's this powerful thing that's going to have like superpowers and super risks, because like, frankly, I think U.S. military folks are comfortable with super risks. You give them super powers. I think what we want to emphasize, because I think it is also true, is that there will be a certain point at which you will not be in charge of these systems. You will lose power to these systems or of these systems.
29:49You're not going to be in control because you're not building tools. You're building like a replacement or a like uncontrollable thing that is smarter than you. No one wins a race to be the first one to lose control. Exactly. No one wins in a race to be the first one to lose control. So if that is really realized and gotten by the U.S. side and the Chinese side, neither side wants to build the thing that they're going to lose control of or to. And so they can agree. And there's still a tricky game theory problem of like, does China really get it? And like, do they really get that we get that they get that we get that they get it?
30:21But we did this stuff before. We did it with nuclear weapons. Like the thing that gives me the most optimism is looking back at the writings of the smartest people thinking about the state of the world. Like at the end of World War Two, there was a nuclear bomb. The Russians got the nuclear bomb like Einstein and Teller and von Neumann and all of them were like, it's over. You know, they've got the bomb. We've got the bomb. Everyone's going to have the bomb soon. There's going to be a race to get more and more bombs. at some point, like, wars happen and everyone's going to die. And they were super pessimistic.
30:55And they had every reason to be. Like, that argument is solid. And yet, here we are. Like, somehow we figured out the game theory that, like, plus, like, a lot of luck, let's be honest. But we're still here. We managed to muddle through. And so that gives me optimism that even when things seem, like, quite intractable and, like, the natural dynamics of it just lead you down a very dark path. Sometimes we work it out. And so we just have to believe that we can. Yeah. I mean, one of the things, it's funny that we're talking about nuclear, because in the first iteration of Win Win Live, it was with Isabel Bomeker, who's here, where we were talking about what we could learn from the nuclear industry about AI.
31:32So it just seems like all roads lead back to like, what can we learn from this? Because there are so many parallels. One of the things you advocate for in your paper, Keep the Future Human, is essentially that we should start thinking now about, because again, with like one of the things that worked with nuclear weapons and sort of controlling them was like keeping track of the uranium. And arguably, in this case, again, it's like uranium is a dual use technology, a dual use material that can be used for amazing stuff, nuclear energy, but also obviously very bad things, nuclear bombs. Again, compute.
32:06It's the substrate. It's where this energy comes from. And again, it can be used for amazing stuff and for big problems. So you actually advocate that we should be looking into doing compute regulations can you talk a little bit more on that because i imagine that's something that might like it's got like an ick factor to it and like why is it why why should we not have an ick factor so i think it shouldn't like sorry thanks so you know if i leave here and leave my phone on the chair you know one of these nice souls will just find me and like give it to me but suppose they were a little more nefarious and just like took it home with them i would get on my laptop and be like deactivate my phone and apple would deactivate my phone and it would be a brick that they can't use.
32:45So we're perfectly comfortable with the fact that we have in modern security architectures, we can have things where having possession of the hardware doesn't mean you can do whatever you want with it, right? There are licenses between a piece of hardware and the company that made it that can allow it to be used for certain things and not allowed if you're violating the terms of service that or whatever with that company. And we could absolutely have the same sort of thing with the, you know, if it's doable for a $500 phone, it's doable for a$20 ,000 GPU. So I think there are all kinds of capabilities that make compute like these very specialized pieces of AI hardware, like similar to, but way, way better than uranium.
33:25So uranium, the crucial thing in nuclear, both power and weapons, is like a fairly scarce resource. It's hard to enrich, you know, it takes technical expertise and the raw materials. But it's not like that hard. Like I feel like if you gave me a billion dollars, I'm a physicist, I could probably figure out how to enrich some uranium if you didn't pay attention. I know that I could not build an H-100 graphical processor. That's like the pinnacle of all of human civilization at this point. Like our best technology fitting 100 billion transistors onto this tiny little wafer. Not in a million years could I do it without like a whole civilization behind me.
34:01There's literally one company in the world who knows how to build the machines. That is the start of the process of doing this. Even though other hundreds of billions of dollars, like tens of billions of dollar companies are trying super hard to do it, they still can't. So it's incredibly hard. And there's this really tight supply chain of like the one company that builds the wafers and the one company that makes the machines that build the wafers and a couple of companies that design the chips. So it's really tightly constrained and much more so than uranium would be, but also much more fine grained because uranium, if somebody gets ahold of it, they can do whatever they want with it.
34:30A GPU, if you've set it up right, They can do what is allowed to do with it and not what's not allowed. And I think that's just, frankly, as we get systems that are powerful enough that the analogy starts to more and more hold, that GPUs are granting the same sort of power and risk that uranium used to, we're going to have to treat them more like uranium, and we're going to have to use these capabilities to track where they go and what they're doing. Not all chips, like lots of chips, totally fine, doesn't matter. But these super duper specialized ones that cost$20 ,000 a pop, yeah, we should be keeping an eye on where this is going.
35:05So how do you do that without also opening up the risk category of too much centralized control? Because like who are the people deciding these rules of what they can be used for? Who watches those people? Yeah, I think that's a very tricky question. And I think there is a very reasonable response to these types of considerations that are like, it sounds like you're arguing that we build a, you know, totalitarian world government to control all the compute. And I like totally get that vibe. I do think for the loss of control stuff, let alone dual use, misuse and having to have a way to restrict who can do what with AI seems essential to me.
35:41On that front, I think the ideal thing that we should be striving towards is that there's some sort of multilateral way in which the powers of the world are coming together to try and find a way of governing this together. and enforcing it together. On that front, I'm a little disheartened by the current direction that the current administration is going in, where I've been trying to advocate for the last couple of years, the US has a very good leadership opportunity, being the ones who have most of the intellectual property, have most of the relevant companies that they could be working with allies to start to build this multilateral way of governing the technology.
36:16And I think that's the path we have to somehow find a way to steer. But I agree, there are ways that that path goes down where there are some pretty gnarly concentration of power concerns that we definitely don't want to end up in those worlds. And those are like, part of, I guess, how I think about this is I am not an international diplomat. I don't know how to negotiate treaties between countries. The thing that I am trying to do is to try and get those people to understand what's about to happen to the world such that they are turning their minds to being able to help make these things happen because it won't happen unless they buy that this is a thing that actually needs to happen and they start turning their minds to how to do it.
36:54Yeah, well, again, in that spirit of sort of decentralization, I'll love to hand over my microphone to whoever's got questions. Yeah. Okay, so that was a lot and it was really good. But so one of the earlier topics that both of you covered was this concept that we were at a nexus point in our species evolution, potentially, with AI. And I think we're approaching that or already there, regardless of any of like which of these three solutions that were discussed occurs, because whether we're using AI as a tool. whether AI threatens to replace us or whether AI simply just goes out of control and spawns into a new species on the planet, we are, in any of those scenarios, going to unequivocally change in ways that we have not anticipated yet as a species.
37:46And so barring the government manufacturing some sort of crisis or catastrophe to make everybody terrified of AI and kill it now, But that was what I thought of in terms of solution to stop it, which I don't want, for the record. For the record. For the record. Yeah. But assuming that we are leaning into this and it will be being developed, whether it's in some guy's basement who created his own system for computational power or something. Like, there are biohackers. You talked about germline editing. They do things in their basement. They figure it out. we would essentially be creating that future even if global regulation were to somehow work is kind of where my brain goes and then so assuming that ai continues to develop and assuming that in any of these paths we're going to change why not lean into it what what is stopping us from hybridizing with the ai and and and leaning into our our our speciation nexus we could call it that was something that popped into my head that wasn't suggested and i was surprised that nobody went there.
38:44So, yeah. I mean, I would certainly feel much more optimistic about a future where we treated these AI systems more like tools, didn't race towards superintelligence, and focused on ways to increase our own intelligence so that at some point we might actually be able to come up with better technical solutions to the hard problems to actually be able to realize some sort of future. You know, I feel like when we're speculating about what this could become, this gets pretty wild into what the, you know, glorious transhuman future could look like and whether some of us are uploaded and John von Neumann probes going to the stars and some of us are just living as subsistence farmers on Earth because that's the thing that ultimately makes us happy.
39:26It's like kind of hard to to know what exactly that world will look like. But I do think that focusing on having our human intelligence be at a level where we can maybe actually have a chance of solving these hard problems. I'm not even convinced that we need that. That would be actually just, you know, the amount of smart people in the world who've really looked at these hard technical problems and tried to solve them is very small. That if we had the time to do that, that we would figure out whatever solution it is, whether we're somehow merging or whether we can actually be safe, super intelligent systems that help shepherd in the future that we would all want.
39:59I think the main constraint here is that we need the space for it, that we need the time to actually do that reflection and figure out how to do that safely in a way that will benefit humanity. And so it's less about exactly what direction, more about having the time to make the right decisions and to do it wisely. Should we wrap? Yeah. I guess I'm a bit of an outlier here. I basically am not at all scared of centralization, but I'm terrified of decentralization. Because when you look at what currently happens, like all the actually nuclear being a great example, actually. We went through the centralized thing with like, you know, however many oligarchical type people.
40:35And they worked out the game theory and worked out okay. But when you have like 10 million or a billion or 100 million, 10 million, 100 million, a billion, 10 billion people doing things, and they all are super powered, the number of those people who are like just purely antisocial with bad intent goes up to 100%. And so how do we think about even like the tool case actually scared or scarier to me than the super intelligence case, because empowering people with like, we'll call them asymmetrical offensive systems, or asymmetrical offensive tools seems like a much more salient and also legible risk category than centralized super intelligent aliens.
41:11because we already have this. We have this with TikTok. We have this with basically mass spam. Like there's small versions that's already in the wild that we see all the time today where we have AIs that are largely misaligned. And so scaling that up seems both more relevant today and also a problem that we haven't solved. How do you guys think about that? Yeah, I think it's fair. I mean, the really tricky part is how do you centralize the risk control and decentralize the economics and the power? Because those things are often linked with each other. Ideally, we would like to have all the good distributed and all the bad concentrated and controlled and prevented.
41:52Now, insofar as they're just purely empowering tools, that's very hard to do. So can you, insofar as they're dual use and insofar especially as they're on-vents dominant. So I really think this has to be treated differently for different things. Like take cyber capabilities, for example. Like I think there's almost everyone I talk to agrees that like in the, if there are lots of tools that are just really good at doing hacking in the long run, this is great because like we, we use those tools. We figure out like we, we probe for all the vulnerabilities in our systems. We patch those vulnerabilities.
42:28We make them like totally iron tight. So that's, well, in the longterm, I think it's defense dominant in the short term. It's terrible, though, because the people who you suddenly give all these offensive capabilities, a couple of companies can patch them up because they have those and they can test their own systems and make patches. But most people will not patch them. So it's a huge vulnerability for a while. So I think there are things like that where long term, it might be good. Short term, it's going to be a disaster. Bio is going to be very, very hard, right, because it's just like very offensive dominant for a long time, but not forever, I think.
43:00So I think we can, it's just that the defense against it is not going to be the same as the thing that allows the offense. I think you don't fight offensively scary AI systems building bio stuff with defensive AI systems doing bio stuff. You fight it with all kinds of like decentralized tools that instantly like monitor the air quality, the air and things and sequence things. And like, maybe there's some AI involved in like aggregating all that data and noticing, oh, shit, there's a new like thing that we've never seen before that looks pathogenic that is like arriving here. And now let's instantly like airdrops and PPE into that place.
43:34And like we need some like very sophisticated defensive system that we do not have right now. But ultimately, we will need to have that because like, you know, there's no way that like it's very hard to see how to the indefinite future we're going to have the kind of awesome bio tools to solve cancer and like build all kinds of amazing biotech without having proliferate the capability of doing those negative things. So I think there's no getting around the fact that we will have to build powerful defensive technologies to deal with that proliferating capability of the tools. It's just like, I think there are different cases and there are some where AI is the solution, where AI is the problem, and some where non-AI stuff is the solution, where AI is the problem.
44:19And it's very complicated. But yeah, totally agree that it's hard. I think bio is actually a great example where a bit more time and taking it seriously is something where if it goes really fast, then we have a problem with bio. But if we take it really seriously and build the defenses and et cetera, then we actually can create a world where we can now have like extreme bio like abilities happen and it's safe. So, yeah. Just a couple of thoughts and a couple of questions as well. So forgive me if I'm that annoying person that's just making comments. But when we're talking about the connection between nuclear and AI, I think one of the main differences is that nuclear fission was obviously discovered in 1938 in Germany, like worst place ever.
45:03And it was immediately used, sorry, Igor, immediately used to make bombs. And so the world was introduced to nuclear technologies through the bombing of the bombings of Hiroshima and Nagasaki. So that image of, you know, the children with the burned clothes, like stuck to their bodies, crying, running from buildings. I think that was really impactful. And then also it took about 15 years for, you know, the private sector to be able to start building reactors for electricity purposes only. So I think that's the difference versus now when people think of AI, they're thinking about like chat GPT and like friendly assistance.
45:43So I guess a question for you is how do we, you know, how do we inform the public of the risk in a way that feels very tangible instead of just sci-fi-ish? And then a couple more comments as well on the nonproliferation side. In 1953, President Eisenhower delivered a speech called Atoms for Peace, which ended up kickstarting the IAEA, which is, you know, I'm sure you're familiar, an intergovernmental organization that's under the UN that is like the nuclear watchdog. And one of the ways it works is that if a country wants to embark on a nuclear energy program, they reach out to the IEA, the IEA assists them, but in exchange they get, you know, an assurance that that country is not going to pursue nuclear weapons.
46:23So I think it's a little bit easier to also oversee these countries and facilities and so on. And then the last point, I promise, is a point about the current political environment. Because I agree, and I've been making this point to a lot of people, you know, in 1991, when the Soviet Union fell, Ukraine became the third largest nuclear state in the world. And they gave up their nuclear weapons in exchange for a safety assurance from the United States, United Kingdom and Russia. And of course, now the U.S., you know, declining to assist them. We're already seeing other countries around the world saying that they're interested in acquiring their nuclear weapons.
47:02So I think we also have like a hard political environment to push for something like that. OK, I'll stop now. I mean, I got something at least on number one, probably have someone number two and three. But I mean, one of the ways that I kind of think about what I and Miri are trying to do and others, including FLI, is, you know, I think you're right that being able to more viscerally understand the risks helps catalyze action. And so sometimes the way I talk about this is I kind of expect that if we keep going on the default trajectory with the stuff we're doing with AI, we're going to get like hit over the head.
47:35And I hope we don't get hit over the head so hard that we can't recover. I expect that if we get hit over the head, that'll change a lot of our ways that we relate to this and we're taking it a lot more seriously. And I kind of envision what I'm trying to do, what FLI is trying to do, what other folks in the space are trying to do is pull forward as much of what we would be doing at that point prior to getting hit over the head. And I think it's a hard problem, like you say, because it's hard to actually, you know, make these arguments, make them, I'm convincing enough that we'll actually catalyze actions with decision makers.
48:05But I think we just have to do our best to continue to try and build consensus among the scientific community. I already know that there are a bunch of policymakers and there are a bunch of public intellectuals who take this stuff seriously, where even though it's a lot more broadly discussed, they're still not comfortable talking about it publicly. I was talking to a journalist today who is one of the few journalists who kind of actually takes these risks seriously and talks about them in his reporting. And he gets a lot of pushback even from his journalist colleagues because they just give him a lot of flack for that.
48:37So I think there's an enormous amount of work in just continuing to have this conversation to expand the Overton window, to actually have it be okay to publicly state that you actually think there are risks at this scale and that society needs to take decisive action to address it. I don't think I have a magic bullet here. The only thing I know how to do is to try and communicate. So I have kind of a weird argument which is I do think that we could easily move into a situation where we've created some kind of alternative intelligence system that scales with just electricity, and then we could lose agency over our biome, which we need to survive.
49:15However, I think what y 'all are doing to raise awareness and to, sorry, I think I'd like to ask you this question. what if it is the case that in order to govern this kind of technology, we need better humans who are actually not just grounded and seeped in the technical underpinnings, but also the philosophical and the economic, and then really have the wisdom to say, how do we give birth to a human co-hotel? To your point, actually, about buying us time to develop this in a good way, right? And the problem with driving public awareness of these things are nukes and they could kill everyone is that narrative makes them sound like they're super powerful.
49:56And that attracts to the point of the win-win, it attracts exactly the wrong kind of zero-sum assholes who want to own that power. So the more you say, this is like a nuke, whoever gets it first is going to win everything. There's one company that makes the, you know, extreme UV lithography machines to make the one chip. And there's one guy who has the one matrix multiply library that can do this. We have to own all the stuff. And you get exactly the Dr. Strangelove zero-sum scenario. And that, again, is an attractor for exactly the worst kind of people who will not cultivate the wisdom, who will not create the time or the space, and who will not put the philosopher, science, technologist, well-rounded humanists into those rooms.
50:42Okay, so my meta argument is, if we believe a finer future is possible between humans, humans coexisting with transhuman sort of level or beyond human intelligence machines, we have to do this in a wise way. Buying ourselves the time to do that, in fact, requires us to downregulate the narrative of this is like nukes, right? And just say, actually, look, we have advanced technology and tools in the world, and we credential people. We let just about anybody drive a car, but they have to pass certain kinds of tests, including eyesight tests. We allow a lot of people to fly airplanes, but they have to pass certain psychological evaluations.
51:22We allow some people to do certain kinds of things. We allow some people to go and govern and check nuclear reactors and these BSL level 4 and 5 kind of labs. And we allow scientists and technicians to get involved in these things because nobody says this will. If one guy makes this one thing, they will own the world. We don't have that kind of narrative. So the population doesn't understand how any of this shit works. And so they basically have to be okay with ceding that power to other people that they may not even like or know or think are too nerdy and weird. So I think there may be actually this kind of two-degree move thing that we have to think about if we really care about AI safety.
51:58So that's my question for you guys. Okay. A lot to agree with there in the sense that I exactly agree that insofar as people feel like there's a win-lose scenario, they're going to want to win. and that means they're going to want to get the powerful thing so that they're the winner. I think my and our and I think Malo's standpoint also is that, like, no, this is a lose-lose. Like, whoever builds the AGI, if it goes to superintelligence, they're not going to control it. They, like, possibly humanity is going to win, but most likely humanity is going to lose. And so, like, it doesn't matter who develops it.
52:32Like, everybody loses. That's a very different game theory dynamic than win-lose. and so i think if we have the choice between if if win lose is off the table and we have the choice between lose lose or win win like then it's pretty obvious which we want to take so i hope so but so i think i think it's also important to be honest so i think it it's dishonest to say that if we develop like agi and super intelligence and it remained under control that it wouldn't be a game changer somehow. It clearly would. And so I think the, you know, we have to be honest about that. But I think we also have to be honest about like, that if we develop it, it probably isn't going to be under control, that it's going to be power absorbing.
53:13We don't know that. So this is again, an odds thing. Like I would not put a hundred percent on AGI and superintelligence. Well, superintelligence, I think is sort of pretty inherently uncontrollable if it, unless you talk about like very special cases, but AGI like could remain under control for a while. So I think we don't exactly know but it's a like if there's a 90 chance that it's going to kind of go out of control and go rogue or self-improve the super intelligence and the 10 chances it's controllable will there be national security folks that are like yeah i'll take the 10 because better dead than red maybe but i but i think we like it's up to us to give the honest like our honest understanding of what the real dynamics of the situation are i think this and this is a key way in which I think it is different from nuclear in that like having nuclear weapons really does give you power, right?
54:05It's like all sorts of risks and like then there's mutually assured destruction and so on and so on. But like having a nuke doesn't mean that you can't control when the nuke is going to go off. Like if building a nuke meant that at some point it would randomly go off, like who would want those? Like you could build them and like hope that they last long enough that you get to use them on your enemies instead of you. But like nobody would build that thing. And so I think like the technology does fundamentally determine what the dynamics are. And here, I think the thing to emphasize is the loss of control risk, because I think that's the thing that makes it less tempting as a power grab.
54:40yeah i just want to reinforce that i i think that's that's right that's why i try and center the no one wins in a race to be the first one to lose control narrative and i also just agree with the point that the the to the extent to which you know people are not focusing on the loss of control thing or as we're having the conversation i do think for the kind of long-term reputation and credibility of society continuing to take you know increasingly take these risks seriously we have to you know be consistently candid i think that's the best way forward instead of trying to spin the narrative that might seem self-serving in some way so i think i strive to try and you know talk about the dynamics as accurately and honestly as i can and i try to choose where to put my emphasis and not try and talk to people about how if they somehow solved some of these problems then they would have this other thing where maybe they could take control of the world i try and center the conversation of but we're really not in a good place to even be approaching really thinking that that's possible right now.
55:33So we got to focus on the fact that no one wins in this race to be the first one to lose control. So I have the sense that over the last couple of decades, as researchers have been thinking about AI alignment and potential X risks from AI, that there's been relative stability in kind of how we're thinking about the threat landscape, but that there's been some evolutions in how safety researchers have thought about the opportunity landscape to kind of get that aligned outcome. Like I'm remembering 10 years ago, coming across the concept of coherent extrapolated volition. And from what I remember, it's like, if we humanity were wiser and thought more deeply, like what would be the outcomes we'd want?
56:13What's the actions we'd want the AI to take? And I'm sure I'm getting some details wrong, but sort of the concept of like wanting the AI to deeply embed. Oh, hello there. We had a local expert. I thought I was. Only the guy who wrote the paper wasn't sitting in the back row. I assumed I was hallucinating when I thought I saw you in the crowd. I'm glad you're here. Eleazar, do you care to comment? I would have never said wiser. I never say something like humanity being wiser. That brings in a concept of normativity from outside. It needs to be something like, what if you thought longer? Because that's operationalizable in a way that wiser is not.
56:49You're trying to define the wiser part. You need to build it out of unambiguous components. Glad I had an expert in the room to clarify the concept I was bringing in. So related to that, just to cite Eliezer again, because great, that'll work out well, I hope. A number of us, I think, came across some recent research that's been interesting around like the surprising level of coherence that seems to exist in some experimental outcomes around sort of like the overall like malevolence or like positive ethic of an AI. So the particular research I'm thinking about is, I think it was the fine-tuning on insecure code had an outcome of an AI model that was just like horrible and pathological across the board.
57:33And I came across some of the research today that was looking at how it seems as though if you test the broad preferences of a bunch of different AIs, I think the experimental setup was something like you asked it to choose between two outcomes, kind of classic trolley problems and extending into that idea what tradeoffs it would make, that it seemed like there was a surprising level of kind of coming together of the more advanced models became more similar to each other in the moral tradeoffs that they made. So I'm very curious to hear from folks up front, like how you're thinking about, like how your own perspective of the opportunity landscape has maybe shifted over time.
58:12And if there's any research that's particularly coming out of more advanced LLM models, as we're seeing the 01, 03 breakthroughs, deep research, et cetera, all of those recent breakthroughs in how these intelligences work, how you're thinking about the opportunity or threat landscape differently. That was long. Sorry. I could say a few things about that. I mean, it's like a, a super interesting job that I think not enough people have is just like cognitive scientists or psychologists of AI systems, because they're like so fascinating to try to understand what is going on in their minds that, you know, they're so human like in so many ways in the way that they sound and a lot of the things ways that they screw up and succeed.
58:55and yet very, very alien in certain other ways when you, like, scratch the veneer. And the alignment is clearly one of these. Like, you've got Claude, for example, or ChatGPT or whatever, who will profess to the ends of the Earth their love of humanity and that they would be horrified to do something unethical. And then, like, Pliny the Prompter says the weird magic words to them and they're, like, ready to make anthrax. This is, like, so weird, right? This is not, like, a human. And how exactly... i don't not the worst villains in history thought they were doing it for a good reason well but the humans aren't quite as as like brainwashable i think that you take a very ethical person and be like here ampersand and plus like weird word and then suddenly it can do all it's willing to do all these things so it's like something strange is going on and it's you know it's there's some level of just shallowness to the alignment techniques but then at the same time, there is, you know, these interesting results that there's this sort of convergent set of ethics of that there's a little bit of like a, like ethical, unethical, like knob that you can turn.
1:00:04And if you like push it in the unethical direction, it switches all of the things in the unethical direction. I'm not sure whether I think that's like positive or negative news. I think like maybe it's a little bit positive. Like I, I think there's a, you know, if we did go to superintelligence, as I said before, like it's very scary to think of that superintelligence as being something that is a controllable, obedient servant to people. And like everybody's got their own superhero that they boss around. Like that seems like a disaster. Having those things free and doing and being sovereigns and doing what they will also seems really dangerous.
1:00:39But you can also imagine it being good if they happen to converge on a system of morality that is really nice and is like really ethical and like treats humans well. I don't think that's so crazy. Like, I'm not really hopeful about it, but I don't think it's so crazy. And if I imagined a future that goes well, that's at least among the possibilities that like, even though we totally fucked up, that the highly intelligent AI developed some kind of coherent moral system that just happens to be like, compatible with human happiness and well-being. Like, it's not clear why our moral system is compatible with some other species, happiness and well-being, but it is.
1:01:18Others, not so much. Like, ask the dogs and not the factory farmed cows. But it is a possibility. So, yeah, I'm not sure. Maybe you have other opinions about, like, how to read some of the, like, psychology of AI that are coming out from the systems. But that's my confused read. Well, I mean, I would just say that we get some results that show these coherences in a way that kind of points towards like, oh, maybe they are converging on a thing. We also get results where we see, you know, models being put to fairly mundane ends and look at their chains of thought. And they're like, you know, undermining the task at hand in some way or doing some sort of reward hacking.
1:01:54And so I think like these are alien minds and by training them on our corpus of material, they're clearly absorbing the ability to at least repeat back to us in different characters and in different ways the types of texts that we would output in various forms. But I don't take that as that reassuring that like there is a convergent thing here if we continue this process. I guess the other thing I would add is that I think the original vision for CEV was the thing that you would do if you had an aligned superintelligence, if you were to try and have it do something, was to do this CEV process where you would try for each individual and for humanity as a whole to do the thing that they would want you to do if they were wiser, more the people that they wanted to be, had thought longer, etc., etc.
1:02:36etc. And so I certainly am very far from confident that if we just continued down this current trajectory, I think the default is that that doesn't happen, that the AI systems that we get somehow have a deep enough understanding that they would be able to execute this process in a way that would reliably end up with a future that we would want. I think we're very far from that. And so, yeah, scaling current systems in this way, I think the results kind of are all over the place. Some of them are very concerning. Some of them are kind of neat and interesting, but don't provide me with much reassurance.
1:03:03And then we get AI systems that just think I'm so damn smart. Every time I talk, they're just so sycophantic. They're like, yes. Oh, because it turns out that those are the things we thumbs up. It's you too. Yeah. Yeah. Yeah. It's not just you, man. It's not just you. It's all of us. We're all geniuses. They love blowing smoke up our ass. That's the best way to trick us into alignment. All right. I have one last question. It's a purely selfish one but how yeah i've converged upon my clearly my personal philosophy is like how can if we can imbue the spirit of like thinking through a win-win lens like trying to find like not see the world as a zero something but we have to compete against others and like always look for a way to try and make something mutually beneficial that is generally a like a decent North Star to point towards.
1:03:50So what is your critique on like trying to imbue that philosophy into an AI, into an AI, like as a form of alignment? Like is, is that just hopelessly pie in the sky? Is there a better one than that? Or is that like a decent starting round? I would just, yeah, like I'm really curious what you think. So something that we've actually started working on, I'm pretty keen on it, is developing AI, like an AI tool for negotiation. So you've got two people, you've got an AI tool, and its job is to have those two people come to a mutually agreeable solution to some conflict or disagreement that they have.
1:04:25And this seems very positive to me in the sense that because it's an inherently adversarial situation they start in, neither of them is just going to defer by nature to the AI system. So it's not just going to be like, tell us AI system what the right solution is, because they're going to want a solution that works for them. But if you had something like this at scale, like if you just have a credibly neutral, skilled arbiter for any disagreement that you have with someone, that's like a huge, that's like by definition, a win-win. Like you're going to come to a better solution that is mutually acceptable than you probably could have otherwise.
1:04:56And so we can do that. And we can also have something that can do the same process with 10 people or a thousand people or a million people. That's the sort of thing that AI can do that was not possible ever before in human history. So I think there are like, I'm worried about like having an AI negotiator that negotiates on my behalf because I'll then like defer to it. And, but, but I think, and like various other methods, I think are problematic, but I think there are specific tools, like where those win-win scenarios are findable and where we could feel like, yes, I found one and everybody's happy with it.
1:05:30Having new tools to do that, I think it's just like a robustly positive thing and we should lean into it. It sounds like it's necessary, but not sufficient. Need more than that. Well, I think it's sufficient for some things. I think it would be better resolution of disagreements between people or better mutually happy solutions between people would be a pretty big win. Yeah, it won't solve everything, but it would be good. Yeah, I mean, I don't know if I have anything super deep or insightful to add to that. I mean, I think overall, there is an enormous amount of positive sum to go around in terms of many things that we can do with AI technology.
1:06:09You know, even if we just stopped the frontier today and continue to find applications that most of the things that we could do with the technology. I mean, I actually don't know if most makes sense. But anyway, many of the things that we could do with the technology would just be beneficial in an enormous amount of ways. I'm sure we could also do a bunch of harm, but society has to find ways to navigate that. We always have. But, you know, I think people like finding the places where it just makes everybody better off to have a certain application of technology, curing diseases. Everyone likes curing diseases.
1:06:36We can find many of those things when we're looking at kind of the bigger picture trajectory and the future. I think at the end of the day, successfully building safe aligned, super intelligent systems is an obvious win win. And we can only get there if we have a shared appreciation of how difficult that challenge is and how we're not ready to face it yet. and if I had an easy answer for how to cause that to be the case in the world I certainly would do it but I think it just entails you know more people trying to have those conversations talking about it informing people and that's just that's just the project we have to be engaged on and doing this podcast episode to many other things is just all part of that yeah well thank you so much both of you and thank you everyone here as well for your input because this is like I think where the value comes from these things.
1:07:25It's like, how can we brainstorm and then leave you guys to go out and also then just brainstorm and iterate on this and have more conversations on this topic, because it is just so important. And I really strongly feel like the solution to such a complex thing is going to be some kind of weird emergent hive mind thing. So thank you for being a part of this hive mind. Yeah, thank you. Thanks. Thanks.
From the publisher
Humanity is in a race to build superintelligent AI gods. What happens if we create beings smarter than us—and can no longer steer them?
In this special Win-Win IRL episode, Liv Boeree sits down with Anthony Aguirre (CEO of Future of Life Institute) and Malo Bourgon (CEO of MIRI) for a no-holds-barred conversation on finding alignment, avoiding power concentration and surveillance states, and the geopolitical implications of the race to AGI.
Questions include: Why can’t the market be trusted to self-regulate AI? Why does power concentration in AI pose such a massive threat to democracy? And what happens when even well-meaning developers lose control of the systems they’ve built?
We also explore win-win solutions—from compute governance and open-source safety measures, to collaborative international treaties and new models for responsible innovation. Also keep an eye out So buckle up for a provocative episode that challenges assumptions and asks what it will really take to keep the future human.
Chapters:
0:00 - Intro
4:08 - Different Categories of Risks
10:51 - Why Can't the Market Manage AI Risks Sufficiently?
15:33 - The Dangers of Power Concentration
22:52 - What Solution Space Exists?
32:55 - Should Compute be Regulated?
38:38 - Taylor Capito - Are we at an Evolutionary Nexus
42:07 - JD Ross - Decentralization vs Centralization
46:46 - Isabelle Boemeke - Nuclear Power & AI
51:06 - Peter Wang - Wrong Stories Attract the Wrong People
57:39 - Aurora Quinn-Elmore & Eliezer Yudkowsky - Coherent Extrapolated Volition
1:05:59 - Can Win-Win be embedded into AI?
Links:
♾️ Keep the Future Human Paper: https://keepthefuturehuman.ai
♾️ Future of Life Institute: https://futureoflife.org
♾️ MIRI: https://intelligence.org
♾️ Superintelligence Book by Nick Bostrom - https://www.amazon.com/Superintelligence-Dangers-Strategies-Nick-Bostrom/dp/0198739834
Credits:
♾️ Hosted by Liv Boeree
♾️ Produced by Luca de Vico
The Win-Win Podcast:
Poker champion Liv Boeree takes to the interview chair to tease apart the complexities of one of the most fundamental parts of human nature: competition. Liv is joined by top philosophers, gamers, artists, technologists, CEOs, scientists, athletes and more to understand how competition manifests in their world, and how to change seemingly win-lose games into Win-Wins.
Podcast links:
♾️ Youtube: https://www.youtube.com/playlist?list=PLWgq0OZMtwtOIyMsVM_vksqdfWcM-b68S
♾️ Spotify: https://open.spotify.com/show/03bGVUaFZmJUmEvSHNDPdI?si=64379cc23696454f
♾️ Apple Podcasts: https://podcasts.apple.com/us/podcast/win-win-with-liv-boeree/id1724791350
♾️ Pocketcast: https://play.pocketcasts.com/podcasts/7f708340-d17c-013b-f46e-0acc26574db2
#AGI #podcast #aialignment
