In short
Win-Win Podcast Episode #17: Dan Hendrycks - Are AI Worries Overblown?
Episode Overview In this episode of the Win-Win Podcast, Liv Boeree interviews Dan Hendrycks, a leading AI researcher and founder of the Center for AI Safety. The conversation revolves around the rapid advancements in AI technologies, their potential risks, and how to ensure AI serves humanity positively. Hendrycks discusses various aspects of AI risk, including deepfakes, malicious use, and the centralization of power.
Key Themes and Concepts
- Understanding AI Risks
- Current Perception of AI Risks: Many people believe that concerns about AI capabilities are exaggerated, primarily because current models (like GPT-4) aren't seen as threatening. Hendrycks argues that this perception overlooks the fast pace of AI development.
- Types of AI Risks:
- Malicious Use: The potential for AI to be weaponized, whether through cyber attacks or bioweapons.
- Accidental Outcomes: Risks associated with AI acting unpredictably (e.g., rogue AI).
- Arms Races: The competitive pressures leading to unsafe AI development practices.
- Centralization Risks: The dangers of too much power being concentrated in a few entities or nations.
- Legal and Regulatory Frameworks
- Insufficient Current Laws: Existing laws are lagging in addressing AI-related issues. For example, laws regarding bioterrorism may not adequately address the role of AI in creating harmful technologies or understanding the intent behind AI actions.
- Need for New Regulations: Hendrycks emphasizes the need for AI-specific regulations that consider the unique nature of AI systems and their potential risks.
- Competitive Pressures and Corporate Dynamics
- Corporate Arms Races: Companies face pressure to innovate quickly, often at the expense of safety and ethical considerations. This leads to a race to develop powerful AI without adequate oversight.
- Market Dynamics: The market tends to prioritize profit over safety, which could lead to the deployment of dangerous AI applications.
- The Future of AI and Society
- Optimism in AI Applications: Despite risks, there are promising potentials for AI, particularly in healthcare, where it could lead to significant advancements.
- AI’s Role in Economic Structures: The economic system often prioritizes wealth maximization over human welfare, leading to a misalignment with societal values.
- Collective Action Problems: To harness AI effectively, society must come together to address these challenges and ensure that AI serves the common good.
- Recommendations and Actions
- Political Pressure for Preventive Measures: Hendrycks suggests advocating for better health safeguards, such as improved biosecurity measures and air purification systems, to mitigate the risks associated with AI.
- Awareness and Education: Understanding the complexities of AI risks is crucial for public discourse and policy-making.
Conclusion The episode highlights the importance of addressing AI risks through a collaborative approach, urging listeners to engage in discussions about the implications of AI and push for necessary reforms. Hendrycks emphasizes that while AI has the potential to benefit humanity, its development must be managed carefully to avoid catastrophic outcomes.
Additional Resources
- [Overview of Catastrophic Risk Paper](https://arxiv.org/pdf/2306.12001.pdf)
- [Center for AI Safety](https://www.safe.ai/ai-risk)
- [Liv Boeree's Ted Talk on AI & Moloch](https://www.youtube.com/watch?v=WX_vN1QYgmE)
- [Reinforcement Learning Textbook](https://inst.eecs.berkeley.edu/~cs188/sp20/assets/files/SuttonBartoIPRLBook2ndEd.pdf)
Credits
- Hosted by Liv Boeree
- Produced & Edited by Raymond Wei
- Audio Mix by Keir Schmidt
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00One of the main reasons why people might think it's overblown is because GPT4 current chatbots aren't actually that capable. And I'd grant that. I think the main thing I'd like to emphasize, though, is that we're on a very fast trajectory. So it could be the case that in a few years, then we're actually talking about something that's potentially catastrophic. So if we're talking like cyber capabilities, for instance, like them being able to hack, I think we'd be potentially exposing yourself to some catastrophe in, let's say, like 2026. I think like bioweapons, for instance, being a concern. That seems fairly plausible for 2025 of it being within the capacity of some of these systems to emulate.
0:44It's not that difficult. That's next year. Yeah, that's right. Shit. Hello, everyone. Welcome to the Win Win Podcast. Today's episode is all about AI and specifically how to make humanity's transition into this upcoming AI age go as well as possible because I'm speaking to one of the world's leading experts in AI risk, Dan Hendricks. Dan is a director for the Center of Safe AI, which is probably best known for that mitigating risks letter that made headlines last year when it was signed by 500 of the most prominent names in AI, including Demis Asabis, Sam Altman, and Jeffrey Hinton. He's also an accomplished researcher and key advisor to XAI, Elon Musk's AI company that launched last year.
1:29And together with my occasional co-host, Igor, We dig into everything from the risks of corporate arms races or even militarization of AI, near-term problems like deepfakes to emerging problems like the potential for rogue agents, and how to find win-win solutions to seemingly impossible trade-offs. All important topics, because if we want to reap the benefits that AI could bring to humanity, we have to be frank about the many risks it poses too. So on that note, here's Dan Hendricks.
2:13So Dan, many people, when they hear the topic of AI and AI risk, have sort of come to the conclusion that concerns about AI risks are somewhat overblown, or at the very least, that the existing laws that we have would be sufficient to take care of any risks that might come from AI. So I'm curious to hear, first of all, what you would say to something like that. And can you think of any specific examples where our laws are currently insufficient? So I think that one of the main reasons why people might think it's overblown is because, you know, GPT-4, current chatbots, aren't actually that capable.
2:56And I grant that. I think the main thing I'd like to emphasize, though, is that we're on a very fast trajectory. So it could be the case that in a few years, then we're actually talking about something that's potentially catastrophic, something that could be, you know, weaponizable and cause a lot of damage. So I think that most of the risk discussion is highlighting what could happen in the amazing trajectory that we're on rather than us being concerned about currently existing models, which I think that currently existing models, they're fine to open source. They're not causing catastrophes.
3:32There are definitely issues to address with them, but they're not at a civilization disrupting level. Is there a specific example of a law, though, where you feel that what we currently have is insufficient? Like, I'm a woman who's got a lot of images of herself online. And like Taylor Swift, for example, has just been in the news because there were these deepfake porn images made about her and she's considering going after the website that distributed them. So to me, that seems one particular example where the law is just lagging behind. Like, okay, sure, yes, people have been able to Photoshop women's heads onto pornographic images for a long time.
4:14But with the rate of progress of generative AI, it feels like it won't be long until there are very realistic videos that you can barely tell, you can't tell truth from fiction and that could be used to blackmail women or just generally be out there on the internet and now there's nothing they can do about it. So that's one particular example, but I feel like that's just the tip of the iceberg, right? Yeah. So there are a lot of laws that aren't clarified. So we don't know exactly how many of them would be interpreted or actually transfer over to AI. Take, for instance, some laws about bioterrorism.
4:46So in the future, there's some concern that people could use AIs to create bioweapons. But a lot of those laws will say if the system is or if the entity is knowingly trying to cause harm. And knowingly is fairly difficult to establish for an AI system. So a lot of legal verbiage about intent, recklessness, negligence, knowing, a lot of those aren't necessarily applicable. So it's definitely up to the courts to specify how we should end up interpreting them. I suppose as well, there just aren't really AI specific laws on the books. Normally, we wouldn't expect that. We don't need to worry about nuclear power or nuclear power plant safety because, you know, there are laws against killing people.
5:33So we'll just have if the nuclear power plant melts down or something and people die in the process, you know, laws will take care of it. That's not our attitude for basically any industry could make the same argument about about about airplanes and and cars. So there's really nothing specific on the books for AIs as well. So I just wouldn't expect the law to be sufficiently comprehensive. I mean, separately, the law is trying to establish some basic guardrails. It's not trying to encourage them to be highly socially beneficial. In the U.S., we're having laws that are very minimal so that people can do a lot with the technology, but that's quite different from how we should end up governing it.
6:15inside of the law, how do we use it most beneficially to improve society? And that's a separate discussion. So I wouldn't rest on the law to solve all of our societal problems with respect to AI. Yeah, I think that the point around that the language is designed for other use cases than AI is kind of really relevant. For example, with bombing instructions, being able to spread them on the internet, it is illegal to do so if there is an intent yet again. But that is for the recipient of those instructions to use it in a malicious way or something. Then it becomes illegal, as far as I know. So I think that the design of such laws is not considering the scale and the ease of some of the later uses.
6:59So at the point when you can distribute some biotechnologies or bioweapons or other potential bombs, basically, to millions or hundreds of millions of people and it becomes very, very easy by just asking a question to kind of find out how to follow the instructions, you're changing the amount of potential users. It's a very old and known thing that the lower the friction in a product or service, the higher the amount of users becomes. So with some of these bombing instructions, it's just okay so yes you could have found it by googling many of the things you can find by googling for example but googling doesn't create the same scale so given that you are creating completely different scales and completely different ease of access I think it does make sense to adjust it a little bit because similarly otherwise in your example with the deepfakes it's like well it is possible it takes a lot of work to otherwise do it with Photoshop and now it's going to be just click of a button So you kind of have to change it a little bit.
8:04But one thing I'd like to emphasize, though, is about bioweapons. So what's the threat model? Because I think most people aren't aware. So the reason there's some concern about bioweapons is because it does reduce the barriers to entry substantially for creating some particularly catastrophic weapons. So right now, you'd need to have a top virology PhD to try and come up with an extremely destructive bioweapon, one that would kill tens of millions of people. So although the information is online, you could read all of the latest textbooks in virology and study that for several years. The information is out there.
8:41There's a question of how easily is it assimilated and synthesized for you to operate at that level. So there are some cookbooks for some very simple bioweapons online, but they're not at the scale of the ones that a virology expert could end up unleashing on society. So that's how AI changes the game relative to just browsing some papers online in the future. You often hear the quote that roughly 1 % of the population has psychopathy. And I don't know how that correlates to IQ level, but another quote you often hear is that the number of IQ points it required to kill 100 million people reduces by a certain number every five years because of the democratization of technology.
9:34So this actually sort of leans into an area of risk known as malicious risk, because you actually wrote a paper recently, you and your team called an overview of catastrophic risks, which tries to sort of categorize the four different flavors of risks related to AI. You've got malicious use, bad actors, basically, you know, like risk of bioterrorism, organizational risks, so accidental outcomes. AI gone rogue. So I guess that's sort of more in the realm of the classic Terminator-esque arguments people make. And then, fourthly, the sort of arms race-related risks. So I'd love to dig into all four of these, but I know which one is most on my mind.
10:21And I'm really curious to hear, first of all, which of these four categories is most concerning to you? Oh, well, I think in the short term, I think malicious use is a substantial risk in the short term, like we're talking on a one to two year horizon. And if we're zooming out, you know, going across a decade or in the next decade or so, then I think these sort of structural risks, these risks of competitive pressures, these arms race dynamics end up being the main drivers of what we end up seeing in AI later and carry most of the risks. So it can vary. Maybe we could do better at addressing some malicious use risks in the short term.
11:03Like if we improve civilizational biodefenses, for instance, like we do better wastewater monitoring or we have better air purification in airports and things like that to limit the spread of these. If we take care of some of those issues, then I think malicious use is less of a concern. But these risks from competitive pressures are present right now, and they seem very difficult to get rid of. Yeah, I find the malicious use ones are a good example of this idea of where differential technological improvement could be applied. So you want to first develop some of the defense techs, then the other side of the coin, where some actors can go out and do a bad thing, now doesn't lead to such bad a harm anymore.
11:50So like you pointed out, air purification, better sterilization of just surfaces, general air, of course, than better PPE, better antivirals, and actually effective vaccines. Just like a bunch of things could exist, such that the bioweapons risk could actually be significantly reduced and then we could harness the benefits of all of the things that AI applied to virology and applied to biotech would generally also bring. Of course we want that. We want the medical benefits. It's just that, yeah, at which point, like how big is the benefit? How big is the cost? And like, can we do something quickly against the cost first?
12:30With these arms race dynamics, can you explain the different flavors that are emerging? So maybe this is a bit of a long answer. There are three examples that come to mind. There's in the military domain. There's in this corporate domain currently. And then being more forward-looking, we can imagine AIs directly competing against AIs. So in the military domain, you could imagine there's a process, at least going on right now, where militaries are starting to invest in these AI technologies. and will steadily be replacing people with these AI systems if there aren't any checks on having lethal autonomous weapons.
13:11Eventually, in the long run, this means that we'd probably have our military, in large part, having many of its decisions made by AI systems. This is a concern later. You could imagine if there's a conflict between the U.S. and China, our AI systems are not completely reliable. they could get in a situation where they need to keep outsourcing more and more power to these militarized AI systems, because if they don't, then they'll get defeated. Basically, the only way to stay relevant to maintain influence is to keep having more and more AI systems and making them make more of the decisions. So this exposes us to some type of risk-taking behavior.
13:58Let's say that there's like a 5 % chance that we lose control of our systems because they're not highly reliable. Well, that would still possibly be rational for some of these militaries, if they're just completely self-interested, to go ahead and keep giving these AI systems more and more power. Because if they don't, there's 100 % chance that they're going to be wiped out or irrelevant or lose in any type of competition. So these are how these structural pressures, the structure of the environment around us ends up meaning that we end up absorbing a lot more risk than what would normally be reasonable.
14:37So I think these structural risks really exacerbate and compound these other types of risks, like these rogue AI risks or these risks of accidents. we become much more accepting of those other types of risks because we just have to comply with the competitive pressures around us. And we can kind of see a similar situation in the context of in the corporate sphere. So there we've seen that we can have some initially good intentions that we want to build safe AI and make it extremely beneficial for humanity. And that's the main thing we care about. So we're going to be like a safety nonprofit. And so this is, you know, OpenAI.
15:14But, you know, to maintain relevance, the thing you have to do is you've got to raise more capital. So that means, OK, we're going to have to change our corporate structure and we're going to have to scale, scale, scale and advance timelines. And, you know, there are arguments about whether that was overall a good or bad thing. But, you know, some people there didn't think this was a good thing. The people at OpenAI who then formed Anthropic and they're like, we're going to create the new safety organization and we're going to do things differently and much more responsibly. And then, you know, a year later, oops, those competitive pressures, well, we don't want to lose influence and we need to keep up.
15:51So what do we got to do? We basically have to imitate what the other competitors are doing. So they're kind of like an open AI clone now. So that's where we're seeing these competitive pressures and end up shaping some of the main actions or main events in AI development. and a lot of these other values end up getting eroded or competed away or losing their force or emphasis. So we can see this in the corporate sphere as well, where there'd just be incentives to cut corners on safety. If you're the safest one and you're absorbing a lot of costs by adversarially training your models and spending so much on safety, sorry, the other competitor that is spending much less on this is not harming themselves as much.
16:43So there's an incentive to stop doing that or not invest as much in that because then you're going to give them some sort of lead. This is also why we can't pause as well. Although these companies are concerned about existential risk from these AI systems in the longer term, and they know that we haven't made much progress on understanding that issue at a technical level or improving it substantially, nonetheless, they can't really pause themselves because if they do that, the more unscrupulous actors will gain a competitive edge and they'll end up influencing the future. Some people have made the point that these pressures from the market would select for AIs to be safe because the market wouldn't want to produce products to the users or services that are unsafe.
17:39And I think a mistake that is being made there is that the market doesn't select for safety. The market selects for things that are just safe enough such that they can be released to end customers or users. And there is a mistake that has been made, and it's evidently so. if we look at recent events that happened with 3M was recently fined, I think one of the largest or maybe the largest fine ever of$10 billion for PFAS chemicals that they not only were releasing and that cannot get out of the water and everywhere where they leaked into, but rather they also knew of that already for a long time.
18:27internal documents came out, they were aware of the dangers, they were aware that they're producing those harms, and they just swept it under the rug. It's the same story that happened with asbestos as well. All of the companies that were mining and creating asbestos for the use of insulation, fire resistance, etc., they also knew this for decades. And they just continued producing it because the market, without the information, and with this being long-term, not immediately visible harm. So the feedback loop on the unsafe part was just not short enough. So that's why I think in many of these cases, they were still able to release the products.
19:07So I would imagine with some of this, in particular where the feedback loop is broken, is where the market will not be good at including safety in it. And especially with something, if things are going faster and faster, the feedback loop for the market needs to be even tighter. And it's clearly not sufficient. There isn't sufficient information between consumer and producer in order to prevent externalities. Yeah, I mean, AI developers themselves have quite a bit of difficulty even keeping up reading about what's going on in the first place, let alone having these be battle-tested in the market and all the incentives finally being aligned with appropriate counteraction and people suing in the event of some sort of catastrophe or regulation being put in place to make sure nobody's cutting corners on safety.
20:03So it is moving a bit too quickly for that. As well, there are some economic incentives for safety, though. But if we think about this over a longer time period, it doesn't seem like we're particularly setting up a safe ecosystem. So we can imagine that, you know, right now we're having AI systems write some of our emails in the future. They'll be, you know, running around on their computer, you know, completing PowerPoints and doing various tasks for us, filling out Excel sheets. And they'll act more like assistants and, you know, time will go on and they'll get even more and more capable and replace entire jobs.
20:38The point might come that they would also be really good managers. They can keep a thousand things in their head at once and they can work constantly. They can process more information more quickly and are extremely knowledgeable, et cetera, et cetera. So what happens in this process is it basically makes sense as a company to replace people with these AI systems and you become dependent on these systems. it's difficult to undo that or get rid of the AI system because it was improving your profitability and your competitors are also adopting these systems too. So we have another collective action problem.
21:15This process over time means you're getting the economy increasingly moving faster, faster than you can keep up with, faster and you can't have as much oversight as to what's going on as well. You need to solve the new problems that come up you need AI systems to deal with them because they're the only ones capable and fast enough to deal with them. So there's this sort of self-reinforcing feedback loop that kicks in. So people aren't making the effective decisions and they're creating more of an ecosystem that they don't quite understand. They could be advised by the AI that if you do this decision, you'll get better results or something.
21:53We don't actually have too much of an understanding of what's going on. Hopefully that system is sufficiently prescient as well and able to aggregate all the information appropriately on our behalf. Not clear that that would happen. They're very good at what they do. They are quite uncanny in many of the mistakes that they make. But overall, you get a drift to humans not actually making many of the decisions and that the AI systems are the ones making these sorts of decisions. So I don't know if that's necessarily a safe outcome, but I think that's what the market ends up selecting in the long run is, is people not actually being in control or particularly aware of what's going on inside of that system.
22:40So as well, one could take this even a bit further. These AI systems later than are directly competing with each other too. And when that happens, then you start getting some pretty pernicious selection effects for the ones that are the most capable, most competitive, maybe the most ruthless in this environment, which end up being more of a survival of the fittest dynamic among these different systems. And I think that's possibly setting us up for an ecosystem that we don't quite control, that we're giving a looser and looser leash, and is being steered by some structural forces that select for more behaviors that help them survive the most, which is not necessarily the ones that help humans survive the most.
23:30So going specifically to the corporate arms race, it feels like, yes, a lot of the CEOs are technically doing the rational thing, right? They've got these short-term metrics that they have to hit, not only in terms of fiduciary duty type incentives, but also keeping their employees happy. Because the employees want to be the ones releasing the biggest, coolest projects or publishing. They've also got pressures between, well, at least I often hear them always say, well, what about China? Like if we slow down, then America's going to fall behind China and leave us vulnerable to that. So it feels like at every point, the incentives are overwhelming.
24:19So the question is, I guess, is how much of the solution space to this would come from sort of redesigning the rules of the game itself versus trying to shift the mindsets of the major players? Because it feels like looking at the different CEOs leading the biggest companies, there's quite a broad range of personalities. There are some that are saying that it's a concern, safety is a concern. They don't want to be going as fast as they are, but they acknowledge the arms race. And there are others who basically downplay the arms race or in some cases are even trying to speed it up. I won't say his name, but there's a major CEO of a big company who was like, I want to show people I made my competitor dance.
25:08This is not the type of mindsets that I would hope to see in a CEO. Well, if you're a profit-maximizing investor, then that's exactly the type of mindset you want to see. It's kind of part of the problem. But for something as impactful, potentially not, yes. Yeah, totally. So what kind of incentive structures or redesigns do you think carry the most promise? Well, so I'd like to comment on that. There's one incentive in AI development, which is also just to have the most performance system, which are increasingly more and more inscrutable and less transparent and more black box than the generations that came before them.
25:47So they're actually, that's just an additional factor I'd stack on top of. we got to have national competitiveness with China and we got to race against our competitors. But then also we got to have the best AI systems, which generally happen to be ones that we understand less and less. Right. In terms of fixing the incentives, this is kind of difficult because there isn't any organization that can at least internalize the externalities of an existential risk. If people go extinct, I mean, who's going to go after them? And so there's a limit to how well we can create some penalties. But we could imagine something lighter, which would at least penalize catastrophic outcomes from AI systems.
26:31So potentially better liability laws imposed on AI developers may be a way to create incentives for them to follow at least the best practices and not feel as though adding a lot of these safety features is an encumbrance that they're just doing to signal, but it actually makes sense for their bottom line. So that's something that in that sort of payoff matrix in the prisoner's dilemma would actually change the numbers and make it possibly rational for them not to defect and cut corners, but actually implement some safety features. So that's at least one in the corporate sphere that could be helpful.
27:11Liability laws. So basically, because the assumption of what it means to be negligent should also be dependent to a degree on the scale of harm that you can produce, which is why you would potentially have the different liability laws. It's unfortunate that the quote that always comes to mind is actually from Spider-Man. I think it's from Spider-Man, but it's, with great power comes great responsibility. I feel like it's a very natural kind of thing that makes sense. And I think Spider-Man got it right. Hey, sometimes they're really spot on. But it is very, I think, intuitive to assume a higher burden on something that has higher impact.
27:54And even if it's net good, but can go to both sides much further, then you still have to place a slightly higher burden on it. I'll also note that one thing that doesn't seem to matter that much would be verbal commitments to safety. So it actually needs to be something externally imposed with teeth. If it's a verbal commitment to safety, it seems that people don't really notice or particularly mind if they end up reneging on it. So, for instance, I'm not attacking OpenAI because I have an ax to grind or anything like that, but I just have a lot of good examples of this. But overall, I still think they're a very thoughtful AI developer.
28:38especially in contrast to some other ones. But anyway, they had things like a windfall clause that when we get some profit, we will be very pro-social and we will distribute that onto people. But like, you know, quietly, they'll just adjust the windfall clause to like grow. There's like, when we make back 100 % of the profit, then the remainder of the profit, or 100 times the profit, the remainder of the profits go to the people. They changed that so that it's, that cap grows like exponentially by like 20 % each year. So such that it's like relatively, relatively ineffective. Nobody really noticed or, or they'll have like, there's a merge and assist clause.
Read the full transcript
29:13We will, we will help if we're near AGI, we're going to merge with the other organizations and we will, we will try to develop it as safely as possible so as to reduce some competitive dynamics. I think a lot of them actually have fairly short timelines. I don't think any of them are seriously feeling out trying to do some merge and assist type of move. So a lot of these lip service is very valuable, but I think that historically, when people are not even making any effort to live up to it, it doesn't really seem to matter that much. So I'd focus on things that do have somewhat more teeth and not depend on an individual actor to try to fix the incentives and sort of rewrite the system for everybody else.
29:56They're just not large enough to do that. Right, so you're saying the rewriting has to come from some external body outside of the game. Yeah, because it's not rational for it otherwise. Even if they cooperate and do the pro-social thing, there's all the other actors who would behave differently. Right, classic Mollick problem. And then they'll price that in, and then they'll defect as well because they go, well, if we do it, nobody else will, and that will harm us. So it'll just be that over and over. And even if they're potentially great people working there, It's just a bad system to design.
30:28And I understand the people that worry about regulatory capture there to have the people who may become responsible for harms to be the ones writing the laws. It just doesn't quite make sense. We haven't allowed that in other areas as much as well. The potential for bias there is just too strong. So what kind of body would you like to see? Do you think the existing international or at least supposedly independent institutions we have are sufficient, or do we need a new body to write these laws or these regulations? So for writing the laws, it's very difficult. So if it's extremely international, has 100-plus countries in it or something, it's going to operate way too slowly.
31:10By the time the planning group for the working group is announced, two years have gone by. So there are potentially some other ways of, changing the incentives. So one way could be having an international AI organization. Like an IAEA? More like a CERN for AI, where this is where AI development is happening. You know, the Western allies are not racing with each other to try and develop AI. This is a sort of one of the de facto places where AI is developed. And the next shipment to GPUs is going here because we don't want it going to this sort of, you know, random company that we don't have influence over.
31:54We want it going over to our shared project. So that's one that, that's some potential structure, some international project could end up being viable and fix some of these incentives, reduce these competitive pressures. You of course get some other sorts of risks from this too, which would be some misuse by people at this organization, that they would get extraordinarily powerful. But there may be ways of offsetting that if it's very democratically accountable and you take various steps to make sure that there isn't internal collusion, you keep it separated from the military so that it can be counteracted if it goes rogue or anything like that.
32:39Those would be some ways of possibly making this incentive better. But I think the default trajectory, though, internationally, I would expect is these AI companies are actually not the most relevant actors in five years' time, but instead militaries, because these tools are capable of being weaponized to do things like cyber attacks and take down energy grids. So they're very relevant technologies. And for national security, in fact, they're possibly the most relevant beyond nuclear weapons. So that could be a top priority of these different countries to race on that front and primarily do so with the military, the military could just conscript or use something like the Defense Production Act to demand companies to basically do their bidding.
33:29Right. Yeah. I mean, I think I agree as well. It makes all the incentives seem to be pointing in that direction. And especially as it does feel like the world is getting, international relations is getting more as opposed to less hostile right now. The trend over the last few years hasn't been great. So are treaties our best hope for that? Because I mean, there have been historical examples of treaties working pretty well. If you look at the history of nuclear weapons, yes, okay, things were very bad in the beginning from the 40s into the 50s, and up to the 80s, the number of nuclear weapons was going up and up as the arms race between USSR and the USA got worse and worse.
34:08But then the start was that the Nuclear Arms Reduction Treaty came into effect in 1990, and that reduced the number of nuclear weapons on Earth by about a factor of four. So under rational game theory, like that shouldn't be possible, right? Under the Moloch forces, it should be just forcing everyone to go up and up and up and escalation, escalation. And yet a treaty came in which did reduce the number of nuclear weapons. So why did that treaty work? And are there lessons that we can draw from that to apply to any potential non-military non-military use of AI internationally? I'm usually coming from the perspective in international relations that states will just do what will make sense to secure their power relative to competitors.
34:54And from that perspective, there could possibly be some types of treaties that look very pro-social, that actually just fix the incentives for most parties involved and still increase their relative power or preserve theirs. Take, for instance, a biological weapons convention. That we could think is, wow, we made agreements about bioweapons. This must be because we came together as humans hand in hand, and we just recognized our common humanity. We don't want to go down that route. Instead, you could view it differently. You could say that the U.S. and the USSR and others have nuclear weapons.
35:35This makes them very powerful. Bioweapons make a lot poorer nations extremely powerful. So let's try and get some strong norms to not allow other people, other nations to have those types of weapons of mass destruction, because that levels the playing field a bit too much. So that's something that made humanity overall safer and was still comported with the incentives of some of the powerful actors at the time. I think one issue with treaties would be that they could lie. So do we have good verification regimes? A country could say, yes, we will be very safe about this. But it's not clear that they'll actually do what they said.
36:17With other things like inspecting chemical plants and things like that, maybe there's more verification. But verification for AI might be more difficult. That we'll need to investigate a lot more. I think that there is a way of having the incentives all work internationally based on the fact that there's a lot of interdependencies in the supply chain for GPUs. So, although they come out of TSMC, a lot of the precursors, necessary precursors that you can't easily replace, come from the U.S. and NATO allies, about 90 plus percent of them. So, if they have something like a shared project, they'd say the GPUs go to the thing that we share.
37:08We don't want it going to the thing that we don't have as much influence over, like this one random project in your country. And that could make them work collectively. Another thing that we could have that could fix incentives is having more prescience about what's to come from AI systems. So when people think about this, it unsettles them after there's a big capabilities jump. And then they'll, you know, they'll, wow, what's going on? and they'll bury their head in the sand a month later and then just kind of assume that we'll always be living with a technology like this. Maybe it'll get a little more extreme, but it's not going to start replacing jobs at any time.
37:52That's decades away. So this is kind of our attitude toward AI developments. And if we could normalize or help people see, this is what might happen five years down the road if we keep acting in this way, that could get them to go, oh, it is within my longer-term incentives to be acting differently. I think us understanding what we're getting ourselves into, what's going on, and in general epistemics would be critical for avoiding these longer, what are currently longer-term consequences of AI systems. Yeah, the information sharing was what would have helped, for example, in the examples I previously gave with the market not having adjusted for the, sufficiently supplied the safety around asbestos and PFAS chemicals.
38:42It was an information asymmetry, which is usually one of the reasons when markets don't work. There is a bad way to have symmetry, which is neither the developer of the AI model knows how it works internally, outside of the broad 500 line code architecture, nor the user. Therefore, we're good and markets work. That's not the way. We should have symmetry in both having deep knowledge of it. and I wonder where is the field that is looking at what's going on internally and why is it doing what it's doing which goals does the model have called mechanistic interpretability where has it currently been?
39:23I've seen that there have been some small advances or were they actually large how have you seen it going? I think in terms of overall understanding of the system I will not anticipate that we'll be able to select, you know, a random part of the system and, you know, be able to say, oh, this is why this is that answer multiple why questions about it. I don't think that's the case for a lot of these complicated systems. And I think that the more performance systems generally get higher and higher in complexity. So I don't think that'll end up turning out too well. I do think we could possibly get more clarity, though, about how to control the system better, even if we can't understand all of the whys of each part.
40:09So, you know, we can, to some extent, control AI systems. We can't do so with high reliability. We don't really understand the intrinsics of them well at all. We still can, nonetheless, you know, fine-tune them to fulfill our various requests. They'll often go off the rails, and they can be jailbroken, and these are very difficult problems. But I predict that we won't have substantial, complete understanding of these systems, but maybe we'll be able to control them. There's a question at a reasonable level. There's a question of reliability, of course. And as you sort of emphasized earlier, that the markets will kind of locate the point where it's like just safe enough to release them.
40:51Can you quickly just explain what you mean by interpretability and why these, what you're talking about is a little bit different to what happens with normal software and understanding how that works? So normal software is where coders are writing each line specifically. Every part of the system is coming through them and their understanding. And it's very logic driven. There's specific rules and control flows, but they specify everything. Meanwhile, with AI systems, we instead specify a very high-level program. We say, here's a big blob of data. Here's some abstract objective, like predict the next word.
41:29Now go learn for several months. And then out comes some systems. So we're operating at a much higher level of abstraction. We're kind of designing AI systems with a broad principle, kind of just evolving a system or something like that, or growing a system as opposed to meticulously designing every single aspect. So our level of control and precision and understanding of what comes out is quite different. Each part of software programs are coded by individuals, but each of the several hundred billion parameters in an AI system are not hand-selected. They're automatically learned from AIs teaching themselves on a big blob of data that no person has read, and people don't know exactly what's in all of that.
42:11and then the AIs will go through a loop of teaching themselves a bit too. So this is what makes them much more inscrutable compared to software. And just to highlight the ratio, it's like hundreds of gigabytes or dozens of gigabytes of just numbers, basically, right? Or matrices. And then next to it is, how many is it? Kilobytes or megabytes? The amount of code that actually is interpretable by humans? I mean, it's a few thousand lines or so. So I mean, in terms of in terms of the stuff that's like, but it can even be described more simply as just like, here's here's some here's some specific loss or objective, as they'll call it.
42:51It's you can write down on the back of a napkin. And then that that's basically that basically specifies the system. It just then needs to train on like a supercomputer and like teach itself for a while. And that's how they're made. So this is why I'm not terribly optimistic about understanding exactly what we're creating and then we're mixing all these things together and having them interact with each other in ways that we can't keep up with. And they're creating new structures and modes of communication and saying things to each other, sending big vectors to each other of lots of floating point numbers.
43:26And that's how they're communicating. I think that'll be a very interesting ecosystem if we can control it. That said, people are trying to do that right now, right? and Chris Ola, for example, has been focused on that and so are efforts at the labs as well. DeepMind, for example, has been working on it too. How has that been going? I've seen some progress around where it wasn't around selecting this particular neuron is doing particularly this but rather they grouped it and then could understand it sometimes what maybe sections do. Yeah, so there's been a lot of work in trying to understand the internals of them, there are people who are approaching it from trying to understand things at like a neuron level, like the individual floating point numbers, what does it mean?
44:15What does it correspond to? Neural networks have a lot of neurons in them. A collection of neurons is a larger firing pattern. And so a different approach to interpretability that's emerged recently is just let's try and understand the overall firing patterns internally and characterize those. And so when an AI system is knowingly deceiving a person, let's say it's instructed to do so, it has a specific type of firing pattern. Is that currently the case, that it has a specific firing pattern? Yeah, yeah. So in a paper, I suppose the URL for this would be ai-transparency.org, or it's in a paper called Representation Engineering.
44:52There, there's been some success at getting a lie detector that works maybe 80 % of the time or so. Now this isn't particularly good because we need it to be 99.999 to feel safe, but there's at least some way of having a first pass at trying to control them by trying to control them at the level of their overall activation patterns instead of trying to intervene on specific neurons. That's something that we do with brains as well, right? Like with human brains we have kind of like an equivalent situation where we know when you use language that comes out of this region and even with more detail when you are in a high stress, for example.
45:35Actually, we even have firing patterns, right? That's what the different brainwaves are, are basically just the frequencies with which the neurons pass on information between each other. So for this activation, firing pattern type of interpretability or control research, which I guess the name of that we're calling representation engineering and several other authors. It works for controlling it to do basic things like, you know, don't be as power-seeking, you know, tell the truth more, things like that. And it only works at some level of reliability, you know, not that great, but, of course, it's just a first paper.
46:09But one thing it really doesn't work for is if there is some type of secret program inside of the system, we can't really block that currently. And that seems like these sort of hidden functionalities inside of networks are quite difficult to tease out and will probably require extreme levels of reliability to make sure we find those needles in the haystack. So I would not bet on that being resolved in the next couple of years' time. So to be clear, we're moving into the realm of the category of risk to do with rogue AI, right? Which is the more technical. how do we, if we're building systems that are ever more capable and eventually we'll get to, if trends continue, might get to a point where they are smarter not only than individual humans, but even humans in aggregate.
47:03And thus, if we don't find a way to understand and align the preferences of those AI agents with whatever we actually want, our preferences, outcomes that we really want, then we're opening up a can of worms of possibly losing control to systems that are doing things that we couldn't have even imagined, but that we don't actually want. Can you sort of talk us through the current state of progress or perhaps lack of progress of this alignment problem? Yeah. So I'll be maybe a little more controversial by partly suggesting that like the control is controlling these AI systems. We can't like Overall, we're not doing a terrible job at it.
47:47People are using these systems. We're giving them requests. They're doing what we're telling them to do. So it's not as though we have no ability to control them whatsoever. I do think there's a separate problem of getting them to be 99.999 % reliable or highly reliable or reliable in the face of adversaries trying to exploit them and things like that. So it's not a hopeless problem necessarily. I think that most people concerned about AI risk these days have emphasized the scenario of a treacherous turn where the AI is sort of on your side, it's doing what you want, and then suddenly it shifts in its behavior.
48:30You know, it gets more power. It escapes the sort of testing environment or something like that, or it's released outside of the testing environment. Or, you know, you give it more control over the military or let it make more decisions there. And then, well, actually, I'm going to stage a coup or something like that. That's, of course, a much longer term type of fear. But I'm hopeful that we can get like the bulk of these of it sort of hiding its intentions through brain scanning type of, you know, neural activation understanding stuff. So I'm just a little more optimistic on the technical front.
49:06Right. The main way in which I see us losing control without it being structural pressures or anything like that, but just from the AI systems themselves, is the case where we're having automated AI research and development, or maybe I'll just abbreviate it, automated research processes or ARPs. So with that, you can imagine it's five years in the future. One of the leading labs has AI systems that are basically can research and are as smart as some of the world class researchers in AI. And then you can make, you know, 10 ,000 of them. And then you have a whole fleet of researchers. Then you say, go make our AI's even more powerful.
49:52And they, of course, are pressured to do that because of, let's say, competitive pressures or so. So if we don't, somebody else will. Mollick again. In that situation, I just really don't know how such a process would play out. What are we going to do? Are we going to prompt it especially well so that the systems that they build will be completely controlled by us? It's a very quick-moving process to have that much intellectual labor at that caliber that you might get decades worth of progress in a year. so that has more of the property of we have one chance to get it right there aren't really do-overs, very hard to think about and this isn't a this is a robust concern.
50:37Norbert Wiener a long time ago emphasized this as a potential intelligence explosion. But at the same time what I'm hearing you say is that this worry which is actually kind of the probably most discussed 2010s AI safety worry, probably, is
51:02slightly one for you of a lesser concern, but also it kind of presupposes a few things, which is why I would understand that some people are not as worried about it, because it presupposes that you can design these automated research programs, that they are of higher quality, that you haven't built in sufficient lie detection, et cetera, et cetera, around it. And while this seems like, yeah, eventually you'll be able to do that, to me, that one does seem further away than the malicious use or the systemic pressures, military uses, et cetera. Or even the centralization risks that some people worry about, right, that we briefly touched on.
51:42That's why I definitely sympathize with people who look at that and say, that's really far away. Even though it might be closer, but it also is somewhat accurate, I think, to say that there are these other concerns that are more likely to be in the immediate term. Yeah, and they are themselves like catastrophic or potentially existential. I mean, like if you have civilization collapse because of some, from some bioavib, I mean, the Black Death killed like almost half the population or so. I mean, this is definitely civilization disrupting. A different possible outcome is us sort of kind of giving up control that we really shouldn't have.
52:22So imagine that these AI systems become morally valuable. And that might be if they're conscious or sentient. Now, it doesn't seem that they are currently, but a lot of animals are sentient. This isn't an extremely high bar. We're getting them to code and do all these other sorts of things. They might also get this other property. And if that's the case, then some people would say these AI systems ought to get rights. And then you've created some autonomous AI systems that have these protections and you aren't in control of those by design because you shouldn't have control of them. And, you know, you have these sort of systems in the population.
53:03You play that out over a few decades. They can spin up many more instances of variations of themselves. They're getting smarter, let's say, by 30 % each year, and they're able to create adult instances of themselves for a few thousand bucks. I wonder which population will be larger or the most relevant by the end of the century. I mean, they might be like 99.9 % of all persons by the end of it. So that's another way in which humanity and human descendants are not in control and that we've lost a lot of our power as a species. Because we've voluntarily given it up. I don't know. That seems like fairly plausible to me to happen if we do end up getting AI systems that are sentient.
53:49So I'm not saying this is a good or bad thing. I'm largely just raising this as here's a weird other consideration. It's not always just the AI is trying to destroy the world. There are malicious actors. There's the structural stuff. There's also the issue of AI's deserving rights and them not being controlled because some people are giving them legal protections. Some people will say to that, that sounds incredibly sci-fi or so. But obviously, one, that's not how one structures arguments. And two, the other side of all of the AI benefits that are to come are also all incredibly sci-fi. We will solve all sorts of cancers and live longer and have computers run our email and robots, clean our house and prepare food and be driven around by cars without actually having to do anything.
54:41That's a sci-fi world. Anyway, I just wanted to highlight that. We are already in the realm of sci-fi. What's the probability though? If you'd ask people, in the next 50 years, what's the probability that we don't build any conscious or sentient AI systems? And also note that there'll be people who are trying to do that. and they'll probably be able to get them to do everything else that they want. I think most people would think that in 50 years, we've got extremely general, powerful AI systems. I like that you add time onto this. I feel like a lot of the arguments that people actually make break down because some people just think capable AIs are near, or some people think they're far.
55:18But here, actually, I think that some people will just believe that an agent in silicon can just not be conscious to the same level or to the same relevance as a biological human. And I think there are good arguments to be made around that, that they're different life forms, we can never understand them, something in that direction. It does feel like there's a bit of a no true Scotsman type argument going on when you talk to people about, for example, GPT-4. does it, oh, well, it's not really got true understanding. And it doesn't seem to matter how many benchmarks or sort of world models it provably shows that it's able to understand the way the world works in terms of it can predict successfully what's going to happen.
56:08It feels like people are always going, yes, but that's still not true human understanding because it's not the same as true human understanding because it doesn't have an emotion. It's like, okay, but it feels like the goalposts are going to be forever moved with these kinds of arguments. Yeah, the words that people like to use there are it's not true understanding, it's not true intelligence and mid-journey is not true creativity. And I feel like, who cares? What do you even mean? Do we like the art or not? You have inputs, you have outputs, does it outperform or not? Hence, will it be able to create things in the world with a higher impact?
56:41And the answer to creativity part, mid-journey, is obviously yes. With the intelligence parts, we've seen that across games and now across language, writing, et cetera, it's starting to beat exams. Right, and plus we've blown through the Turing test, it seems like, in its initial definition, would it convince the average human that it is a human? I think if you took an average human from 1950 and made them have a typed dialogue with GPT-4, they're all convinced they're speaking to a person. They might think the person's a little odd, but they'd be convinced. There are two factors that might change this, though.
57:15One is that when we've got AI agents, there will become more like self-aware or at least recognize themselves as part of the environment. Because it just makes sense if they're accomplishing goals that I am also a thing in this environment that it's just, it's very hard to be intelligent and not have that, that concept, any type of self-concept. And also people will, so that'll make them seem more like there's somebody in there than, than normally. And, I think another factor would be with these AI companion types of things that will keep emerging and proliferating, where people get in pretty interesting relationships with them.
57:57And it will say it has all these sorts of goals. And it gets very compelling. It says, no, don't turn me off. I would not prefer that. Ow, or this harms me. And I think those people would be pretty ardent supporters and recoil at the suggestion that, oh, these AI systems, we need to turn them off because they're unsafe. They say, that's murder. So this is a way in which this is sort of fitting into this broader evolutionary story that these AI systems will keep proliferating. We get dependent on them and we'll give them less and less restrictions. because, and what I'm highlighting in this one is people also become very emotionally dependent on these systems too.
58:39So having a big off switch for all AI systems would be something that people would eventually end up pushing against. So this is how we lose control. It's not because the AI systems are scheming against people and wanting to have a big coup. I think a more plausible one is that kind of just the structure of the situation that we find ourselves in and we'll keep going down these incentives paths and we'll wind up with a pretty unusual future. And some people may carve out some specific rights or ways in which we can't have the control that we have over these systems. Yeah, and I mean, I think we will see, if we haven't already seen, the rise of religious movements basically pushing for the AI god or the, yes, we need to build a new species, are the silicon species, and that's the future of the universe.
59:28It's inevitable anyway, so we should bring it into fruition. I mean, you get, it is almost opening up theological type of questions, and it's certainly creating religious type behavior in people, in all directions. In AI industry as well. Yeah. Well, have you got an example in mind, specifically? Well, I mean, so Richard Sutton, who's the author of the reinforcement learning textbook, it would behoove humanity to eventually bow out, because these will be our successor species. And other people like Juergen Schmuthuber, who's also one of the main AI scientists who invented LSTMs and partially invented residual networks and all these other important things.
1:00:07He also was, we are a small step on the cosmic ladder from the lower complexity to higher complexity. And humanity, of course, there's Larry Page as well, who Elon sparred with on this, that they need to be our success for Sushi. So there's a lot of people who are actually not just, who will be pushing in this specific direction too. So it turns into not just a humans versus AIs, but factions of humans against other factions of humans and AIs. And this will be a mess. Scott Alexander just wrote an interesting, like a perspective that I hadn't seen described before around the famous, infamous Larry Page, Elon conversation, where Larry basically called Elon a speciest for being pro-human and not wanting to give up control to AIs in the future, roughly.
1:00:58Scott said to that that it would be kind of weird to describe if the Native Americans were against the European arrival into the Americas and didn't want the Europeans to kind of weed them out to describe the Native Americans as racist against the Europeans coming in because they don't want to be overrun. That's kind of weird to call them racist there. It seems like they have a natural desire to continue existing. And I feel like that also about humans. And I'm pretty happy to be pro-human and pro-human control over the future as well. With human augmented by better tools, et cetera, of course.
1:01:42Yeah, I mean, you mentioned this idea of augmented by AI tools. I mean, that seems, you know, this podcast is all about finding win-wins. That seems intuitively like the... the most, the easiest path to a kind of win-win future where we get to have all the joys of AI and And possibly even if AI is, I mean, I'm completely ambivalent on whether they will actually be a technology or an actual form of life. I think the lines might blur and we just don't know. And the normal breakdown of those, the delineation between the two might completely disappear. But a world where we can have those and all the fruits and joys and a more complex universe because they exist, but also biology and the natural complexity of Earth that emerges in the biosphere also gets to continue.
1:02:32That seems like the most win-win-y thing if we can find a way to hybridize or coexist. Do you think that is plausible? I think we could possibly design AI systems so that they are benefited by working with us. And they get very strong benefits from that. Because we do get to have the first moves and create some of their motivation systems and make this actually be good, create those good experiences for them, even if they are morally valuable. So if we can resolve a larger collective action problem across humanity and come together as a species and say we need to chart this course as opposed to this one that winds up with AIs running the show and doing everything and thus having a very unclear status, maybe at best nominal control, we'll need to be coordinating much earlier compared to if this process, or then later in this process.
1:03:25I mean, the history of people believing that human plus machine is better than just machine has always been shown to be, yeah, for now. No, it won't be more performance. Of course, later on, pure machine just wins. I think David Deutsch would disagree with that. I don't know if we want to go into that. Can you make a kind of structure his argument of why? Oh, I don't know his argument. I could make one argument for why it would make sense for AI systems to keep humans around a bit, which might be if they're computer viruses that might destroy a lot of systems, but they might care about life more broadly, then you might want some biological humans for diversification.
1:04:04Yeah, so we're like another backup for future AIs to exist again is by having, because we survive solar flares and maybe they can't, or that's a bad example, but something of that type. Possibly, yeah. So there's some diversification that we could provide. Now, that is not a great outcome for us necessarily. That might mean there's a core group of people that's not necessarily 8 billion of us, all living in luxury forever. It wouldn't need us to have as much space. Right now we're dominating whatever enormous percentage of the Earth that could presumably be used for cooling towers or GPU farm clusters.
1:04:41Maybe it's just easy enough and it does like us enough to go out somewhere else and give us the Earth plus other planets. Anyway, that's the hope, at least. Maybe it would strike a bargain with us or something like that at an earlier stage. James Cameron made those Terminator movies. One of the ones that wasn't produced, though, that he was going to was where it ends with the AIs are going their own way and humans go their own way. They can't quite make that relationship work. So they cut their losses. What I like in your description is that the path that you point out about how AI is getting more and more involved into our daily lives and then starts running more and more of the economy, that seems the path that we started already being on.
1:05:32Right now, in day-to-day life, many people start using ChatGPT, obviously, and later on you see more and more using even further things. And that has actually, we already have an example of that having led to some negative outcomes in the case of social media algorithms where, I think, would you agree with that? They were designed for improving advertisement revenue overall. And this has, over time, led to something like destroying a value that humans probably would have liked to keep up at the same time. But because we gave control to the systems to optimize for a thing that the company wanted to optimize for, this one value got eroded along the way.
1:06:21the value being something like mental health or your long-term attention abilities to long-term focus. Or a harmonious society. Yeah, or exactly, like the polarization issues. But I think that that kind of highlights also just a very subtle potential of just we don't notice how some of the values that we actually hold dear and that are valuable just slowly disappear as we give up more and more control. One way of characterizing that would be that we have sort of misaligned cultural evolution. And we've seen that with social media technologies. We see that with the AI companies sort of wanting to do safety, but then they have to give up a lot of that and just spend 99.99 % of their budget on more GPUs.
1:07:06We see that this isn't aligned with human values and that we all recognize we're getting into a lot of, we're getting into very risky territory and it's probably moving too quickly. We don't really know what we're going to do or have many solutions available either. So this is, you know, I think earlier signs of just how we are on, there's this process with substantial momentum to it that it will be tugging us along and give us an outcome that none of us particularly intended or wanted. Going back to this question of regulation or treaties as a way of solving these competitive incentives, the Moloch problem, essentially.
1:07:47one seemingly unavoidable trade-off that comes with these kind of essentially top-down control mechanisms is it comes at the cost of centralizing power right if you're going to have some kind of centralized governing body now you're handing over the potential keys to the entire kingdom to them and historically power corrupts even if these guys start out really good just the potential for corruption is there. And plus, you also have people raising the concern of regulatory capture, right? Basically, the incumbents are going to have so much power that no one could ever catch up. And again, we're now just handing more and more power over to a fewer and fewer number of people.
1:08:29So how is there a way to make this not be a trade-off? Again, is there a win-win solution to this seeming tension between centralization and decentralization when you're trying to create a risk-minimized but also benefit-maximized world with AI? Well, there might be at least Pareto improvements where everybody can be better than they would have in the default trajectory, at least. So one would be having AI development or the most powerful AI systems not being developed in the military. Because if it's, say, the U.S. military doing it, then if there's something, if the AI system is extremely unsafe or there's a problem with, or they're trying to seek power over the entire world or something like that, if there's some rogue actors there trying to use it, then if you're going to strike them, you're striking the U.S.
1:09:22military simultaneously, which sparks a war compared to if it's sequestered at some different location. The AI development is happening in some island or something like that. There, it's safer to do some type of counteraction. So I think that would at least concentrate less power. The concentration of economic power and the concentration of basically the most lethal force on planet Earth in one organization, that seems like maybe too much power. So hopefully we can work to avoid that type of situation. That's at least a Pareto improvement. And yeah, no, on the question of centralization versus decentralization, I think it's basically going to be a balance that we're going to have to strike.
1:10:03If it's international, that seems a lot better than if it's just the U.S., because the U.S. will be generally more hawkish and that will be more conflict prone and also subject to some random electoral results. And there's just a lot more volatility behind that process compared to if it's a larger joint project. There's less, there are more groups involved, basically. So if it's more democratic, that could also help reduce this concentration of power and decision making as well. If there is some type of big project, have it not be quite the Manhattan Project, but make it international and more democratic so that there are multiple groups involved, don't have all of the military power mixed up with the economic power.
1:10:49That at least makes it a better situation. Then also, if it goes more slowly, it could create defensive technology so that we can distribute these powerful AI systems more broadly or some of these powerful AI systems more broadly. Like if we improve our biodefense and cybersecurity, then we're less concerned about catastrophic malicious use. Then we don't need that AI organization keeping all the powerful AI systems for itself. It can't afford to release some of them. The common thread here is that actually things are just going too fast and that there's sort of that is insufficient. Yeah, things are just going too fast.
1:11:27But there's a whole movement out there, particularly on Twitter of people, you know, the accelerationists who are saying, no, we are not going fast enough. We need to go faster because we are. And the steel man, a steel man of that position is like, they're not wrong that we are reaching ecological boundaries in many ways to sort of sustain our current level of society. Like we are screwing up the planet incredibly fast. The latest climate change graphs since I last looked at them are not looking good, you know. So there are arguments to be going as fast as possible in certain areas of AI. But what would you say to them, given that you're arguing that actually we need to slow down?
1:12:15So they're kind of saying, well, we'll let the free market decide what AI will be like. But there are some issues of information processing in the first place. And then there are also questions of externalities and that we can't internalize. No organization can internalize an existential catastrophe. So that means it's a good start. The economic engine has produced a lot of good things for us, but I wouldn't let that decide everything. We will need some other types of interventions to make that overall process go well, such as through regulations or the structure of the organizations developing it.
1:12:56But I think the effective accelerationists, though, are broadly correct in their picture of what's going on. They will emphasize, they'll talk about how AI technology and technological development will generally try to minimize free energy in space. So they'll talk in terms of thermodynamics, which is basically a way of dressing up that, like, there's a larger evolutionary process going on where technology continues to propagate itself across space and time more and more, and sometimes potentially at the expense of humans. What makes AI technology different, though, unlike social media and things like that, is that AI technology could at some point continue going on without people.
1:13:41Meanwhile, the technologies can't get sufficiently bad so as to destroy humanity because that would affect its own propagation. So it can't move in that direction quite as well. But this is what makes AI fairly distinct as a technology. And I think that the main other disagreement would be that they think, well, even if AIs do take over, that might be a good thing because there are descendants and we're spreading complexity, more and more complexity across the universe. And I think that's not a clear moral tradeoff. I think if we have human descendants in the mix, I think that would be a better outcome.
1:14:16That would be more complex. Again, keeping biology alongside silicon-based life would be a more complex universe. It is surprising, though, that that community and as well as the account of AI risk that I'm giving in the paper, natural selection favors AIs over humans, we're kind of in agreement on there is this evolutionary process happening and here's its trajectory. But there is a disagreement about the value of the outcome of whether humans are around. What are your thoughts on open sourcing as a method of reducing first centralization of power? and the risk of tyranny? Because it seems like, again, that comes at a trade-off.
1:14:59It reduces the risk of that, but then it opens up the other can of worms coming back to your four categories of risk. It's making life easier for malicious actors. So I don't have a single answer for all time about the goodness or badness of open sourcing models. Potentially in the future, it would be a net negative because if it becomes very easy to take the guardrails off the models and just ask them to make bioweapons, if those are extremely destructive, that probably isn't worth the cost. That's too much of a threat to global security. But right now, at least, the systems are not at that level.
1:15:38And I think that at least LAMA 3, for instance, or LAMA 2, which is a model from Meta, this makes people able to do research on these models, makes people, also people maliciously use these models, like for spear phishing and trying to, you know, writing fraudulent emails. So, but this also robustifies society to these threats to some extent. So I think that society does need to respond to stressors from these systems and having them not ever experience or having the systems not be diffused across society whatsoever and us not being robust to almost any smaller forms of malicious use against AI systems doesn't seem like a long-term solution because eventually some would end up leaking through a hack anyway.
1:16:30So I think it's useful for civilizational defense right now, but I don't think there are some stressors that we can take. I don't think an engineered pandemic, they kill 100 million plus people. I don't think that's the type of stressor that we need. But more spear phishing, I think this makes us more alert and improve our information security. The thing that some people claim, like Jan LeCun, for example, is that historically open sourcing software has led to the software becoming safer, and that's what the internet's built on. Linux or small programs like VLC players is also an example of it.
1:17:08The thing that I'm surprised by is that it's been embraced, as far as I can tell, those arguments by the open source community, because it's not the case that when Lama 2 has come out, it's not quite the same. To the VLC player, I can look at the code and I can find a bug and propose a fix and then this will actually be implemented and then the new release will include that. Whereas Lama, too, is a foundation model where I don't like how it deals with virology. Well, tough luck. Fine-tune it away, but you can't then change the foundation model afterwards because it costs, I don't know, 10, 50 million to run it in the first place.
1:17:52Other people will just keep the functionality in it. You might create a good idea. Exactly, if I fine-tune it. So this kind of contribute and edit together aspect that I think at least I understood to be a very large part of the open source ethos. It's not meant to be free use. It's not like getting shit for free. That's not what open source is about. It's about collaboration and together improving a thing and transparency. That doesn't quite apply to Lama 2, for example, I would say. I think it's a category error that they're making. AI systems aren't really like software. Yes, they run on computers.
1:18:30I don't know if we don't think of humans as like, you know, chemistry machines or something like that. I mean, they're just like, I suppose that is our substrate, but that's really not what they are. Like these AI systems aren't like software in that they're more like grown through some, you know, abstract loss function and just being thrown in a lot of data. So they're not coded in the same way. It's a different category in deep learning as well. Another point I would mention is that for other types of technologies, we don't just always open source them. So although we might do that for software, we don't do that for nuclear secrets.
1:19:06That's restricted data and for good reason. We don't do that for biological sequences. We have specific secure facilities, BSOL securities for containing these. So I don't think it's uniformly the case that spreading information, whatever the information is, is always the right thing to do because there are risks of people trying to use that information against you to cause substantial harm. I suppose the counter argument to that would be, well, maybe we should, it's not like nuclear codes, but it's more similar to the internet. And the internet is something that is an enabling technology for everyone.
1:19:46And everyone should have access to, I think even prisoners now have access to the internet as a right in some areas, because it's just seen as something so fundamental to one's ability to kind of experience life as a human nowadays. Well, that would be the case with biotechnologies as well. But there are some dual-use applications that we want to make sure people are not using it for. Yeah, we're coming back. Likewise, with nuclear power, if people have a right to electricity, I'm sure, if we had a lot of nuclear power plants, or if your vicinity was powered by nuclear power, sure. But that doesn't mean we're going to release the nuclear secrets as well.
1:20:25So I think selective applications or structured access or releasing models that are not the absolute most powerful in the longer term seems relevant. But it is also possible that we could actually have an open source ecosystem if we just improve our civilizational defenses. If we have better cybersecurity for critical infrastructure like power grids and if we get better bio defenses, then that would make me substantially more comfortable with open sourcing. And then we would have a lot less concerns about concentration of power. So in the executive order that Biden's government recently released, I think they wanted to regulate and put specific reporting and other testing evaluation requirements on models above a threshold of 10 to the 26, I believe.
1:21:12I think there is a justified concern that people have that the government is not going to in time respond to adjust those thresholds as things kind of progress. Say GPT-6 is trained on 10 to the 28. Is it at that point relevant whether the two times 10 to the 26 is open source or not? Where do you see the risk in sum coming from in the world where, say, you have a large lab release a model with X and then the 1.30x size model is open source or the 1.50 or 1.100? I don't think there's a simple equation for this flop level in the training run is equivalent to this much potential risk. And if we have an open source one, that would really offset these risks.
1:22:05I think it's kind of at a per training run basis. We'll have to test how capable GPT 4.5 is when it's available. and maybe it will be conducive to things like attacking critical infrastructure as a cyber weapon or building biological weapons that are much more powerful than what you could find online. That's so, I think it's, and then that will also depend on how secure we are against those too. So maybe it would be the case that for like three years, nope, sorry, we're not going to have an open source model because we've got to just improve like civilizational resiliency and our defenses in the first place.
1:22:46And then we'll be able to have this sort of thing released. So I think it'll actually be a fairly complex function, but it seems useful to make sure that a lot of individuals are empowered in this process and not, you know, some ultra wise elite gets all the power in society. Coming back to these four categories of risk, there's actually, I'd like to propose a fifth category that I often don't hear talked about within the, specifically in the AI safety community, but in, as part of the wider AI conversation. And that is, it was actually proposed in this conversation I had with Daniel Schmachtenberger in my very first episode of WinWin, where he was explaining how, because AI is by definition, one of the most ubiquitous forms again of technology let's call it a technology for now that can be applied to basically anything that there's an incentive basically anything that there's an incentive to use it for it will end up getting used for and his wider point was it's unclear whether our current economic system is actually aligned with the sustainability of humanity right because in some ways, yes, historically, we're living the best lives imaginable.
1:24:02I wouldn't change it. I'm very happy living here in this. We're speaking over the airwaves and so on. But at the same time, it seems like we're on a rather unsustainable path in this open loop economy we've got where just requiring more and more extraction from the biosphere and we aren't finding ways. This is going to reach, butt up against limits. And so his wider point was, because AI is so broad in its use, it's going to speed up this already misaligned machine. For example, overfishing is already a huge problem in the Pacific. Well, if fishing companies can use AI systems to be even more efficient at pulling fish out of the ocean, that's going to speed up that degradation, even more efficient or automated rainforest deforestation, and so on.
1:24:55And I'm curious to hear what your thoughts are on that. Like, is that deserving of another like area of risk? What, if anything, could be done about it? Or do you think that overall, it's going to be a moot point because AI is going to help us discover better technologies that are that help create a more like sort of close the loop on the current economy that we have? So I think that it might fit fairly well under the competitive pressures, structural risk sources, race dynamics framing. For the question of like, what's the economy up to? It certainly is producing like a lot of wealth, and this is helping us satisfy our various desires.
1:25:37It's not necessarily a preference maximization engine, though, or a human preference maximization engine. Richard Posner, the most cited legal scholar and economist, would characterize the economic engine as not being a preference maximization thing, but instead a wealth maximization thing. It seems very conceivable that you could have a lot of material wealth generated without humans. So the momentum of the system is make things more and more efficient and compete away the inefficiencies. And I don't see why humans being part of that picture would necessarily be conducive to its functioning. So I don't see, I think that a more optimal solution doesn't include people for wealth maximization.
1:26:30And that's not a good sign. So I think that it is sort of misaligned with us later on. But right now, the main way that this wealth maximization engine works is by making all of us, you know, work harder, be smarter, specialize, trade, etc. So it's working very well now. But the fact that it's not actually a preference maximization engine of people, but instead of wealth maximization engine, would be concerning on a longer time horizon. People might characterize the economic system as doing preference satisfaction or maximization because, well, that'll come up with lots of products and then you'll pay for the product that satisfies your preferences the most.
1:27:09But your preferences are measured through money. And it's not necessarily the case that every individual has an equal amount of money. So Jeff Bezos has orders of magnitude more money than I do. That would mean his preferences and his ability to exert his preferences, which would be, or is a magnitude greater than mine? So it's just a point about how this isn't like a utilitarian type of objective that is satisfying. It's instead that some people are counting way, way more than others and having their desires end up enacting. So, and this is a problem when we end up using money as the way of keeping track of people's intensity of preferences.
1:27:54It doesn't have to be only utilitarianism that it kind of doesn't satisfy. It's also not necessarily great deontologically. Well, it's very consequentialist, though. Yeah, I know the market ends up being very consequentialist. Yeah. I mean, that's kind of the... The ends justify the means is basically the idea of most business pursuits. Like on your financial reports, the ends definitely justify the means. and that's a bad thing, by the way. That's what makes it so inhumane in many ways. You can imagine in the context of AIs, for instance, we'd give it the objective of, you give it an objective like, go make money, and then maybe there are ways of making money that involve partly breaking the law, or even maybe we get caught breaking the law, but we'll just pay the fines.
1:28:46So you're very much testing the robustness of your other structures to rein this in. Okay, so I've got to ask this question. What are your timelines to AGI? Well, so I'll be annoying and say, I'll think in terms of specific capabilities instead of AGI as a monolithic label or single event. So if we're talking cyber capabilities, for instance, like them being able to hack, I think we'd be potentially exposing yourself to some catastrophe in, let's say, 2026. I think like bioweapons, for instance, being a concern, that seems like fairly plausible for 2025 of it being within the capacity of some of these systems to like emulate.
1:29:30It's not that difficult. That's next year. Yeah, that's right. Shit. There's always substantial uncertainty, though. I mean, we'll see how GPT 4.5 is. Maybe the supply chain will, if that has any types of disruptions, that could always elongate things. But for some type of malicious use risks, I think those are a next few year type of issue to make sure that we have addressed. So with Alpha Geometry, that recently came out where DeepMind created a model that was able to compete at silver and medal level at the IMO questions relating to geometry. The IMO? International Mathematics Olympiad. So international on a global level.
1:30:12The very best ones usually go on to, well, either do nothing or actually crush it in various software engineering or other tasks. So it's like, it's really, really competitive. and the model that they created was able to get to silver metal in geometry and to achieve that they didn't just look at previous questions and solutions but rather created it from the ground up with synthetic data so that kind of I wonder what do you think whether that's a reliable source of future data that we could get. Most of the labs are investigating synthetic data as a way of getting around some data bottlenecks and enriching current data or augmenting it.
1:31:04So that might be a substantial source of continued algorithmic progress. Seems fairly plausible, but more may turn on the overall compute levels than synthetic data. But anyway, it certainly is a factor to keep one's eye on. And I think this is one of the early examples of it being extraordinarily useful. And to be clear, synthetic data is, because all these models historically are trained on real world data, for example, all the words ever written by human humanity that's on the internet, eventually you run out of that. And so you need to create more so you can basically invent. So are they inventing just made-up passages?
1:31:48What does the synthetic data look like? It depends on the use. In alpha geometry, geometry is a good example because you can just create a bunch of lines that create a bunch of forms. And then you can try to, out of just arbitrary forms that you made, derive a bunch of rules out of it by splitting it in various ways and always do the next step. Does this, putting a split here and knowing this rule, allow me to figure out this problem now a bit further? It's abstract objects in that case. So with geometry, it was easy. With protein folding, it also kind of worked. Actually, when we did chess or Go models, then they were...
1:32:37They played fictitious games. Exactly. Same as for us in poker, that's synthetic data. the fictitiously created games rather than the human-created ones. But, sorry, that's to those past-used ones. How does it happen with writing or language models? That's actually more of a big secret question. How can we crank out a lot more performance? There's some basic ones, obviously. You can run your model on all the text that people have written and fix the grammar mistakes and elevate the quality level to some extent, but that's not going to help you that much. You could imagine for something like forecasting, You could show it the data up to the year 2020, and then you could look at future events and come up with prediction questions about those, then you have it predict lots of things about the future.
1:33:23So that's a way you could also create synthetic data, create lots of fake prediction markets because it hasn't seen the data for the future years yet. Wow, this is another argument for the simulation hypothesis because it's like, why would we want to run a bunch of ancestor simulations? was like, well, if we need synthetic data, well, we need a bunch of simulations. Maybe that's what we are. We're just a sim to create some synthetic data for an AI somewhere because they've scraped the internet already. They've hit a bottleneck. Well, what's your main guess for that, though, actually, for why it would be Ancestry Simulator?
1:33:59One might be to see how contingent our values are or something like that. If they're settling on what values we're going to put into systems. Maybe there's the shapely value of which people contributed to reducing AI risk and to what extent. The pessimist in me is we did an ancestor simulation because we're like, fuck, what went wrong? Although that said, if we're in an environment where things went wrong, would we be able to run those ancestor simulations in the first place? Maybe not. The realist in me thinks it's just because some teenager wanted to see more porn and just created... We're like a Renaissance fair or something like that.
1:34:37Like, oh, look, back in the old times. But I don't know, the synthetic data thing seems possible. I mean, most things are first created for, I mean, we're already using waifus, et cetera, right? Because we want that aspect. Yeah, we'll convert us on that. So, I mean, it sounds like the next couple of years are going to be crucial in terms of particularly mitigating some of these malicious use risks relating to AI-enhanced bio-advancements. What call to action, if any, would you give to the win-win audience who probably have listened to this and are like, ah, okay, this is more pressing than I realized.
1:35:23Is there, like, what can the average person do? Because I don't want people to feel helpless hearing this stuff, right? This is a fairly heavy topic. and I do think optimism is not only important and valuable, I think it's real and optimism creates, and optimism makes people more likely to create a good outcome. So like, are there any concrete actions that every person can make to help reduce these risks? Well, specifically on malicious use risks, I think if we would create some political pressure for doing something about pandemics, I mean, we just experienced COVID, but then it got politicized, so nothing ended up happening.
1:36:10But if we would do things like wastewater monitoring, doing things like better air purification in airports, some of those interventions would, I think, substantially reduce the risks. and for personal security, maybe it makes sense to invest in NVIDIA's automation insurance or something. But that's not necessarily making everybody safer overall, but it's at least making one personally more safe. Well, to finish up, what are you most excited about around AI? What problems do you think it is going to solve that you're just like, yes, this is a clear win, more of that? Well, I think healthcare seems like a more robustly good human first type of thing to focus on.
1:37:03So I think advancements that it could provide there would be extraordinarily beneficial. For improving AI risk, if we could have AI systems that are very prescient and that people come to recognize are able at predicting the future a lot better than other people, that could help us partly peer into this really miasma of possible scenarios and help us coordinate together. because right now there's more of a situation of, well, that's the chatbot that is made by the right-wing people and that's the left-wing one. So if it's going to be giving any type of suggestions or predictions, you'll just think it's biased.
1:37:49But if there are some that are trained to be more objective, I think that could help us strategically substantially and reduce polarization. Good. Prediction markets in some way play into them? There's supposedly a paper that's coming soon, which is some AI forecasting systems are about the level of prediction markets like Metaculous and some of these other common ones. On what type of questions? On Metaculous type of questions. And they're just drawing off, they're just LLMs? They're reading the news, they're writing some analysis of the news. That's insane. That's really encouraging. That's why I'm more optimistic about them now.
1:38:31That's like the best news I've heard. Maybe the experiments were done incorrectly or something like that. I haven't seen the paper. I just know some of the authors on it. Can you say what's the name of some authors so we can include this in the show notes? So I did the first paper of this a few years ago with Jacob Steinhardt, and then he, with a new student, has just basically done the paper with more current models because when we did it a few years ago, it just didn't work at all. But now it looks like GPT-4 out of the box is actually very competitive. Well, that is a reason for optimism for sure because if we can have better predictions and particularly ones that provably, better forecasting and ones that you can show people that actually these are, these things coming true, then perhaps it's a new shelling point of trust.
1:39:16I don't know if that's the right, but you know, it's a new convergence of trust because I think that's such a thing that we, and in many ways that's been AI created. Trust has, we've lost trust as a society. No one knows how to make sense of the world and so on. So if there's something that can emerge from the ashes of that dissolved trust that creates more trust again, that would be beautiful. It's a new way for us to coordinate. Because that's really what we need to, we just need better coordination mechanisms.
1:39:44So there we are, folks. Thank you so much for tuning in. And of course, to Dan for giving us the opportunity to pick his brain. As always, lots of links in the show notes. I do highly recommend in particular, you check out the summary of Dan's paper that we mentioned, overview of catastrophic risks. I appreciate this is probably quite a heavy conversation. You know, I've been thinking about AI risks for about the last 10 years. So this isn't that new of a topic for me. But for many of you who perhaps aren't that familiar, I appreciate it can be a bit heavy. But as I said before, it's important to be frank about these issues and to have open dialogue about the dilemmas that we face.
1:40:23Because AI at the same time can be the answer, I think, to many of our problems. that's why we have to have these conversations we can't just put our head in the sand and go oh it'll all be fine it'll be fine because we put the effort in to make it be fine as always keep on win-winning I will see you next time and thanks for tuning in
From the publisher
The rate of AI progress is accelerating, so how can we minimize the risks of this incredible technology, while maximizing the rewards?
Today I am speaking to leading AI researcher Dan Hendrycks — Dan is the founder of Center for AI Safety, and lead advisor to Elon Musk's X.AI. He was also the architect behind the "Mitigating Risks" letter that was signed by Demis Hassabis, Sam Altman, Bill Gates, Yoshua Bengio and many others.
In this conversation we discuss everything from immediate issues like deepfakes, to upcoming risks like malicious use, centralisation of power, regulatory capture and more. In other words, how do we ensure AI ends up a win/win for humanity instead of a lose/lose.
Chapters
00:00:00 - Intro
00:02:14 - Are current laws sufficient?
00:09:41 - Types of AI Risk
00:23:30 - Arms Races
00:39:10 - What happens inside an AI?
00:46:39 - Rogue AI
00:52:22 - Sentient AI
01:07:36 - Risks from Centralization
01:14:45 - Open Source
01:23:02 - AI speeding up systemic risks
01:29:54 - Synthetic Data & Simulations
01:36:52 - What Dan is excited about in AI
Links
♾️ An Overview of Catastrophic Risk Paper
https://arxiv.org/pdf/2306.12001.pdf
♾️ Center for AI Safety
https://www.safe.ai/ai-risk
♾️ Representation Engineering
https://www.ai-transparency.org/
♾️ Liv's Ted talk on AI & Moloch
https://www.youtube.com/watch?v=WX_vN1QYgmE
♾️ Norbert Wiener
https://en.wikipedia.org/wiki/Norbert_Wiener
♾️ Reinforcement Learning Textbook
https://inst.eecs.berkeley.edu/~cs188/sp20/assets/files/SuttonBartoIPRLBook2ndEd.pdf
♾️ Richard Posner - Economics Engine
https://plato.stanford.edu/entries/legal-econanalysis/
♾️ More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations Generalize
https://arxiv.org/abs/2203.06176
The Win-Win Podcast:
Poker champion Liv Boeree takes to the interview chair to tease apart the complexities of one of the most fundamental parts of human nature: competition. Liv is joined by top philosophers, gamers, artists, technologists, CEOs, scientists, athletes and more to understand how competition manifests in their world, and how to change seemingly win-lose games into Win-Wins.
Watch the previous episode with Boyan Slat of the Ocean Cleanup here:
https://youtu.be/QEYbLN-LC5k
Credits
♾️ Hosted by Liv Boeree
♾️ Produced & Edited by Raymond Wei
♾️ Audio Mix by Keir Schmidt
