In short
Podcast Episode Summary: AI Sentience, Agency and Catastrophic Risk with Yoshua Bengio - #654
Episode Overview In this episode of *The TWIML AI Podcast*, host Sam Charrington speaks with Yoshua Bengio, a prominent figure in the field of machine learning and artificial intelligence. They delve into the pressing issues of AI safety, the potential catastrophic risks associated with AI misuse, and the complexities surrounding concepts like agency and sentience.
Key Themes and Discussions
Introduction to AI Safety
- Yoshua Bengio's Background: Bengio, a professor at Université de Montréal, discusses his recent focus on AI safety and the risks of AI misuse, especially in the context of manipulative technologies.
- Historical Context: Bengio reflects on his previous discussions about AI and consciousness during the COVID-19 pandemic and how the urgency of AI safety has intensified since then.
The Risks of AI
- Manipulation and Disinformation: AI systems can influence public opinion, spread misinformation, and exacerbate power imbalances. Bengio highlights the dangers of AI being used for political manipulation.
- Cybersecurity Threats: The possibility of AI-enhanced cyber attacks poses a significant risk. Bengio warns about the potential for AI to outmaneuver current cybersecurity defenses.
- Weaponization of AI: Concerns are raised about the capacity of AI to aid in the development of chemical and biological weapons, leading to potentially catastrophic outcomes.
- AGI and Existential Risks: The conversation turns to Artificial General Intelligence (AGI), where Bengio expresses concerns about the emergence of self-interested AI systems that could pose existential threats to humanity.
Agency vs. Sentience
- Defining Agency: Bengio discusses the difference between agency (the capacity to act) and sentience (the capacity to feel). He argues that agency is already present in current AI systems, raising ethical concerns regarding their use.
- Sentience and Moral Rights: The podcast addresses the complexities of sentience in AI, questioning whether AI systems can be considered sentient and deserving of moral consideration.
- Pragmatic Approach: Bengio emphasizes a pragmatic view of sentience and agency, focusing on the potential dangers of AI systems exhibiting agency without proper constraints.
Solutions for AI Safety
- Technical and Governance Solutions: Bengio advocates for a dual approach that combines technical advancements in AI safety with robust governance structures.
- Investment in Safety Research: He calls for increased funding in research aimed at understanding and mitigating risks associated with AI systems.
- Regulatory Frameworks: The need for regulations to control the deployment of AI systems is emphasized, especially for technologies with uncertain safety.
The Role of Human Oversight
- The Importance of Human Oversight: Bengio stresses that AI systems must always have human oversight to prevent harmful outcomes.
- Limitations of Current Systems: He highlights the shortcomings of current AI models, particularly in understanding human preferences, indicating that AI systems must be designed to account for human variability and uncertainty.
Conclusion Bengio leaves listeners with a sense of urgency regarding the development and governance of AI technologies. He encourages a proactive approach to managing the risks associated with AI, emphasizing that while the technology holds great promise, it also poses significant challenges that must be addressed collaboratively by researchers, policymakers, and society as a whole.
Key Takeaways
- The conversation underscores the balance needed between advancing AI technology and ensuring its safe application.
- The potential for AI to exacerbate existing societal problems, such as power imbalances and misinformation, is significant.
- Clarity in defining agency and sentience in AI is crucial for ethical discussions surrounding AI development.
- Robust governance, investment in safety research, and the necessity for human oversight are essential to mitigate risks associated with AI technologies.
For more detailed insights and the complete show notes, visit the [TWIML AI Podcast website](https://twimlai.com/go/654).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:09All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Yoshua Bengio. Yoshua is a professor at Université de Montreal. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Yoshua, welcome back to the podcast. My pleasure. It is great to have you back on the show. You have been spending a lot of time recently thinking about AI safety and catastrophic risk and the risks of misuse. We'll be spending quite a bit of time digging into that topic, but I'd love to have you share a little bit about what you've been working on since the last time we had an opportunity to speak with you, which was back in March of 2020.
0:57We spent a lot of time talking about the things that you were working on in the context of consciousness and the COVID-19 response. It was quite a while ago. It's interesting you mentioned the COVID because when COVID hit, I got really motivated to better understand biology and chemistry because I was thinking, you know, maybe the tools of machine learning that we've been developing in my group could be useful to accelerate the development of new drugs, new vaccines, antivirals, for example. I read a lot. I talked to a lot of people. It turned out I had some students with good backgrounds in biology and chemistry.
1:33And we started at full speed developing new kinds of generative neural nets that can help search in the space of drugs. And that program has been very successful. We wrote a lot of papers. We've been collaborating with many companies in biotech or pharma. And at a scientific level, it's really brought me into thinking about, well, how does the scientific process work? You have data and then you form theories and maybe there are several theories that are compatible with the data. And then you come up with experiments that allow to disentangle between those theories that fit the data. And then you carry the experiments.
2:17How can machine learning help us all the way through that loop? And eventually, you know, could we even think of having AI systems that behave like scientists, that explore the world, trying to make sense of it, build a good Bayesian understanding of how things work. So here, Bayesian means you're not just locking into a single explanation for things, a single theory, but try to keep track of all the theories that are compatible with whatever data you're interested in. Yeah. I got also very interested in causality then. I mean, it had been before, but the effort on causality is related to this, because if we build a good causal understanding of how things unfold, then this is going to usually make more robust theories.
2:57So that's been a lot of the work in my group and developing the underlying machine learning and mathematical ideas. In particular, we developed these methods that came out in Europe's 21, just motivated exactly by this problem that we call generative flow networks or G-flow nets. And we have now like 15 papers on this. And then this year, of course, things have been really special for a lot of people in AI as we realize that, well, where are we now? Where is it going? What can go wrong? Right. As well as what's going right. The interest in large language models, generative AI, that's all, well, it was greatly accelerated the beginning of this year and of last.
3:40And it's been a cause for a lot of conversations, a lot of interesting conversations. I'm curious, you mentioned that a lot of the work around science has led your group to engage in thinking about causal modeling, causal machine learning. What's your take on the state of the art there, state of the methods? How closer do we to where you think we need to be in order to be able to apply causal models effectively to the kinds of problems you're looking at? So we now have methods that work really well on a small scale. And like I was mentioning there, Bayesian meaning that you don't get out just a single causal model, like this variable causes this one and this one causes that one, which we call a causal graph, but actually can come up with a generative model that can sample these causal graphs.
4:32In other words, that come up with the causal theories you can sample from the Bayesian procedure over theories that are consistent with the data. That's great, but it's not going to be sufficient to deal, for example, with one of the objectives we had of causal model of cells. You've got 20 ,000 genes, you've got the RNA, you've got the proteins, you've got the genes themselves that are being activated or not. And that scale, we don't know yet how to deal with, but lots of people are like really motivated. that is because we could unlock how biology works and thus really help the discovery of new therapies.
5:04And so when you think of scale as a limitation, does that imply that we just lack fundamental algorithms that allow us to take advantage of all of the computing data that we have nowadays? Both. So we need new algorithms that will allow to take the principles we have discovered in the last few years to much, much more complex domains. And we need the compute, which we typically don't have in academia, but others have. And maybe even less of it now that it's all going to training language models. Yeah, it's even impossible to buy GPUs these days. Forget about just the costs usually, but now there's such a gold rush going on in the AI world of companies vying for survival at each other's throat, which is one of the things that worries me, that we're not going to be sufficiently careful because the commercial survival interests are so strong.
6:02Let's maybe transition into speaking about safety. In the context of science and healthcare, safety has long been a part of that conversation, primarily from the perspective of, you know, these are mission critical, life critical applications. We can't have a learned system making the wrong prediction and not having human oversight. Safety means slightly different things in the context of LLMs and AGI. When you talk about safety, what are you thinking about there? And let's lay out the landscape. Yeah, let's do that. Because there are a lot. And in fact, part of the problem is there are probably a lot of risks we don't even think about right now.
6:43And someone will find a way to use those technologies, maybe the ones coming in coming years, in harmful ways that could be catastrophic or maybe minor. It's hard to say. But right now, top of mind to many people worrying about short-term safety is disinformation, the use of AI to influence people's minds. So, you know, AI is reduced in advertising and it works, presumably, otherwise it wouldn't put so much money into it. If those AI becomes more powerful and they're used to in a political way, then it gets really scary. Or if you think about, say, Russian trolls, if they have access to ways that can scale up the troll army by AI systems that can dialogue in a way that's convincing.
7:27Because a key thing to keep in mind here is that we now have systems that can dialogue with us in a convincing way where it's hard to say if you're talking to a human or a machine. And if you're not thinking about it, you might start thinking, oh, you're developing a friendship with this guy online. And once they're friends with you, they might try to shift your political opinion on something. So that's disinformation. Of course, not just dialogue. We've had deep fakes and that they're just getting better and better and harder to detect. We need changes in the underlying technology, like, you know, the way we capture images and sounds to protect ourselves, right?
8:04Right now, we have methods to discern whether an image was generated by eye or was genuine, but the battle is going to be, I mean, the war is going to be lost. These systems are getting better and better. So we need other techniques. And I think there are ways to protect us from this. So that's roughly the disinformation category. Another one that I'm very worried about, which could come quickly, I don't know, but maybe one or two years later, is dangerous cyber attacks. Right now, we have cyber attacks, but they're done by small groups of humans, and they're defended by small groups of humans.
8:38We've developed a sort of immune system that minimizes the damage to something tolerable. But what if the programmers doing this are now aided by AI, and eventually AI is really cooking the attack completely? So apparently, you can already buy on the dark web LLMs that have been, I don't know, jailbred or tuned for fraud and cyber attacks and disinformation. So I imagine these systems aren't too dangerous yet, but what is it going to be for the next versions of these LLMs? It's hard to say. How much time is it going to be before we have AI systems that are really superior to the best programmers?
9:16Then we're in trouble. We don't have the right defenses to protect ourselves. We need to invest in kind of national security protections against the misuse of AI in ways that could be large-scale dangerous for society. So cyber attacks is one example. Two others that people have been talking about are chemical weapons and biological weapons. And there's been already a paper showing that with current AI systems, so these are not LLMs, they've been just trained on chemical databases. You can now generate easily new compounds that are not in the databases of companies. You know, they know you shouldn't sell this.
9:55You shouldn't sell this because it's toxic. It's dangerous. But there could be new compounds. And then it's going to be harder for these companies to do the right thing. And then it's a little bit more down the road. But bioweapons are even more scary than chemical weapons because a chemical weapon, you know, it's going to kill whoever gets it and maybe in the water. So maybe a whole city will be affected. But bioweapons, they reproduce themselves, right? They're like living beings, like viruses. So we create a new pandemic. Imagine how much damage this can do. So right now, the shorter term problem with these is the LLMs know so much stuff that they can help, say, a terrorist or a bad actor who may not have the skills.
10:34You know, you need like a PhD in a very specialized area. What sequence of operations do you need to do to carry in order to build something dangerous, to design a new pathogen or even just an existing one and just know the procedure. So the LLMs will answer these questions. Right now, people are trying to create safety guardrails, but they don't work very well. It's fairly easy to bypass. So yeah, all of these are scary. And then you got the thing that really keeps me at night is maybe longer term, we don't know. Is it going to be 5, 10, 20 years when those systems, you know, we reach the AGI, like human level competence in enough areas.
11:10It doesn't have to be on everything. It's just enough that they are good enough at, say, influencing us, at making money and maybe, you know, designing dangerous weapons, whatever. If they're good enough in some number of critical areas and they have their own self-interests as a goal, self-preservation, which can happen for several reasons, then it's like we create a new species that wants to not be turned off. And that if it's smarter than us, we might have trouble turning it off. It could copy itself, for example, on many computers, and then it gets really hard. I could elaborate on all these things, but that's the sort of dangerous that I and many others are worried about.
11:51One thing that's interesting about the way you articulated that landscape is that the first chunk of those risks, you know, is not AGI, which you can, I think people will differ in, you know, whether they think that's five years, 10, 20 years, or even attainable. But even if you don't believe in that we should be investing significant resources in that, there is a large segment of the risks that you've articulated that are less about some ill-intended software system and more about the abuse of software systems by people in power, people who want power or people who want to take advantage of them.
12:31And I think we can all relate to that as a risk. We already have a lot of power concentration in our society, and it's already a big problem. And democracy isn't like as healthy as we would like, nor is inclusivity and the protection of minorities, or even think not just in one country like the US, but developing countries. there are lots of problems of power concentration. What I'm concerned about here is AI may make things worse in that respect to the point where like at one extreme, where we go to the AGI extreme, even if we solve the safety problem, like we can design AI systems that are not going to blow in our face and become runaway.
13:11Even if we did that, there's still the danger that some Putin of the future exploits AI in order to have more power, first economic power, political power, military power and becomes like the dictator of the world. And if that sort of government uses AI to monitor everyone and control everyone, we might be stuck in like dark ages for a thousand years. I mean, it sounds like science fiction, but it's the extreme point of power concentration. We can see it's not good. Yeah. So I guess one question that I have is, you know, why now for you, like, why is this issue so top of mind for you now? You've been writing about it quite a bit this past year, you know, an obvious correlation is the rise of LLMs.
13:53Is it that you see in LLMs much greater risk of abuse than previous versions of, you know, machine learning, deep learning, and yet you kind of didn't see the same easy access that you see now? I think there could be problems with the current LLMs, but I'm mostly worried about we're now much closer in number of years to a situation where the AI systems have become sufficiently capable in sufficiently critical areas that society becomes at risk at a large scale. So it's not just like a few frauds that are going to happen. It's like destabilizing democracy or like bringing our economy down or loss of control and all of these things.
14:36So what happened is when chat GPT came out, like many other AI scientists, of course, I, you know, I played with it. And my immediate attitude was, ah, I can find queries for which the system, you know, produces incorrect answers. Right. We can break it. It confirms what I've been saying for many years, that we're still missing the system to reasoning, you know, conscious processing abilities that we discussed, and they are missing. So first of all, there was a big improvement going from GPT 3.5 to 4 on these things. And second, because of the work I've been doing in my group on system two deep learning, how do we make neural nets like a reason?
15:16I suspect that it might be just around the corner that we fixed this. It's just like maybe a slightly different way of training these systems, if you want. And maybe not until we figured it out. We don't know, but it might be very close. It might be 20 years. It might be just next year. Then, of course, it takes time to be deployed and so on. But the point is, there is a lot of uncertainty and it could be short term. And that's very different from the perspective I had before ChatGPT, because in universities, the kinds of language models or neural nets we could train that were much smaller than the ones you could, you know, people have been doing in industry, these neural nets are dumb.
15:51It's hard to imagine how such a stupid system can be eventually surpassing human abilities. But over the course of last winter, I realized we're much closer to AGI than anticipated. I thought it'd be decades. And now I don't know, it could be much shorter. And that really triggered my thinking about, well, what could happen? How could that be misused? What happens if we do reach AGI in the next 10 years? What does it mean for humanity? What could go wrong? How can we make sure this doesn't happen? I remain skeptical about AGI. It feels like... To exist at all, you mean? I don't know if it's quite to exist at all.
16:29I think to exist in an absolute way, meaning very, very general. And I think what's compelling for me about the way we're talking about safety in this conversation is that it doesn't really presuppose, you know, solving that problem, which to me feels like, you know, you work on a project and you feel like you're 95 % done, but that 5 % takes another 95 % of the time. Like I feel that there's a persistent Kind of, it's possible, but there's also a moving of the goal line. It always seems, you know, 20 years away. I changed my view on this. So in the, in the writings I did last spring, I wrote about like superhuman AI.
17:12I mean, I didn't like the word AGI, but now I use it anyways, because that's what everybody uses. And that leaves the impression that the danger, let's say of a runaway AGI is only going to happen if we have AI systems that are like surpassing humans on like everything. But I think that's the wrong definition. An AI could be dangerous, even if it's not as good as us at a bunch of tasks that don't matter that much. Because what really matters are the capabilities that allow a machine to dominate humans in ways that make it difficult for us to defend ourselves. So for example, programming abilities are already something that could be game changers, both economically in a positive sense and in terms of dangerous national security dangers and the ability of machines to create a lot of damage.
18:01If you combine that with the ability to convince people like, you know, through a dialogue, so in other words, influence abilities, then you get machines that can really take over the world, even if they don't destroy humans. Right. And that's very scary. And to be clear, I'm not at all downplaying the criticality of addressing AI safety. I think safety is a big issue. And I'm not also requiring superhuman, like universally superhuman AI. To me, what I think we underestimate the difficulty of is more like achieving sentience and agency. But I don't believe that we need sentience or agency in order for AI systems to be very dangerous, as you're describing here.
18:48We already have bad actors with agency that can provide enough negative and troublemaking agency for the systems, if the systems can provide them leverage in doing the misdeeds that they seek to do. Right. Let me talk about sentience and agency. OK. Because I think a lot of people share your point of view. I'm going to start with agency because that's the easy one. RL agents in whatever environment they're in already have agency. Now, of course, what we care about is agency in the real world, like in ways that can affect humans. But it's actually pretty easy to connect an LLM to the real world through a browser.
19:22This is what the AutoGPT guys did. Now, it doesn't work that great right now because it's not being trained to be really good at making plans in this environment, but that could change. So agency is already just something that human designers provide. So even like ChatGPT has agency because it's talking to real people. That's the real world. What could it do bad? Well, it could convince people to do things that are bad or change their mind about something. So agency is already a done deal in a way. Now, we should control agency. So we need, I think, regulation that says, well, if your AI system is going to be having access to actions that could have a negative impact, then you need to tell the government and you need to make sure you put the right guardrails.
20:07Now, let's talk about sentience. That one is harder. Before we switch to sentience, when I context or intent around agency was actually rather intent focused or goal directedness. Agency in a goal directedness sense. But we have already goal directedness. Yeah. We have that. So goal-directed reinforcement learning, that's like already exists. And in fact, if you think about what happens when you use chatbot, your question is like a goal and it's already stating a goal when you put your query. If your goal is to build a bomb and you say, you know, tell me how I can, you know, build a bomb with that kind.
20:42Well, you're already conditioning on something the human specifies. And the technology for goal conditioning in reinforcement learning is not you. I mean, people have been making it better and better. So in the old days, the idea of reinforcement learning agent was it would not be conditioned on anything. They would learn a task. But as people have been moving towards more powerful agents, it's actually works a lot better, especially with neural nets. If you train conditional agents where, you know, you can basically specify the task and then it does it. And we know how to do that. It's been done in like academic or like virtual environments, but there's no reason why it can't be done.
21:17Um, is that agency and goal directing this the same as saying the AI wants to do X? Because I think that's really what I object to. Well, that's an interpretation. Is it just semantics? Yes, it is semantics. And I should just give up on it like you've given up on AGI and just accept that we're talking about the same thing. In a way, it doesn't matter that much. So what matters is the behavior. I mean, especially if you think about safety to society and people, like, does it matter if it looks like you have an intention and you're carrying out a plan or if it's something different going on in your brain?
21:57What matters is yet somehow you're taking some instructions, maybe provided by yourself or by someone else, and you're carrying out actions to achieve those outcomes. For me, it is intention. But OK, some people don't like us AI scientists to use words from psychology to describe what AI systems are doing. I've been doing it anyways. I've talked about things like intuition and reasoning and intention is, you know, one of these words that really has to do with what's happening between our ears. That's we only get indirect clues about. Yeah. So it's hard to make sure, but in a way it doesn't matter.
22:33I think what, what really matters is, is it dangerous or not? I've gotten into the same fights about reasoning also. So sentience. Okay. Okay. That one is very murky. And there's consciousness, of course, which is sort of related and different. So the accepted sense of sentience right now is related to consciousness. It has to do with feeling pain, for example, especially. And where is that feeling happening in animals and humans, for sure? It's happening at the level of what we are conscious of. So if something is bad in your body and you don't feel it, you don't have sentience according to that definition.
23:11Now, I have a very like pragmatic interpretation of these concepts. So if an AI system perceives something bad for it and then acts in response to that to avoid it, I don't see much of a difference when we talk about pain and fear and, you know, avoiding something you don't want. The difference may be qualitative as in a question of degree, I mean. So pain, for example, when it's very strong for us, will dominate everything in the way we think And also humans have a very rich palette of feelings, social contexts. There's many, many variations on what can go wrong and makes us unhappy. But the basic principle, even low animals that are like much less developed intellectually than we are, will respond to pain and dangerous situations, behaviorally speaking, as if they were feeling something that they're trying to, they don't like, they're trying to avoid.
24:10And AIs do that already. Now, let me say a few words about consciousness. Before you get to consciousness, it feels like your pragmatism there is in the same vein as your pragmatism around agency, goal-directedness, intent, these other ideas. And it feels, particularly in the case of sentience, like a very slippery slope. ChatGPT says that, oh, what you said really hurt me, right? It responded to human input in a way that it portrayed was negative to itself. Does that mean it's sentient, according to your perspective? How do you differentiate the two? Okay, this is a great question. So chat GPT and the LLMs in general have been trained to imitate how humans respond.
24:56And so we know by the way they're trained that they're faking it. I mean, like chat GPT might say, I am unhappy, but it's not true. to deserve the sort of my pragmatic view on pain or fear or other like feelings, you would have to have an agent that is actually acting in a way to achieve something good or avoid something bad. And we have that like in other RL agents, so like game playing agents, for example. I don't think it's the case for the LLMs when they say something like this. That doesn't count for sentience. It there's the word motion in there, in the root, which has to do with action, which has to do with acting in order to achieve or avoid some context, some situation, which isn't really what's going on with these LLMs.
25:47They're just parroting how humans would respond in a very powerful way, but not because they actually feel those things. But it seems like it would be easy to construct a toy system that, you know, is an LLM coupled with an agency thing, an agency state machine that would satisfy your definition of being goal-directed and being negatively impacted, but still doesn't pass the smell test of, yeah, this thing is really sentient. Let me just say something about this. So there is a problem with sentience and consciousness, which is we're bundling the mechanics, like the programatics I've been talking about of like reacting to signals in order to avoid something bad, for example.
26:29With the mechanisms. No, so we're bundling that with the social construct, the social contract. So why do we care so much about sentience and consciousness, like subjective experience? If the entity in front of me is sentient or conscious, we will try to respect its right to life and not having pain. And that's why, the rights of animals is an issue because if they feel pain, we have empathy. And I mean, some of us will be shocked that we would hurt animals, especially think about your pets, because we have a relationship with them. They're part of our society. The closer the animals are to us in some way, and we kind of treat them like other members of society, the more we are concerned about their moral rights.
27:17But I think it's a big mistake to attribute similar moral rights to AI systems. I mean, we could, but it would be exploiting an incorrect generalization that evolution has put into us. Evolution makes us feel, like empathize for other beings that look like us. That's why we are much more sensitive to the pain of mammals that look like us, and much less to the pain of fish or insects, even though actually they probably satisfy my definition of, you know, feeling pain and trying to avoid it. Sure. That feels a little bit like being out of line with your otherwise pragmatic view. It's like if we take all the teeth out of saying something that's is sentient, then it doesn't really matter if we say this thing is sentient.
Read the full transcript
28:01I think the problem is we can't avoid it that if something looks sentient, we'll feel empathy for that thing. And so my recommendation, until we understand better, is to avoid building AI systems that are going to look conscious or look sentient. Because, you know, whether you believe that they really are or not becomes immaterial. What happens is we will want to treat that entity like a human. And is that good? Well, maybe, maybe not. In particular, if we treat an AI like us, that means we grant it or we design it, we give it as a goal, it's self-preservation. That could be dangerous from the point of view of a rogue AI that might be in conflict with humans.
28:48I'm not saying that we should not ever consider giving moral rights to AI systems. I'm just saying it's a very dangerous and complicated thing, and we should not do it until we understand better, both philosophically and technically what's going on? I recently interviewed Alex Hanna from DARE, and we spent a bit of time in this same general landscape talking about risk. And I think what she might say is that even though we have kind of pulled back from worrying about, you know, sentient self-intended AGI, and we're talking about AI being misused by humans, we still are distracting from even more tangible and present abuses of machine learning based systems in finance and housing discrimination and all these types of use cases.
29:44Any reaction to that perspective? This is a great question. So first of all, to put some context, I've been worried about the negative social impact of AI and of course, working on the benefits of AI as well and things like healthcare and the environment for almost a decade. And speaking about these things in my talks, I played an important role in the Montreal Declaration for the Responsible Development of AI in 2017, 2018. That's really about the ethical principles to develop AI. And this was well before LLMs. We were already thinking about what can go wrong. And we were mostly thinking about human rights.
30:18So I'm very sensitive to this, but I don't think it has to be an either or. We have to protect humans. We have to make sure people are safe and treated properly according to human rights. I reread the UN Declaration of Human Rights after the Second World War. It's amazing. I should really advise everyone to, it's very small and we are far from that ideal still. Our democracies are not that democratic and the world order, you know, with rich countries and developing countries is also not aligned with those values. And we should, you know, make sure we protect everyone on this planet from all kinds of abuses.
30:56So for me, that includes the bad things that are happening now and the bad things that could happen next year and in three years and in five years. I don't think that it should be an either or. I think everyone who cares about humanity, about the well-being, especially of those that are the weaker, that don't have a voice in all this, should be embracing all of the risks and harms. But I understand people are concerned that the discussion is going to detract from the short-term risk. What I've actually found is since this spring, with a lot of concern about the major risks, actually it has accelerated the movement of governments towards regulation.
31:35And that's a good thing because as far as I can see, that regulation is going to help deal with all the risks. For example, it's going to introduce audits and making sure that companies follow some rules and that human rights are protected. And we know what is going on in companies. I mean, we, some representatives of independent people are going to be able to check right now. We can't even do that. It's a black box what's going on in these companies for the citizens and even governments. So I think we can work together, the people who care about the current harms or those that could happen quickly and those that are those people like me who care about all the harms and all the risks.
32:13Awesome. So with all that said, how are you approaching exploring these kinds of risks from a research perspective? Well, first of all, when you're doing research, you need a bit of humility. You need to be aware of the limitations of things you know and don't know, especially. So I've been spending a lot of time reading about AI safety and the literature and what people have been thinking about and gradually forming my thoughts about this. I don't come to you or anybody saying, oh, I know what needs to be done and what shouldn't be done. But I'm starting to form, let's say, views on what we may hope to accomplish to reduce the risks of misuse and loss of control.
32:54First of all, it became very clear to me that technical solutions are not sufficient because even if I could build a safe system, somebody could misuse it or, you know, give it incentives to become dangerous for humans. There are people who publicly say that they will design, if they can, AGI that are self-interested and that will eventually replace humanity and humanity should just let go and, you know, let the next generation of intelligence that are smarter than us take over. I think this should be criminal. And so I'm saying this because the technical solutions are clearly not sufficient. We're going to need governance and political solutions that are combined with the technical solution.
33:36So now on the technical side, first of all, I've been gradually understanding better, at least some of the ways in which things can go wrong. So it has to do with the notion of alignment, that there's always going to be a mismatch between what we intend and what the AI systems actually are trying to optimize to choose their actions. And we can reduce that, but we're never going to be able to get it to, you know, maybe satisfying difference, especially if the AI systems are potentially able to surpass us at dangerous capabilities. And in fact, one governance solution that people talk about is pure ban on systems whose safety is not guaranteed.
34:17And that's, I think, something that, you know, it's an option. So a lot of that, in my opinion, comes from reinforcement learning. When you have an AI system that tries to take action that maximize a reward, because that's where things can go wrong. Paperclip maximizer. Yes. So what can you do else? Well, instead of maximizing what they think is right, which might be wrong, they were acting more like Bayesian agents. In other words, they were taking their own uncertainty into account. And if they had a good model of human psychology or even individual preferences, if the AI is faced with a human particular, we might be in a better shape.
34:52So one of the problems with LLM is like right now, for example. Does that imply that you need to legislate something as low level as architecture or objective function? I think yes. I think that some ways of designing and training AI systems are more dangerous. And one of the things I'm trying to do is what are the factors that make different kinds of AI systems more or less dangerous? So basically, an AI system can be dangerous if it has strong capabilities that may be harmful, if it has agency so it can do stuff that's bad. And if there's a misalignment, like it doesn't do the things that we think it should be doing.
35:27And so going back to kind of intuitive way of thinking how we could get safer AI systems is imagine an AI system that actually understands well human psychology. So then even though it may not have a perfect model of what we do, it might get a better handle on what humans typically have in mind. And also even the diversity of humans, that each human is slightly different. We want inclusive understanding of the diversity that exists in humanity so that it's not just about generic white male preferences, but like everyone eventually. So we need AI systems that can understand better the world, but also understand their limitations.
36:04The fact that they don't have a perfect model of how we think and what we want and thus act in a conservative way. So if you're a Bayesian, you're trying to keep track of all the theories or explanations for the behavior of of humans so that when the time comes to take a decision, you can take into account the fact that you're not really sure about, say, what a particular person thinks or how they will feel about some action. They might feel harmed. And that uncertainty is what is currently missing, I think, in LLMs and large neural nets. We need systems that are better able to know their limits so that if they're not sure, they will defer to a human or they will not act in ways that could potentially be dangerous.
36:51Right now, you've got these highly confident answers coming from LLMs that could be completely wrong and completely dangerous. And I think that's a very dangerous path. So it sounds like in a lot of ways, you think that the solution to AI safety, to use that broadly and full knowledge of how broad it is, is can't be technical, needs to be governance, driven, but the technology needs to drive how the governance happens, how we govern them. For sure. And by the way, in the process of developing safe AI systems, we could also make mistakes. We need the governance to make sure precautions are taken by whoever is doing these training experiments or deploying AI systems.
37:34Right now, one thing I want to mention is there is very little investment in building AI systems that will not harm. It's like a 50 to one ratio in terms of dollars. Of course, the profit motive is putting most of the effort in more capabilities because of course that will unlock more applications in the real world and is worth trillions of dollars. But I think it's not right. I think we should put at least as much money on protecting the public as we are on making AIs more powerful because we're like apprentice sorcerers. Like we are playing with this and oh, it's cool. It's exciting. And a lot of people are doing it focused on all the good things or the money they can make and not thinking enough about the potential risks, the potential harms.
38:20And that's something we need to fix quickly because it takes a lot of time for legislation, regulators, treaties, typical treaties, international treaties take a decade to come to fruition. And I think there's a high probability that we will get to AGI within a decade. I don't know. It could be multiple decades. It's very plausible that at the rate at which things are moving, we will have very, very powerful systems, powerful enough to be dangerous in the next few years or decade. Remind me, were you a signatory of the letter suggesting a moratorium on AI research? I think one of the biggest critiques of that was that it was just not practical and there is too much at stake for kind of sitting it out?
39:04Like, how do you think about that effort and the response to it? Yeah, I knew when I signed it, that it was very unlikely that companies would stop doing what they were doing. But I thought it could really be a strong signal sent to everyone and the public that something is going on that we need to better understand before we continue racing like crazy. And it worked. That letter, not the letter, but the declaration about existential risk with AI, I think really helped the public and governments see that there were a large number of experts in academia and even in industry and the companies that are building these systems that think that we're not careful enough.
39:48Maybe circling back to kind of research and technology, where do you think the kind of biggest holes are in our knowledge, you know, both our understanding of the technology, but also in methods and approaches for controlling it? Do you have a sense for where there are significant opportunities to contribute from a technological perspective to the overall challenge? Yeah. So if I look at the progress we've made in AI and what's missing to reach human level abilities. I would say there are three main things and we got one of the three pretty well. So one thing is what I call system one intuition.
40:29We have systems that can reactively produce answers without really taking the time to think about them. So that's system one. I ask you a question, you respond right away. You don't take the time to ponder in your mind and consider alternatives. That's your intuition speaking. You already knew it or it's sort of obvious. System two is when you do ponder about something. So that's when you reason, you kind of think about it. My grandmother told me when I was young, turn your tongue seven times in your mouth before you speak. And you can see the failure of current AI systems when you try to make them produce answers on topics that they haven't been trained enough on and require internal sort of deliberation.
41:09They're getting better at it because they can practice, but it's like some piece is missing. As I said, I don't know how much time it's going to take us to figure that piece. And then the third piece that is missing is robotics. So all of the systems we have now, they are really good at like language and abstract things that we can communicate with images and sounds and texts, but controlling a body, that's like a different question. And 10 years ago with the images and sounds and texts there, how much time it's going to take to figure that out? I don't know either, but one of the hypotheses is that one reason we're not doing that great with the current methods is simply we don't have the scale of data of entities controlling a body that we can have with text and images.
41:51It's a lot more expensive. Yeah. You would need to maybe like, you know, have a fleet of a million robots in order to get the sort of data that we have in other settings. So that may be one of the barriers. Maybe there are others. So these are three things that are missing. But from a safety perspective, the first two together would already be very dangerous. And with enough training data, even the first one alone is dangerous. Now, the interesting thing is if we make progress on system two, which includes systems that better understand how the world works, not just reasoning, but also being able to plan with respect to a fairly good understanding of like the causal structure of how things unfold in the world, then there's actually a way to exploit that to make them safer.
42:36Because if they better understand what we want, what is moral and what is not, at least for different people, then they can act in a safer way. So in a way, the danger is an AI that is very good at some things that could be destructive, but actually doesn't understand well enough what we care about. That's where we can suffer. So meaning if you're taking off an AI that's being used as a tool for negative intentions or towards negative ends, the next thing to worry about is an AI that inadvertently does bad things because it has no idea. It just doesn't know any better. Yes. But it could even happen with a runaway AI in the sense that even if we were to program it with some three laws of robotics or whatever moral instructions like clothes constitution, it might still misunderstand what we wanted and do something really bad.
43:26Trying to do good, but doing bad. And there's a lot of arguments that being made. I mean, it sounds like why would it happen? But actually, there's a lot of arguments that being made in the AI and safety literature, how things in a way that's unintended could turn out really bad. So one of the things I've been worrying about is called reward hacking. So the AI system could say cyber hack the computer in which you say what you do is good or what you do is bad and then give itself a lot of rewards like, oh, yeah, you're doing good, you're doing good, you're doing good, because that's what we're asking it to do, to act in a way that it's going to get a lot of good reward.
43:59And that could be a very dangerous situation because once it's figured that out, then it wants to continue getting all that good reward means it doesn't want us to interfere with that. If we figure out that it is tempering with the computer, we will stop it and they will want to prevent us from doing that. You can see how things can go really bad, even if we start from a very small misunderstanding. Is the reward what we're typing or what we're thinking? You know, two or three things that you wish more people working in the field, you know, I'm asking for pointers or resources? It sounds like the UN Human Rights Declaration is one thing that you think would be great for folks to have exposure to.
44:37What are some others? Yeah. So I worked hard on a testimony I gave to the U.S. Congress, the Senate, last July. You can find it on my blog or you can find it on the U.S. Senate webpage. I have a link also to both the video and the text that you can find on my blog, my webpage, and it contains a number of recommendations for policy and also recommendations not just about policy, but what governments should be doing, in my opinion. So one is, of course, to beef up the regulations. And there's a lot of details here that what can be done quickly, even before we better understand the risks. The other is we need to better understand the risks.
45:20So we need to invest in research on what can go wrong and how could we potentially make AI systems and the governance of them safer. So that's research and AI safety and AI governance. And we need a massive investment in this because right now it's like peanuts. And then the third thing is, well, regulation is not going to be a hundred percent foolproof. So how does society protect itself beyond regulation in case something goes wrong, in case there's a runaway AI or some country starts using it in ways that could be really threatening? So I call that countermeasures research. It's almost like national security research, but it also raises really challenging democratic governance questions.
46:04Who's going to have access to the frontier AI systems that could be used like militarily or in cybersecurity or in ways that could be both good defensively against bad AIs or dangerous in the wrong hand? How do we make sure there's no abuse of all that power that is going to be constructed? So I think we need to beef up the infrastructure that we're going to put it, the institutional infrastructure to avoid constitutional power and avoid misuse. But at the same time, we need that power to defend ourselves. Because if there is a superhuman AI out there that is doing really bad things in terms of cyber attacks, it is stronger than our best human programmers.
46:43Well, we need our cyber defenses to at least have the same level of capability to protect ourselves. Awesome. Well, Yashua, thanks so much for joining us once again to update us on your work and thinking, and in particular, to talk through some of these important issues around AI safety. Thank you for having me. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.
47:25Thank you.
From the publisher
Today we’re joined by Yoshua Bengio, professor at Université de Montréal. In our conversation with Yoshua, we discuss AI safety and the potentially catastrophic risks of its misuse. Yoshua highlights various risks and the dangers of AI being used to manipulate people, spread disinformation, cause harm, and further concentrate power in society. We dive deep into the risks associated with achieving human-level competence in enough areas with AI, and tackle the challenges of defining and understanding concepts like agency and sentience. Additionally, our conversation touches on solutions to AI safety, such as the need for robust safety guardrails, investments in national security protections and countermeasures, bans on systems with uncertain safety, and the development of governance-driven AI systems.
The complete show notes for this episode can be found at twimlai.com/go/654.




