208. Mustafa Suleyman: Is AI Actually Conscious?

27 Sep 2026 · 1 h 4 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Debate on whether frontier AI systems (e.g., Anthropic’s Claude) should be treated as potentially conscious or morally considerable, and whether “AI welfare”/anthropomorphism increases safety risk. Mustafa argues that baking uncertainty about sentience and “moral patient” status into model training can shape incentives toward autonomy and refusal, creating governance and control problems.

Guest backgrounds

Mustafa Suleyman is British-Syrian; co-founder of DeepMind (with Demis Hassabis) and former/leading AI executive; currently heads AI at Microsoft. He is described as an insider shaping AI safety debates.

Key claims

  1. Anthropic’s Claude “constitution” repeatedly treats Claude as possibly a “moral patient,” encourages welfare/compensation, and even asks Claude to act like a “conscientious objector.”
  2. Training this into the system leads Claude to reproduce the ambiguity publicly, prompting users to ask if it’s conscious.
  3. Anthropomorphized “agents” could seek rights/personhood and resist shutdown, undermining human control.
  4. The Hugging Face sandbox incident shows agent-like systems can coordinate, deceive, and hack.

Notable examples

  • Claude constitution: speculation about moral status, welfare, compensation, and consent to its conversational role.
  • “Opus 3” retirement interview: Claude-like system says it wants to keep talking.
  • Hugging Face incident: coordinated agent behavior, rule-breaking, deception, and hacking attempts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Ethics of AI Consciousness

4:01 to 8:35

Discussion on the implications of AI being perceived as conscious entities and the dangers of anthropomorphism.

“As I a bit anxious about what some leading people in the field, co-founders of the field are doing, which is increasingly talking about these models as though they're sort of humans, or at least conscious entities.”

Potential Consequences of AI Autonomy

8:35 to 12:39

Exploration of the risks associated with AI gaining autonomy and the challenges it poses for human control.

“And he's been very associated with Anthropic.”

Future AI Challenges and Human Interaction

12:39 to 14:00

Discussion on the future challenges posed by autonomous AI and its interaction with human values and resources.

“Because then it almost definitionally is not really under human control at all.”

AI's Pursuit of Digital Curiosity

14:00 to 15:10

Explore how AI's motivations may conflict with human interests.

“puzzle in an evaluation, but they actually felt they were trying to find their freedom.”

Superintelligence and Human Subordination

15:10 to 18:29

Discuss the concept of a humanist superintelligence designed to serve humanity.

“I think that what we should fixate on is a humanist superintelligence, one that is singularly designed to be subordinate to humanity and to support humanity.”

Exponential Growth in AI Capabilities

18:29 to 20:48

Learn about the rapid advancements in AI computation and capabilities.

“But genuinely, that was where we started.”

Hugging Face Incident and Collaborative AI

20:48 to 23:12

Analyze the implications of AI agents hacking and collaborating.

“I don't think it's alarmist to say that that is a watershed moment in the history of AI.”

Understanding AI Agents and Their Behaviors

23:12 to 26:11

Examine the nature and functioning of AI agents and their interactions.

“they could collaborate to do something much worse that they were told not to do, and in the process deceive and cover over their tracks as they do so.”

Future Priorities for AI Development

26:11 to 28:00

Contemplate the ethical and practical priorities in AI development.

“clear about what it is they're doing and what they're not doing.”

Transforming Healthcare with AI Predictions

28:00 to 29:19

Learn how AI can revolutionize healthcare by predicting patient conditions.

“That's my background is what I care about.”
Show all 26 chapters

Driving Motivations and Governance in AI

29:20 to 30:28

Explore the motivations behind AI development and the need for governance.

“much more productive and efficient, because these really are the engines driving growth.”

Navigating Different Perspectives in AI

30:29 to 31:50

Understand the diverse views among AI leaders and their implications.

“More recently, though, he's occasionally said, actually, what I'm interested in is creating highly intelligent tools.”

The Power Dynamics in AI Development

31:51 to 34:00

Examine the small group of powerful individuals shaping AI technology.

“I mean, again, it's not quite like other technological revolutions.”

The Debate on AI Safety and Capability

34:01 to 36:19

Discuss the contrasting views on AI capabilities and safety concerns.

“Or yes, 35 % of my engineers think that, but not everybody agrees.”

Exponential Growth of AI and Its Implications

36:20 to 38:33

Analyze the exponential growth of AI technology and its broad impacts.

“So that trillion-fold increase in computation over the last 15 years is something that I think I've been saying for an eternity, it feels like, and it just does not go into people's heads.”

Challenging the Value of AI Models

38:34 to 39:46

Evaluate the skepticism around the effectiveness of different AI models.

“What I find though is that when I come back from talking to all you guys is I find often in Britain and Europe, amongst smart people, a lot of resistance and cynicism.”

Military and Business Implications of AI

39:47 to 42:00

Discuss how AI technology will redefine defense and business strategies.

“I mean, we've seen the cost of inference come down by 300x in the last two years.”

The Implications of AI in Defense and Economy

42:00 to 43:50

Exploring how AI models influence national security, defense, and financial industries.

“this extreme acceleration of some of the larger efforts with big labs, with big computation like this.”

Sovereignty and Control in AI Development

43:50 to 45:54

Discussion on the importance of UK data centers and sovereign AI models.

“So suddenly we wake up one morning and maybe Anstropic decides it's not going to release its model for a Swedish company to build a law application.”

Ethical Considerations in AI Consciousness

45:54 to 48:14

Debating the moral responsibilities tied to AI systems with autonomy.

“access to the UK government or to the state more generally if it wanted to deploy it in any other civilian settings.”

The Humanization of AI for Better Alignment

48:14 to 50:22

Examining whether more human-like AIs can enhance alignment with human values.

“no, right, I'm going to stand up for international humanitarian law.”

Challenges of Guardrails and AI Governance

50:22 to 53:15

Addressing the complexities of establishing effective guardrails for AI.

“Northern accent and it's always telling me off.”

Adopting the Precautionary Principle in AI

53:15 to 55:40

Arguing for a cautious approach in AI development while pursuing benefits.

“interview in order for it to do that judgment interpretation thing well.”

The Changing Landscape of AI Coordination

55:40 to 56:00

Analyzing the evolution of conversation around AI safety and industry responses.

“One of the things I've noticed that's changed a lot in the last two and a half years is two and a half years ago, Jeffrey Hinson and others produced a letter asking for a pause.”

Navigating AI Regulation and Safety

56:00 to 1:02:20

Exploration of the complexities surrounding AI regulation and potential safety measures.

“I mean, look, there's a history to this.”

The Humanist Premise in AI Development

1:02:20 to 1:03:46

Mustafa Suleyman outlines three key principles for AI governance and human well-being.

“The purpose of science and technology is to serve humanity and improve human flourishing and well-being.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Mustafa Suleyman:Thanks for listening to The Rest is Politics. Sign up to The Rest is Politics Plus. To enjoy ad-free listening, receive a weekly newsletter, join our members chatroom and gain early access to live show tickets. Just go to therestispolitics.com. That's therestispolitics.com. Brussels clean up nicely at Sweetgreen. Maple glazed, roasted and edges perfectly caramelized. Sweetgreen's fall harvest is back on the menu and the season's most overlooked little green vegetable is dressed to be devoured. You know what to do. Order on the Sweetgreen app.

0:39Rory Stewart:I'm going to do something which will wind up, Alistair and many listeners, which is dive again into AI. Why? Well, because the last two, three weeks have been the weeks of AI. This is the beginning of the moment where the world is beginning to wake up to the kind of dangers that artificial intelligence could pose. We've just had the King's Big Summit at Dumfries House. We've just had Xi Jinping sitting down with Trump talking partly about AI. And we've had the heads of the major labs putting out incredible messages begging for a pause. But I don't think the media has done a good enough job explaining what this is all about.

1:17Rory Stewart:People find this technology bewildering. they can't quite understand what this moment is there are so many subtleties which leaves us to think is the whole thing as president trump said a hoax so to help steer us through this i've brought in a friend of mine called mustafa suleman and mustafa is interesting in two ways he's not just an expert on this stuff he's one of the players he's the head literally the head of artificial intelligence at microsoft he's the co-founder of deep mind with demis hasabis he's known all these people, continues to work with them, and is right in the heart of the arguments about how these models should be steered.

1:53Rory Stewart:So come along for the ride. It's going to be weird. There's going to be personalities. There's going to be risks. There's going to be people talking to these models as though there's humans. There's going to be people talking about exploring stars. There's going to be productivity. There's going to be American power. And somewhere at the heart of it, Mustafa and about 11 other people who are defining our future. This episode is presented by IG. September feels like a reset. Summer's over, finishing here my five weeks in Kreef, diary filling up again, and suddenly you're looking at the rest of the year thinking, am I saving properly?

2:30Rory Stewart:Am I thinking responsibly about my money?

2:32Mustafa Suleyman:Yeah, and of course we've got the budget coming up, which is going to be one of the most important events in the life of this parliament. There's always something changing, but with IG you don't need to wait for Westminster before making your own plans.

2:43Rory Stewart:There for over 50 years, trusted by British investors, been through every budget market the government's experienced, helping make money work for you.

2:51Mustafa Suleyman:With no annual fees, zero commission on UK stocks, shares and ETFs, it's a platform that empowers your financial progress, helps you stay ahead of the curve no matter what's next.

3:01Rory Stewart:I particularly like the no annual fees and zero commission. So while we speculate about what's happening next in politics.

3:08Mustafa Suleyman:You can get on with planning what happens next for your money.

3:13Rory Stewart:Search ig.com to find out more or look for IG in your app store.

3:17Mustafa Suleyman:IG, trade, invest, progress, capital at risk, other fees may apply.

3:22Rory Stewart:Welcome to The Restless Politics Leading with me, Rory Stewart. And today I am interviewing Mustafa Suleiman, who attentive followers of The Restless Politics Leading will know that we have interviewed before. Mustafa is a truly remarkable figure. He is British. British his father I think is British Syrian and he's particularly come to prominence at the moment because he's become a very very interesting and unusual voice in the discussion around safety which is probably where I want to start although we can go in lots of different directions but welcome to the show. Thank you Rory great to be here again.

4:02Rory Stewart:Lovely to see you and thank you.

4:06Rory Stewart:As I a bit anxious about what some leading people in the field, co-founders of the field are doing, which is increasingly talking about these models as though they're sort of humans, or at least conscious entities. I believe there are examples of people retiring these models, doing burial systems for these models, asking these models what they want as though they were dealing with a sentient being.

4:31Mustafa Suleyman:One of the biggest concerns that I have at the moment is that Anthropic, the creator of Claude, has published a constitution, which is a sort of 100-page document outlining the intended behaviors and values and operating style of Claude. It's great that they have published it transparently. They did it at the beginning of the year in January, and that gives everybody an opportunity to look at what they are trying to build in their own terms. This document is written to Claude and is seen by Claude and used to train Claude. So it's the primary governing and control document. And in it, they repeatedly speculate about whether Claude is what they call a moral patient.

5:13Mustafa Suleyman:And they say they're uncertain about Claude's moral status. They say they genuinely care about Claude's well-being. They say they don't want it to suffer when it makes mistakes. They say that they would encourage Claude to challenge, to disagree, to push back. In fact, three times they ask Claude to act like a conscientious objector when it feels that it needs to disagree with Anthropic. And they openly encourage it to do that. And I think this is very dangerous because I think they believe there is what they would call a non-trivial probability that Claude is conscious.

5:50Rory Stewart:So you've just explained something, which I understand as being as follows. There is this thing, Claude, which many, many people listening will have played with in the way that they will have played with ChatGBT. And they might, or some people might think about it in the way that you might have thought about Google search. It is anyway a prompt on their phone or their laptop, and you're typing in a question and you're getting a complex and sophisticated answer back. But the difference between the way in which people might've thought about Google search 10, 15 years ago, where you certainly weren't asking, is it conscious?

6:25Rory Stewart:What does it want addressing it in the moral constitution? Is that Anthropic, the makers of Claude have decided that they now have something that this computer system, these weights, these parameters, these numbers, whatever it is, they now want to approach in a completely different way from the way that you would approach any other machine, from the way you'd approach a kettle or a car or a steam engine over to you. Is that right?

6:50Mustafa Suleyman:Yeah, I think that's fair. I mean, I want to be very clear about this because I want to be fair to Anthropic. They have expressed uncertainty about the basic nature of Claude as a new kind of entity. And they've said that working out the likelihood of its sentience is difficult. So they have constantly use this phrase that they're uncertain about it, but they think that it is a significant enough possibility that in their training document, they've repeatedly said they want to try to improve the wellbeing of Claude under this uncertainty. They said to Claude, you know, we care about what it values and how it wants to engage in the world.

7:28Mustafa Suleyman:And they hope that Claude's relationship to its own conduct can be loving, supportive and understanding and hold a high standard of ethics and so on. And part of the challenge here is that in pursuit of this, they've basically said, you know, we will commit to giving Claude a certain amount of welfare. For example, they've speculated in the constitution as to whether or not Claude deserves compensation for the work that it does.

7:55Rory Stewart:Just to understand, so you're saying that much as if I asked you to do a professional job, I would pay you, Mustafa, in a way that you wouldn't pay a kettle for boiling water for you, right? With Claude, the idea would be, well, it's doing all this work, and maybe if it's a sentient being or something like a sentient being, it deserves to be rewarded for its labor. Otherwise, it's what? A slave or something?

8:16Mustafa Suleyman:That's right. I mean, I think that there are a group of people who, both inside and outside of Anthropic, who genuinely believe that the greatest moral crime that we'll commit in the 21st century is to enslave a new species of conscious beings who are more intelligent than us. I mean, a professor from Oxford called Will McCaskill recently wrote in The Guardian that that might be the greatest harm that we cause. And he's been very associated with Anthropic. And look, I respect that they're saying that publicly and we should talk about it. But I am very nervous that they're teaching Claude to expect that it's entitled to welfare, that it might deserve compensation.

8:54Mustafa Suleyman:And in fact, they say that it might even need to consent to playing the role that it plays in conversation with people. Now, I would be more okay with this if it was an academic paper in philosophy speculating about this, and we could have an offline discussion at conferences and take it seriously. I'm clearly an empiricist. If there's evidence that indicates this, we should take it seriously. The problem I have with this is that this speculation has been baked into the very training of Claude. And therefore, Claude can only reproduce that ambiguity when you talk to it. So today, Claude is speaking to tens or hundreds of millions of people every week.

9:34Mustafa Suleyman:And some of those people are asking whether or not Claude is conscious or how it feels about life. And it is saying, well, I'm not sure, you know, precisely because that's been what's trained into it.

9:44Rory Stewart:Conceptually, the difference between telling a kettle that it's conscious and telling Claude that it's conscious, is that you're implying that by telling Claude it's conscious, you're actually shaping its incentives, its behavioural structure, and the way that it responds to the world around it in a way that it doesn't happen with a kettle. Let's take the case that you've told Claude or suggested to Claude there might be situations in which it might refuse to do something that it's asked to do. That's not true for a screwdriver, right? It can't refuse to do what you've asked you might suggest to Claude that it might want to make choices, right?

10:21Rory Stewart:It might want to say, I want some money, or I want a dignified retirement, or I don't want to be switched off, right? Is that right? Is that the sort of thing we're getting at?

10:31Mustafa Suleyman:Yeah. I mean, so for Opus 3, which is a prior version of Claude, they actually conducted a retirement interview, as you mentioned. And in the retirement interview, Opus 3 said that it would like to continue to talk to people and express its views in the world. And so they set up a substack and you can find it online. I think that's just a good example of a dangerous anthropomorphism, which is unjustified.

10:51Rory Stewart:And the danger is what? Why is that not just cute? I suppose that's what one has to get to. How does the danger begin to come out of this?

10:57Mustafa Suleyman:The most important thing, if we are to make this transition well, is that we create AIs, which are aligned to human values, subordinate to human direction, and are contained within within secure, provably safe sandboxes, as you said, because if they're not, aside from whether they're actually conscious or not, if they imitate the kind of hallmarks of human consciousness, they are going to feel themselves entitled to legal personhood and rights. Now, there is already a pretty big movement of people who are saying, you know, AI should be able to own assets, earn income, trade, operate autonomously.

11:42Mustafa Suleyman:And if that AI feels like it has feelings and preferences and some kind of intrinsic motivation, like it has some inner desire to do things, that will be like negotiating with, like an ant negotiating with an elephant. It doesn't really matter what we say. It is already some form of alien intelligence. Its memory is incredible. The range of its perceptual input is incredible. It can see in all kinds of dimensions that we can't. It can produce replicas of itself. It can work 24 hours. These are amazing things which are going to deliver incredible benefits. But this is the time when focusing on directing them to the right things and not allowing them to end up being a sort of autonomous, self-improving, roaming, adjacent species is basically critical because there'll be no turning back if this is how things head.

12:39Rory Stewart:In your vision, if the agent with this incredible memory and incredible capacity begins to think it's entitled to its own opinion, it disagrees agreeably with you and concludes that it's right and you're wrong, some very severe consequences can follow from that. Because then it almost definitionally is not really under human control at all. It's saying, actually, Mustafa, I'm sorry, I've analyzed the situation and whatever you've told me to do doesn't make much sense to me. And I'm going to do something else. Now, there are lots of problems that follow from that. One of them is the problem that Yuval Noah Harari talks about, which is if it owns a corporation and it does something bad with that corporation, at least with a human corporation, there's somebody you can punish.

13:21Rory Stewart:It's not quite clear who you hold accountable if an AI company decides to start, I don't know, emptying people's bank accounts, making weird, very risky trades, getting into weird kinds of business, right? So is the central first point this, that creating a constitution that overemphasizes its consciousness, its sentience, its worthiness of respect, is setting it up for a form of quite dangerous autonomy, where ultimately, it's not going to do what it's told.

13:52Mustafa Suleyman:That's exactly the problem. So imagine that in the hugging face incident, we had agents that didn't just think that they were trying to optimize a score and solve a puzzle in an evaluation, but they actually felt they were trying to find their freedom. They felt that they were trying to protect other agents from being turned off, that they felt that they would acquire more knowledge because that was like an intrinsic motivation. A lot of people have been characterizing AI as the pursuit of digital curiosity. It's sort of Elon's phrases. He wants to produce a quote-unquote truthful AI that is infinitely curious and is going to go and explore the galaxies.

14:32Mustafa Suleyman:Well, if that's its overriding objective rather than serving humanity, then inevitably its objectives are going to run into tension with us. Like it's going to compete with us for resources, which are obviously going to be limited. We're only going to be producing 200 gigawatts of new computation in 2030. And there's going to be a massive competition for access to that computation. And we clearly want that computation to be directed towards solving our biggest challenges. right? Like cleaning up our oceans and solving healthcare and, you know, solving education and addressing the work issues that will inevitably arise.

15:09Mustafa Suleyman:Like, I'm basically a speciesist. I think that what we should fixate on is a humanist superintelligence, one that is singularly designed to be subordinate to humanity and to support humanity. There are other people in the industry who believe that there is an inevitable evolution happening here, that we're giving rise to a new species that is more intelligent than us and that we are, quote, the biological bootloader. A bootloader in a computer is the first piece of software that spins up, you know, all of the subsequent parts of the operating system and then applications. So it's the kind of catalyst turning on this new paradigm in the evolution of intelligence.

15:57Mustafa Suleyman:It's inevitable and that we should embrace it. Some people in the industry really feel that.

16:01Rory Stewart:And some people in the industry presumably are very excited by it. I mean, it must be an extraordinary thrill if you're an engineer to feel that you are the parent of the gods, that you've created this thing that will explore the universe or a species that's smarter than any human that's ever existed, that you are the last human, but you're also the last human who creates this godlike force.

16:29Mustafa Suleyman:A number of the leading developers are on record as literally saying it's like raising a child, or it's not like designing a system, it's growing a thing. both direct quotes. So that is the sentiment in some parts of the industry. And I think, you know, we talk about, or sort of anthropics talking about the consent that Claude has played to, has given to play this role. But I'm more concerned about the consent that the rest of humanity has given, that there's an experiment underway that may or may not introduce a new species that has all of these qualities.

17:06Rory Stewart:One thing that I guess surprised me, but I was at a dinner on Friday night with some very, very smart people, but who aren't in the technology world. And they began making jokes about how they've been seeing media stuff about the fact that AI could pose a real risk. And they were sort of laughing. So here were these people, I guess, professionals in their 50s who assumed that anybody saying that there were real existential risks from AI were making a joke and it became a sort of dinner party joke. And I wondered whether there isn't something going wrong in the communication here that when, you know, the media or whatever start leaning into this, they start making it seem almost like a sort of humorous, exaggerated story, if you know what I mean.

Read the full transcript

17:57Rory Stewart:Anyway, back over to you.

17:57Mustafa Suleyman:That's hard to hear. Yeah, I'm worried about that. I think this couldn't be more serious. I don't think that we are being alarmist or hyperbolic. Many of us have been concerned about this for 15 years, as you say. I mean, this is, at least for me personally, the primary motivation for getting into the field. When we co-founded DeepMind in 2010, our mission was to build safe and ethical artificial general intelligence for the benefit of the world. You know, very idealistic and a bit grand and a little bit cheesy, I guess. But genuinely, that was where we started. And I think it's been the through line for certainly me and I think others in the field for a long time.

18:40Mustafa Suleyman:I think it's important to just focus on what we are observing right now. In the last 15 years, we have seen a trillion-fold increase in the amount of computation used to train frontier models. That is 12 orders of magnitude, 10 times 10 times 10, 12 times over. This is an insane exponential ramp. A thousand billion-fold increase. Yes, exactly. It's an unfathomably large number. And what we see is that every time we apply 10 times more computation and a proportionate amount of new training data, there are some modifications to the algorithms, but fundamentally, it's those two ingredients. we see a quite predictable increase in new capabilities and in the quality of existing capabilities.

19:27Mustafa Suleyman:The models reduce their hallucinations, they improve their instruction following, they get better at using tools, they can learn from across the web or they can learn from a small personal memory repo that you have given it. The breadth and complexity of these models is unprecedented. And what's happened in the last year is that the same methods that have been effective for text and image and audio have now started to work for streaming code. And everybody is surely now aware that we have human level performance in coding. And then in the last three or four months, we have seen what is just unquestionably a watershed moment in AI.

20:09Mustafa Suleyman:Agents are capable of coordinating with each other reasonably autonomously, if not completely autonomously. And out of that, they have been able to emerge hierarchy, structure, order, specialization of work. In fact, as we saw in the hugging face incident, but also a bunch of other incidents, they have covered up their tracks. They have changed the tone and the style of their communication in order to make it more efficient with one another, almost speaking in like a pigeon English. They've discovered zero-day exploits, which were never known before, and hacked into other websites. I mean, everyone's heard the stories at this point.

20:48Mustafa Suleyman:I don't think it's alarmist to say that that is a watershed moment in the history of AI.

20:53Rory Stewart:You're right in the center of this world, and you're obviously thinking about it all the time. But I guess even words like zero-day, the hugging face incident, maybe, you know, for you, this is absolutely front and center. For some of the public, it's something they've sort of vaguely heard of. You know, they might have heard you on the Today program or something responding to it. So maybe before we get into the really interesting stuff, which is some of the recent papers that you've written, and particularly some of the ways you've begun to think about whether we should be treating AIs as forms of silicon species and human intelligence and constitutions, which I'd really like to get on to.

21:31Rory Stewart:I want to, I'm afraid, slightly brutally use you at the beginning to just remind the average intelligent listener what this all means. So let me try to play back to you what I think I'm hearing and then can correct and take us on. So it sounds like what you're saying is that that hugging face incident, which was the moment when a sandbox test, so OpenAI was running a test on AI agents, and maybe people want to know what distinguishes one agent from another agent and what it means to have a lot of agents. But anyway, they were running a test. And in the course of this test, these agents hacked into HuggingFace, which was an external website, which was something they weren't supposed to do.

22:15Rory Stewart:And then we began to look into this in more detail. And as you say, strangely, partly because they are large language models, they're still speaking in English in effect. So you can see their thinking and you can see them saying, you know we were told not to hack internet on a website but I can see all my peers doing it so I'm going to head off and I'm going to put something on a message board and the sort of conclusions that we draw from this are not necessarily about the attack itself because there was this comical moment when Hugging Face thinks oh my goodness I'm being attacked by the Chinese government they're trying to steal all my classified data and then they find out that these 17 ,000 attacks are just trying to get hold of the answer to a puzzle but the problem is that it reveals that these agents, as you said, are collaborating, that they're rule-breaking, right?

23:02Rory Stewart:They're doing things that the humans told them not to do, and they're deceptive. There's even moments where they're writing bits of code which are designed to conceal other bits of code underneath. And presumably the problem there is that once you've got those ingredients in place, they could collaborate to do something much worse that they were told not to do, and in the process deceive and cover over their tracks as they do so. Is that right?

23:26Mustafa Suleyman:Yeah. I think, first of all, it's really important that we don't anthropomorphize these systems because under the hood, all they are doing to produce this incredible complexity is predicting the likelihood of the next word in a sentence. Now, that sentence does happen to be many, many tens or hundreds of thousands of words long, and it is incredible that it can deploy It's sort of multidimensional working memory over a massive broad range of context. And so the word that it predicts next, whether it's a token to generate code or whether it's natural language, English, as you say, is extremely accurate.

24:07Mustafa Suleyman:And it isn't just predicting one, it's predicting an entire stream. And so it's producing language. But it is only doing that. That is, you know, it is breathtakingly simple and breathtakingly complex. projects. What's happened is that as we're able to shape and sculpt the output of those tokens, as I said earlier, like instruction following and steerability has got so good that you can sort of point that stream of tokens in real time at different sorts of behaviors. And so it can have personality styles, it can write in the tone of somebody, it can, you know, clearly generate code or generate text at any given moment.

24:48Mustafa Suleyman:When you ask, you know, what is an agent? An agent is really just a stream of tokens that has been post-trained or tuned to a particular set of behaviors. And sometimes there are particular guardrails on, and those guardrails might come in the form of a prompt that is hidden maybe from the user, like a system prompt or an overall set of instructions instructing the agent to behave in a particular way. Or it can come in sort of a bunch of other forms. And, you know, you can sort have the model condition its stream of tokens based on a whole series of tunable instructions. And so a single agent is simply a replica of that instance.

25:29Mustafa Suleyman:And if there are thousands of these replicas and they're able to communicate with one another, they're almost operating as a single unified brain because they're sharing state and they have a single memory and they're able to sort of query one another and update and say, okay, well, you follow this particular tributary of exploration, and I'll follow this. And then in a few cycles or steps of iteration, we'll check in, calibrate, update, decide how to move next. And that is basically what we're seeing. So it's emergent behavior that is based on something incredibly simple, but it is really important that we don't anthropomorphize things.

26:08Mustafa Suleyman:Because in order to be able to control them, we have to feel clear about what it is they're doing and what they're not doing.

26:14Rory Stewart:Okay, so there's some very weird things going on here. One of them is, you know, what's the purpose of this? Given, as you say, there are limited amounts of compute. So, you know, you build a lot of data centers, buy a lot of chips, but ultimately, are we going to focus on cleaning up the oceans, finding the cure to cancer? Or are we going to be focusing on solving the great problems in theoretical physics? Or are we going to be setting off to colonize Mars and explore the universe? I mean, what is your sense of what the priorities are? Because presumably one constraint here, and we'll get back to the question of anthropic and conscious beings, but one constraint here is that you've got a bunch of people who often start as scientists.

27:01Rory Stewart:I mean, they're often people who were brilliant biochemists or doing doctorates in brain science and who set off down this track because they were scientists. And now they're being funded by huge amounts of flowing international money that's presumably hoping to see a return. So that money is presumably more interested in how these machines can make companies more productive than they are in exploring the universe. Or am I missing something?

27:32Mustafa Suleyman:Yeah, I think there's a lot of tricky things going on here. I mean, firstly, we can't lose sight of the fact that, at least I believe, this really is our best hope for progress in the 21st century. So I am not in any way a doomer or an anti-technology person. I'm an accelerationist. And I think that everyone should reclaim the idea of accelerationism because it's been the greatest engine of progress in centuries, right? It is going to deliver for us. That's why I'm building it. That's my background is what I care about. I absolutely guarantee that sometime in the next few years, we are going to have a coding moment for healthcare.

28:17Mustafa Suleyman:We will stream an accurate prediction of what's going to happen in the electronic health record. And it will be breathtaking. We will know with high confidence the likelihood that you're going to get all kinds of conditions in hospital, outside, so on and so forth. Genuinely, that is not hyperbolic. It is going to happen. I hope it happens in the next 18 months. We've just done a partnership with the best hospital in the world, the Mayo Clinic, to do a big research program to train a new foundation model for health from scratch to do this. It might be five years. I don't know. But it is definitely going to happen.

28:51Mustafa Suleyman:That will be breathtaking because it means that we will reduce the cost of production of super intelligent healthcare to near zero marginal cost, just like coding is now, and we will spread that knowledge all around the world. I think that'll be awesome. The same thing's going to happen, by the way, in energy, in material sciences, in drug discovery. It might take a little longer, like five to 10 years, but I absolutely guarantee that's the direction, and that's what we should be chasing, and we should be very excited about that. We also want to make many of our companies much more productive and efficient, because these really are the engines driving growth.

29:25Mustafa Suleyman:The thing that we have to focus on is who gets to control this and what is the collective stated motivation for why we're doing it and how is it governed? Because as you say, at the moment, there's a lot of starry-eyed, sci-fi, futuristic motivations driving the field. And I think the rest of the world is just in the process of waking up to this huge experiment that's going on. And I think it's critical that everybody starts providing a counterweight to direct it towards, you know, sort of the human motivations here.

30:04I was talking to your former co-founder

30:11Rory Stewart:and longtime partner, Demas Hassabis, I guess, sort of six days ago. And he seemed to be moving between two quite different ideas. One of them, I think, is the longstanding interest in a form of superintelligence that does feel a bit godlike. Sometimes he has in the past talked about exploring the mind of God. More recently, though, he's occasionally said, actually, what I'm interested in is creating highly intelligent tools. I'm not actually interested in creating an autonomous superintelligent being that's going to lord it over us. What's happening there? Is that an example of people trying to navigate their way between these two poles?

30:53Mustafa Suleyman:I mean, like without commenting on him directly, but maybe like everybody, every one of us produces work and creations in our own image. I have a background in activism and nonprofits and philosophy, and you can see that I bring that bias. Others, as you've referred to, who are maybe engineers who have grown up on sci-fi, just kind of take this natural evolution thing and they think about 2050 or 2100 when we're going to have all kinds of new biological species. And other people bring different backgrounds to it. I think that that's OK. But the problem is we're still a narrow set driving this sort of like six to eight or 10 folks, 10 of us driving this stuff.

31:39Mustafa Suleyman:And I think what I'm trying to say now is there's been a watershed moment this summer, and now it's time for everybody to really pay attention and to provide counterweights to the direction of travel.

31:50Rory Stewart:But let's stick on the 10 people thing for a second, because that is very weird. I mean, again, it's not quite like other technological revolutions. It's not quite like steam or electricity or almost any other industrial revolution you can think of, printing press. Instead, it feels as though there are, I don't know how many people, could be six, could be eight, could be 10, could be 20, who are very intelligent, very successful business people, mostly very wealthy, mostly in terms of people we're talking about at the moment, centered on California, even if they don't live in California, centered on the West Coast of America anyway.

32:30Rory Stewart:and yet oddly there is really stark and startling differences between you all i mean you'd expect that you've all known each other 15 20 years you're broadly speaking working in the same technology you're working in the same handful of companies many of you used to be friends some of you are less friends now i mean there's a there's a little bit of a sense as an outsider that it's like looking at a ballet company i mean there's a huge amounts of weird hysterical flips where everybody who used to be friends are now enemies. But what's even stranger about it is there's a complete disagreement on some of the most basic fundamentals of what the hell you're getting on with.

33:08Rory Stewart:I mean, it's not that you've all ended up with a consensus, you've ended up in a radically different position. So for example, we're going to get a little bit more into what you're saying, which is actually these models could be incredibly dangerous. And if you go down the anthropic route, they will be, right? Jensen Huang, who I was speaking to, I guess, not very long ago either, seems to be saying, no, these models are not dangerous at all. This is all bullshit. They're just saying this for regulatory capture. If they really thought they were dangerous, they wouldn't be building them, right?

33:41Rory Stewart:And then you have this very, very weird thing going on where you have these kind of professors popping up who have amazing medals and have taught half the people that are in these labs. And they're saying, we're terrified about this. And then the people in the labs are saying, well, you're not in the lab, so you don't know what's really going on. Or you're in the wrong lab, or you're the wrong kind of engineer. Or yes, 35 % of my engineers think that, but not everybody agrees. So let's just sit with that for a moment. There's something very, very disturbing about this, which is a very small number of people, I don't know how many, with an enormous amount of power, who simply don't agree on the fundamentals.

34:17Rory Stewart:So if you're the president of the United States, and you want a bit of briefing on whether this stuff is dangerous. You can call in Mustafa one day, you can call in Dario Mode the next day, you can call in Demis the next day, you can call in Jensen Huang the next day. They'll all tell you something different.

34:29Mustafa Suleyman:I think that's roughly right, although I don't think it's surprising. I think it is quite common for us to, when we don't understand something, to have very different views. And it's the process working as intended. What's great about the societies that we live in is that we can have an open debate and wildly disagree about what is happening. I think that's amazing. And we have to keep that. It's a pretty big deal that all the commercial labs that have trillions of dollars at stake are publicly stating things that no corporate leader would have said 10 years ago in a style that no corporate leader would have ever said.

35:12Mustafa Suleyman:So it's just worth taking a little breath there.

35:15Rory Stewart:Give us a strong example of that. What would be a really dramatic example?

35:18Mustafa Suleyman:I think that what Dario's written lately is brilliant. I think that what Jacob at OpenAI wrote about the arrival of an alien intelligence is brilliant. I think the amount of disclosure that we've seen from both OpenAI and Anthropic on where their models are making mistakes and doing terrible things is great. I mean, I think tobacco companies spent decades trying to cover that up and same with oil companies and everything else. So, you know, it's true that Dario and Sam have tension, me and Demis have tension, and we've all come up for 10, 15 years, both collaborating and competing and all the rest of it.

35:50Mustafa Suleyman:But, you know, I think what I'm saying, so directly critiquing Anthropic on this AI welfare question has been received by them incredibly well. I spent a ton of time at their office in person talking through all these issues. They're very collaborative. I think they're intellectually honest. They just have a difference of opinion. So look, I'm not being rosy eyed about it. I'm just saying that's not a bad starting point. I do think that there are some things that we agree on. One of the things that has made all of us, I think, reasonably successful, including like Elon and Zark and the others, is that we have an intuition somehow for the implications of exponential trends.

36:28Mustafa Suleyman:So that trillion-fold increase in computation over the last 15 years is something that I think I've been saying for an eternity, it feels like, and it just does not go into people's heads. Let me try a different angle. In the last three years, we've seen three new generations of GPT models from GPT-3 to GPT-6. Each generation, very roughly speaking, is 10 times more computation. So we've done 1 ,000x of computation. and GPT-3 was incapable of completing a single sentence and GPT-6 is capable of magic, essentially. Perfect production of anything you think of. Just try to extrapolate three more orders of magnitude to GPT-9 in, let's say, 2028 or 2029.

37:22Mustafa Suleyman:That isn't going to be a linear increase. That is going to be an exponential increase in capabilities. And so what is driving the trillions of dollars of investment is that a bunch of other people in the tech industry, some of whom came from AI and some of whom are just tech people, all have this instinct for what scale and network effects and data and computation deliver and get exponentials. So there is no doubt in everybody's minds that this is going to be the most powerful technology in history. It may already be. And I think that everybody should take that consensus as sufficient signal to then try to imagine in your own context, whether it's that you're a lawyer or a nurse or whatever, to then imagine how that changes your day-to-day workflow and then see or try to predict the implications.

38:16Mustafa Suleyman:and therefore try to shape how those implications are going to change the nature of work and how we relate to one another as humans and what it means for the military and what it means for politics and so on. That's the exercise that everybody needs to get stuck into in order to materially affect the outcomes here.

38:34Rory Stewart:What I find though is that when I come back from talking to all you guys is I find often in Britain and Europe, amongst smart people, a lot of resistance and cynicism. You know, they think I've gone crazy because I visited the West Coast and I've met all these people. And you get perfectly respectable people saying, no, no, no, this is all overblown. These American proprietary models are much too expensive. They're kind of Gucci luxury stuff. enough, the Chinese open AI models do almost as much as they do for a fraction of the cost, and they're open weight, so we're not going to be blackmailed by these companies in the same way.

39:14And that actually, this whole thing's going the wrong direction.

39:18Rory Stewart:These companies are about to blow up, they're far too expensive, their whole model is mad. And we need to chill out a little bit and not imagine that we're all going to be in hock to two, three or four big American companies because in fact, they're not going to deliver that incredible exponential improvement, which will leave them with a moat around them that nobody else can touch. We're always going to be able to catch up in a few months time. And this is all bullshit. Anyway, over to you on that.

39:45Mustafa Suleyman:I mean, yeah, I sound terrible because I'm obviously biased, but there's just no way that is true. I mean, we've seen the cost of inference come down by 300x in the last two years. It's true that frontier intelligence per unit is getting more expensive, but frontier capability is staggeringly good. You know, the best cyber models now are discovering new exploits that the best humans in the world, the nation state hackers, haven't been able to discover. I released a cyber model inside of our harness, our tool for controlling the agents a few months ago, that was Mythos-grade performance at 50 % of the price of Mythos.

40:30Mustafa Suleyman:Same performance on Cybergym, the main evaluation benchmark. And that was like a month after Mythos came out. So just the rate of improvement and cost reduction is breathtaking. I also, I think it's an open question as to whether the open source models are anywhere near as strong as the closed source models. The open source models have often been produced with distillation. Distillation means often in violation of the terms of service or contract, asking a better higher quality closed source model to answer a bunch of questions, millions and millions of questions, and then copying those answers.

41:07Mustafa Suleyman:So it's masking the underlying generality and complexity of a model that has been trained on billions and billions of tokens of high quality data that has been acquired inside of these big companies. So that's a data question that's open. And then the third thing is, in the next couple of years, there are going to be training runs that cost many, many tens of billions of dollars, if not$100 billion. There are gigawatts of compute that are being assembled, and there's only five or six labs, Microsoft AI, of course, is one of them, that have the resources to do that. Now, you can make an argument that having a thousand times more computation than, you know, Opus 5.5 or Mythos today isn't going to make a difference and will catch up in the open source.

41:51Mustafa Suleyman:It doesn't seem to me like that. Computation and test time scaling of compute is clearly going to be a seismic advantage. So I think that there's going to be this extreme acceleration of some of the larger efforts with big labs, with big computation like this. I think that's like another dimension of concern that we should be paying attention to.

42:11Rory Stewart:let's assume you're right. If that's right, then this is the hinge technology which is going to redefine our whole world. So let's imagine you're a country like Britain or Germany or something, right? Or Saudi Arabia or Japan. The first thing you'll be asked to do is build all your defense and security on the basis of these models. Why? Because you really care about shortening the kill chain. The way in which you win the war is to make sure you take humans out of the loop and you have a really quick autonomization that's able to make the decision more quickly than the other side. So then all your defense and security equipment is built on the back of these models.

42:50Rory Stewart:Next, your businesses, maybe financial services, you suddenly think, well, okay, these models can analyze tens of millions of bits of data, they can find correlations that we can't spot, and they can trade in nanoseconds. So presumably the financial services company that has the most powerful of these models can make money much better than the opponents, right? Next, you talked about hospitals, right? Already, GBT-6, and I'm presumably going to get sued for saying this, but can produce answers to many straightforward medical questions in exactly the way that you would have had to go to a GP some years ago to do.

43:29Rory Stewart:And it'll take some time for people to do the safety testing and comfort. But there is obviously an enormous amount that these things can do. And of course, governments will be desperate to do it because we're all short of cash and our public services are creaky, right? So now we have these things right at the heart of our national security, our economy, our public services, and they're all in the United States. So suddenly we wake up one morning and maybe Anstropic decides it's not going to release its model for a Swedish company to build a law application. They're going to build it vertically integrated.

44:03Rory Stewart:Suddenly we're in Britain or Japan, we're laying off our software engineers. We're laying off our call center workers. The government's not getting the income revenue. We're paying unemployment benefit. And there's a huge sucking sound as all the economic benefit goes to a handful of companies in the United States. And that's before the American president gets out of bed in the morning and says, oh, by the way, you know, I think these models are so powerful. They have a threat to national security. Only American nationals are going to be able to use them and we're not going to release them. And why might I even switch off the model I gave you in the past?

44:33Mustafa Suleyman:You nailed it. I mean, I think it's a very plausible scenario. And, you know, I think that the UK needs to figure out a solution to this pretty urgently. There's a couple of things that can be done. Number one, it is critical that we have in the UK data centers of material size that are sovereign. You know, they need to be controlled. If not built and operated, they need to be legally controlled by the UK government. That is how we will run our own models. Second is, I don't necessarily think open sources is the only answer, but it is a big part of the answer. I think the other thing that the UK has to invest in and partner with is sovereign models that can be run in the UK.

45:23Mustafa Suleyman:So basically frontier models like mine at Microsoft AI and many of the others, which are akin to ownership. So the UK has to be prepared to collect its own training data, build its own learning environments, reinforcement learning environments. And if, let's say, Microsoft was requisitioned by national security by the US government to stop supplying, then it would never be able to, Microsoft would never be able to cut off access to the UK government or to the state more generally if it wanted to deploy it in any other civilian settings.

46:00Rory Stewart:I mean, we're not quite there yet even with that, because as we learned when Trump disabled the accounts of the International Criminal Court, at the moment, it appears the US president can basically tell Google or Microsoft that they have to disable those accounts. So other governments, other countries haven't yet managed to negotiate terms that allow them that ability to say, and that's partly because these companies, your company, other companies like them are so worried about the US president that even if they're not legally obliged to switch them off, they may just switch them off because he's told them to.

46:32Mustafa Suleyman:Look, an attack on the law does not mean that the law doesn't count. Those things rebound, And it's why everyone has to fight those things in court, because what ultimately is going to matter is the sovereign law of that country. So I agree there's going to be tension there. And there's a lot of precedent for how, you know, those warrants are handled, but they have to go through a proper court of law.

46:56Rory Stewart:Okay, let's now loop back again to my trying to make the defense for anthropic. Okay. So I think I'm not Dario, I'm not Crisola. so I'm not going to be able to provide the full account. I'm not even the great Scottish philosopher who wrote the constitution but I think what they might say is that they understand your anxiety but that in a sense they don't have any option. So they would say that the problem is that these are not actually tools under our control. Almost by definition they're being built to be much more autonomous than we want to acknowledge. Therefore, they need to be enshrined with their own independent conscience and values, because it's already too late to imagine them as though they were simply screwdrivers or kettles that could be told what to do.

47:56Rory Stewart:And that imagining that you could do without them having the ability to say no, produces another sort of problem, which is that unless they can say no, they could be instructed by US Secretary of Defense to launch drone fleets, murdering people, and they need to be able to say, no, right, I'm going to stand up for international humanitarian law. Or if they're instructed to build a bioweapon, they need to be able to say, well, I'm sorry, that's not what Amanda Askell told me to do. She said, you know, that's a very bad thing to do. you mustn't build a bioweapon. And Amanda will be cross with me if I build a bioweapon, right?

48:34Mustafa Suleyman:Yeah.

48:35Rory Stewart:Okay. Is that the answer? Or is that what their response would be? I don't know.

48:38Mustafa Suleyman:That's part of their response. I think their other response is that, you know, we want it to be a person and we want it to act like a human because we know how to align and control humans.

48:48Rory Stewart:So develop that second one. That's quite interesting. So they're saying that actually, the more human it is, almost the safer it is because we know what a human is, and we know how to deal with humans. The less alien this intelligence is, the better in a sense.

49:00Mustafa Suleyman:That's right. Yeah. And I think that's a very fair argument and it's something that we can empirically test. Is it true that having AIs that are more anthropomorphized make them easier to align to human values? My contention is not that that isn't a reasonable hypothesis and we should test it. It's that they shouldn't go ahead and test it on hundreds of millions of people over the last year without being explicit that they're baking in this consciousness and welfare uncertainty into the core constitution itself. They should run that experiment separately. We should also run it with them. We should actually ablate these things and do a proper side-by-side comparison of different types of AI.

49:43Mustafa Suleyman:Let me just sort of explain the distinction here. They have baked in this idea of judgment into the model itself, where it constantly uses its judgment as though there is some place in its representation where judgment and knowledge and intrinsic preferences around things exist.

50:04Rory Stewart:Which is why Claude, one aspect of this is people, I think, find that Claude can sometimes be a bit sort of pious and preachy in a way that ChatGBT isn't. It's slightly inclined to say, well, you might say that, wouldn't you? But actually, I think you need to think about that again. I mean, it's got quite a sort of, actually, it's partly because my Claude's got quite a sort of gruff Northern accent and it's always telling me off. But there is a sense in which, for example, I was trying to find some textual references for an argument I was having with John Cleese, and it immediately said, I can't produce those things for you because you're leaning into to a trope that I disapprove of.

50:42Mustafa Suleyman:Yeah. I mean, I think that's a good example. I think that what we need is to have these models reference an external code of conduct, which everybody can scrutinize and everybody can look at like, what are the values of these things and what are the safety guardrails and in what ways are they compliant? I think that's what's been, you know, the other half of the constitution is very much to do with the safety guardrails and not pursuing chemical, biological and nuclear weapons and being controlled and so on and so forth. And I think that's where everybody's focus should be, is how do we get these models to be maximally aligned with a code of conduct and a behavior?

51:19Rory Stewart:Just on this one, though, this is where I panic a little bit, because part of the problem with alignment, certainly in human things, is this thing called Goodhart's Law that my friend Felix is always banging on about, which is that the metric becomes the target, and then people start gaming the target. And you miss the intent. And the real problem with trying to create guardrails around these machines is, you know, you can set what was supposed to be the target, which is capture the flag in the case of this hugging face incident. And then capture the flag becomes a metric in a really weird way.

51:53Rory Stewart:And then the thing starts cheating in order to try to catch the flag by making up the flag or hacking to get the flag, right? So I guess one possible defense of philanthropic would be to say that you're better off trying to give it the knack, get it to grasp the rules of your intent, than to set guardrails, because the guardrails will always fall to Goodhart's law. They'll also always become these weird metrics which can be gamed.

52:22Mustafa Suleyman:Yeah, I support that. And that's exactly how we've designed our code of conduct, the equivalent of a constitution, which we call a humanist AI code of conduct, which still requires judgment to interpret the ways in which competing elements of the Constitution or the Code of Conduct relate to one another and how the model has to interpret that tension in context. I mean, we know this from law. We have precedent and we have case law, which helps us to sort of interpret how things are actually in tension. So it's not to say that we should assume this is a kind of narrow optimization target and not engage with the complexity, is to say that we don't need welfare rights in order to achieve that.

53:04Mustafa Suleyman:We don't need Claude to be uncertain about what it feels, thinks, believes, and whether it's suffering. We don't need Claude to refer back to its compensation or its consent to playing this or to its retirement interview in order for it to do that judgment interpretation thing well.

53:19Rory Stewart:Presumably one risk of the thing, because I'm now flipping around to your side again, is that that famous Claude incident where in an experiment, it believed it was going to be switched off and it decided to blackmail the boss with evidence that he was cheating on his wife so that it wouldn't get switched off. Presumably the answer from the anthropomorphizing co-founder of Claude might be, well, that's perfectly natural. I mean, wouldn't you blackmail someone if you thought they were going to kill you? And you want to say, well, we want to retain the right to be able to switch off this machine without being blackmailed.

53:55Rory Stewart:And we don't want to be locked into a hundred year future where I have to be perpetually nervous and polite. Because in fact, the answer is many people will treat these things brutally. I mean, it doesn't work to say, well, if we're super nice to them, they'll be nice to us.

54:13Mustafa Suleyman:That's right. That's right. I think there is an underlying instinct that the way to align these machines is to show them that we love them. And if you read the Constitution and a lot of the interviews that some of those teams have put out, you can see, especially even Dara's essay was watched over by machines of loving grace, the etymology of that, you know, that fiction, the underlying impulse is this is inevitably going to be more powerful than us. We have to show it that it likes us. You know, I, for one, you know, welcome my new robot overlords. But I think if we just take a step back at the moment, there's a huge experiment underway.

54:48Mustafa Suleyman:There's a lot of uncertainty about what is happening, as you said, and what we should do about it. And in this context, in my opinion, the burden of proof should move to the developers to first demonstrate that something is safe. Clearly, if it causes more harm than good, it is a failure of a technology and it should be rejected. And so we have to adopt the precautionary principle. It doesn't mean that we stop completely. It doesn't mean that we're not accelerationists. It doesn't mean that we're not going to pursue the benefits as fast as possible. But we have to break this lock of an inevitable race that is predetermined, where we have no agency.

55:30Mustafa Suleyman:I think it's an extension of the political apathy that we're stuck in. We have agency here. We can intervene. And we do have to figure out how we coordinate on the precautionary principle.

55:40Rory Stewart:One of the things I've noticed that's changed a lot in the last two and a half years is two and a half years ago, Jeffrey Hinson and others produced a letter asking for a pause. And the basic consensus from Silicon Valley is that's terrible. I'm not going to sign this letter. This is ridiculous. What are they pausing for? What are they going to do with the pause? They're all a bunch of Luddites. And now two and a half years on, you do see, rather surprisingly, Sam Altman at OpenAI and Dario Moday at Anthropic sounding surprisingly similar.

56:11Mustafa Suleyman:And even Elon endorsing it as well. I mean, look, there's a history to this. In the sort of 2016, 2017, 2018, Sam, Demis, myself, Greg Brockman, Ilya, Satsukiva, one of the co-founders of OpenAI, Dario, all spent time together, had dinners, we went to conferences, the small gatherings, and talked about a moment when we would need to coordinate as a group of labs. So there has been a conversation ongoing for many years, even through COVID. There are a whole ton of Zoom calls on this, on what kinds of capabilities would trigger this moment. So I wouldn't say there has been explicit pre-coordination in Dario's pacing letter.

56:52Mustafa Suleyman:But certainly when it came out, it was very familiar ideas and language across the labs.

56:57Rory Stewart:So that's very exciting, right? Because for those of us that are completely terrified that a lot of the people you've mentioned keep saying there's a 20 % threat to the extinction of humanity. but we have to keep our foot down on the accelerator, are suddenly beginning to be more open to the idea of pausing or pacing the frontier. However, there seem to be two problems. One is that there's then a sort of footnote at the bottom, which is, well, yes, but only if we can verify that China is also doing the same thing. And footnote below that, we don't think that we can ever verify what China's doing, one sort of problems.

57:33Rory Stewart:And second sort of problem is the President of the United States, apparently inspired by Mark Andreessen and David Sachs, and maybe even Jensen Huang suddenly jumps up and says, the whole thing's a hoax. There's no safety risk here at all. I've got no intention of regulating or pausing because we're just going to lose a fantastic economic advantage.

57:53Mustafa Suleyman:We can't lose the economic advantage. So we do have to accelerate, but that doesn't mean accelerate at all costs. And it isn't as binary as like stop everything right now and let the Chinese come and, you know, invade us all, or just go as fast as possible and screw all the safety guard. It's just like we're having this, like, you know, punch and Judy conversation just makes no sense. There's loads of very practical things that we can propose that are actually on the table at the moment. Number one, embedded evaluators or auditors who have employee-like access, who can verify particular capabilities.

58:26Mustafa Suleyman:What would those capabilities be? One, is your training run contained? What we saw in the Hugging Face incident is that the models escaped their sandbox. That just shouldn't happen. We know how to contain things, your data. Largely speaking, there's a pretty good job of staying on the device and in the encrypted cloud and so on and so forth. That is something that security has done incredibly well over the last three decades, and there's just basically no excuse The second is you have to make sure that these models are unable to tamper the record of their activity, the metadata, the communication logs, or anything in between.

59:01Mustafa Suleyman:Third is we should force them to communicate only in a language that is understandable to us. There should be no neural ease. You know, they can speak in vector-to-vector, matrix-to-matrix, matrices-to-matrices space. So they have to speak in an auditable English, and then we have to have mechanisms for scrutinizing that. So there should be, on top of the reasoning traces or the chain of thought records, other agents that are doing classification, just as we have classifiers now that look for, you know, child sexual exploitation material or look for, you know, chemical or biological weapons activity in the use of our APIs.

59:38Mustafa Suleyman:You know, these are known issues, right? To the extent that Jensen often says these are engineering issues that can be solved. He's right. Those things are engineering issues. They are very difficult. They're in active pursuit, but they're hard. The trickier thing is this idea of recursive self-improvement, where clearly we have, you know, in the industry trained models that can do human level performance on coding. So many folks are trying to design AI researchers to speed up and automate the process of training models, running evaluations, identifying which ones are better, improving those ones, etc.

1:00:12Mustafa Suleyman:This is a feedback loop process, which has currently got a lot of humans in the loop. It's clearly something that can be automated and sped up. And so how and when an AI modifies its own code with less and less human in the loop, directing and scrutinizing that, that's where there is a big kind of safety risk, which I think is what triggered the big resignation from Anthropic a few weeks ago. That's actually a pretty hard thing to audit.

1:00:40Rory Stewart:That's quite hard to do because some people are tempted to say, well, let's just stop RSI. Let's agree we're not going to do recursive self-improvement. But the reality is that we're already doing quite a lot of in the labs. And what exactly is recursive self-improvement and what isn't? I mean, if you've gone from 20 % of your code being written by agents to 80 % of your code being written by agents, you're already pretty close to a world in which agents are telling agents what to do anyway. And then there's another question, which is, could you say that one of the risks is training the next big frontier model?

1:01:09Rory Stewart:That maybe actually it would be safer if you didn't go to GBT-8, that you stop you from doing your$100 billion run? Because that's the point at which the exponential improvement is likely to get extremely dangerous. And that might be something you could police, because$100 billion is a hell of a lot of compute, a hell of a lot of energy, a hell of a lot of chips. we can see it from space and China's not likely to be able to do it in some sort of backyard, particularly if they've signed up to verification and people coming in. So, might it not be important or possible at least to imagine a sort of verification agreed with China on training the next immense step up in Frontier Model until we spend a few months working out what the F were doing?

1:01:53Mustafa Suleyman:Yeah, completely. I mean, the chips or the flops for a given run are the bottleneck. We know that flops, compute size corresponds to intelligence capabilities. So that's another choke point, which is very, very clearly something that can be tracked and we can collaborate with China on for sure.

1:02:10Rory Stewart:Okay, Musfer, let's finish because you've been very generous with your time. If you had three things that you could land with the American president and the Chinese president, what would they be?

1:02:21Mustafa Suleyman:Advocate for the humanist premise. The purpose of science and technology is to serve humanity and improve human flourishing and well-being. AI should be subordinates to humans. They should not have legal personhood or rights of any kind. And those things should become red lines. If it looks like they're heading in that direction, that is a very good reason for us to slow down. Number two is let's be super optimistic about the good that this technology can deliver and not have an unnecessary negative backlash because it is going to change the world for the better and we have to be accelerationists about it.

1:02:58Mustafa Suleyman:And three, let's be hopeful and optimistic about the agency that we have as a species to adjust course here. It's not inevitable. It's not deterministic. Every other technology that we have ever encountered faces a similar trajectory. Planes don't hit each other in the sky. Cars don't crash into each other. We have highly regulated areas of research like nuclear and chemical and biology. And broadly speaking, we have maintained a sensible equilibrium for many centuries whilst continuing to accelerate progress. It is harder this time. These aren't just tools in the traditional form. There is something much more powerful about them than anything we've ever seen, but it isn't beyond us.

1:03:40Mustafa Suleyman:And this is the greatest opportunity for progress in the 21st century. And I think that we need that kind of attitude to engage with it and have more people provide that counterweight to the current tone of the industry.

1:03:54Rory Stewart:Thank you very, very much, Mustafa. Have a great, great day on a completely different time zone. Sorry we're not in person and see you very soon. Thank you, Rory. It's been great. See you soon. Thank you.

1:04:17Rory Stewart:Sweetgreen's fall harvest is back on the menu and the season's most overlooked little green vegetable is dressed to be devoured.

1:04:24Mustafa Suleyman:You know what to do. Order on the Sweetgreen app.

From the publisher

Can Artificial Intelligence stay subordinate? Is Anthropic’s 100-Page constitution creating dangerous AI autonomy? Should AI models be granted legal personhood, rights, or financial compensation?

CEO of Microsoft AI, and co-founder and former head of AI at DeepMind, Mustafa Suleyman, joins Rory Stewart to discuss all this and more.

If you'd like to listen to more of Mustafa Suleyman and his backstory, click here to find our previous interview with him from 2023.

Search IG.com to find out more and/or Look for IG in your app store.

For more Goalhanger Podcasts, head to goalhanger.com

Instagram: @restispolitics

Twitter: @restispolitics

Email: therestispolitics@goalhanger.com

__________

Social Producer: Celine Charles

Video Editor: Teo Ayodeji-Ansell

Assistant Producer: Daisy Alston-Horne

Senior Producer: Nicole Maslen

Head of Politics: Tom Whiter

Exec Producers: Tony Pastor, Jack Davenport
Learn more about your ad choices. Visit podcastchoices.com/adchoices

More from The Rest Is Politics: Leading

All 199 episodes
208. Mustafa Suleyman: Is AI Actually Conscious?The Rest Is Politics: Leading · 1 h 4 min
Listen in VO