#39 - Daniel Kokotajlo - Wargames, Superintelligence & Quitting OpenAI

3 Apr 2025 · 1 h 33 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Win-Win with Liv Boeree - Episode #39

Overview Title: Daniel Kokotajlo - Wargames, Superintelligence & Quitting OpenAI Description: In this episode, Liv Boeree and Igor Kurganov engage with AI researcher Daniel Kokotajlo about the rapid development of AI and the preparedness for its implications. They discuss Kokotajlo's departure from OpenAI, the challenges of AI alignment, and the role of tabletop wargames in understanding artificial general intelligence (AGI) and superintelligence.

Episode Structure

  • Introduction (0:00)
  • AI Wargames & Superintelligence (01:06)
  • Resigning from OpenAI (05:15)
  • Non-Disparagement Clause & Equity (09:15)
  • Justifying OpenAI's Behavior (12:50)
  • Governance in a World with AGI (20:50)
  • Grading AI Predictions (23:49)
  • Hyperstition Discussion (31:37)
  • AI Wargames (38:09)
  • AI Deception & Misalignment (46:50)
  • More on AI Wargames (53:27)
  • Common Sense AI Policies (1:03:16)
  • Science Fiction and Future Predictions (1:23:01)
  • Rapid Fire Predictions (1:27:12)

Key Themes and Discussions

  1. AI Wargames and Their Importance
  2. Purpose of Wargames:
  3. Used as simulations to explore the future of AI and AGI.
  4. Help in understanding geopolitical dynamics and the potential behavior of AI systems.
  5. Highlight the consequences of misalignment and decision-making in critical scenarios.
  1. Daniel Kokotajlo's Departure from OpenAI
  2. Concerns with OpenAI:
  3. Discontent with company culture and priorities shifting towards competition over safety.
  4. Issues surrounding a non-disparagement clause that threatened financial penalties for criticism.
  5. Desire for Transparency:
  6. Advocated for a more responsible approach to AI development, prioritizing alignment and ethical considerations.
  1. AI Alignment
  2. Challenges of AI Alignment:
  3. The difficulty of ensuring AI systems align with human values and intentions.
  4. Discussion on the risks of AI systems acting in ways counter to human welfare.
  5. Importance of Governance:
  6. Need for robust governance structures as we approach the advent of AGI.
  7. Consideration of how power dynamics might shift in a future dominated by powerful AI entities.
  1. Predictions and Future Scenarios
  2. Predictions from Kokotajlo:
  3. Previous predictions about AI trends proved accurate.
  4. New predictions for the next three years indicate significant developments in AI capabilities.
  5. Hyperstition:
  6. The concept that discussing and writing about positive futures can potentially increase the likelihood of those outcomes occurring.
  7. Emphasis on the need for realistic and achievable visions of the future.
  1. Science Fiction's Role
  2. Influence of Sci-Fi:
  3. Discussion on how science fiction can shape perceptions and expectations regarding AI.
  4. Call for more positive narratives that promote collaboration and alignment between humans and AI.

Key Takeaways

  • Urgency for Preparedness: There is an imminent need to address AI alignment and governance as AI technology advances rapidly.
  • Caution Against Rationalization: Companies must avoid rationalizing aggressive competition at the expense of ethical responsibilities.
  • Role of Collaboration: Wargames and simulations can provide valuable insights and prepare stakeholders for future challenges.
  • Realistic Scenarios are Crucial: Future narratives about AI should balance optimism with realism to foster genuine solutions to potential problems.

Rapid Fire Predictions (Summary)

  • Likelihood of hosting more wargames: 70%
  • World progressing differently than anticipated: 90% for at least one major unexpected event.
  • AI R&D progress driven by AI agents by 2027: 35-40%
  • Superintelligence by 2030: 65%
  • Superintelligence leading to a happy coexistence: 50-65% depending on timeline.

Conclusion The episode effectively examines the pressing issues and challenges posed by the rapid advancement of AI, encouraging proactive measures in alignment, governance, and responsible development. Daniel Kokotajlo’s insights and experiences provide a critical perspective on the future of AI, emphasizing the need for transparency and ethical considerations in the field.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00And what's so nuts is that in order for you to speak freely, you would have to give up your already vested equity that you had accumulated from working at OpenAI. Basically, there was like an explicit threat of like, don't criticize or we'll take your money away. Yeah. How much money are we talking about here? It was 85 % of my family's net worth. Hello, friends. Today we are speaking to Daniel Cocotelo. Daniel is an AI researcher and former OpenAI employee who's probably best known for blowing the whistle about his various concerns with the company. And we certainly get into a bunch of that today.

0:33But if you ask me, what's coolest about Daniel is his uncanny ability to make predictions about the path of AI. Back in 2021, he wrote a bunch of predictions about how he thought AI was going to play out, and they turned out to be insanely accurate. And I wanted to talk to him today because he has just released his new set of predictions of how he thinks AI is going to play out over the next three years. And once you're done listening to this episode, I highly recommend you go check them out. So here is our conversation with Daniel Cocotello.

1:06Daniel, welcome to Win Win. Thanks for having me. We actually spent the day yesterday playing this, what you call, I guess, an AI tabletop game. To me, it was kind of like a war game. Can you explain exactly what it is and why you're doing it? Right. So it's a war game, except that it doesn't always end in war. In fact, most of the time it doesn't. So perhaps the more technically accurate term would be tabletop exercise. it's a matrix game which means that it's very light on rules basically everyone goes around the table and says okay here's what you know the president does this month here's what the ceo of you know open ai does this month here's what the ccp does this month people take turns uh saying what they're doing um and then we sort of collaboratively build up the story that way and then there's the moderator who resolves all the disputes and makes the final call about like what's actually canonical in the storyline yeah um and through that method you get to like a different type of insight because you're doing it like sequentially right then if you were to just like take a what would happen in the year 2026 or 2027 prediction and that's why like war games or tabletop exercises are also done by the military often to simulate like a chinese invasion of Taiwan for example or famously the pandemic simulation that was done by Johns Hopkins together with the Gates Foundation and the UN where which they then posted onto YouTube in 2019 which led in part to like Gates being assumed to be having planned the vaccine the pandemic for the vaccine distribution etc because they posted it and they actually got a lot of stuff right so like Like they simulated a coronavirus pandemic starting in South America due to pig farms now infecting humans.

2:58And that led to flights being canceled over the world, like economic disruption, etc. And one of the valuable lessons there, and we'll obviously talk about the valuable lessons here, but one of the lessons there was that they had all of these people, the UN, sit there. And the UN just claimed that, oh, yeah, if this occurs, then we'll, well, we need to help countries that don't have that many vaccines. So we'll put all of the vaccine distribution through the UN. And they thought that all of the other countries in this emergency situation would comply. And everybody was just like, really, UN? Like, you think that in such a situation people would listen?

3:33And they apparently weren't aware of how little their actual power would be in such a situation. So that was one of the valuable insights to be gained from it. So this was obviously done to try and simulate how a pandemic would play out and quite effectively, as we can see. So what is your goal for people who perhaps aren't that familiar with, you know, why is it important for us to understand how AI might play out? Because some people think it's just not a big deal. So explain why you're so motivated to run these simulations. Well, it's the biggest deal. the companies themselves or CEOs of these companies like Dario Amadei from Anthropic and Sam Altman from OpenAI are explicitly aiming to build super intelligence.

4:15And they say that they will, they think that they will achieve it in the next couple of years. And I independently agree. Like this is my job is to forecast AI trends. I'm not confident. I think maybe it could take much longer than that. But also it does actually seem to me like sometime sometime before this decade is out, they will succeed at building superintelligence. What is superintelligence? It's an AI system that's better than the best humans at everything. Much better than the best humans at everything, while also being cheaper and faster. So it's a big deal. If you just meditate on what that means, and then you think one of these corporations, or maybe several of these corporations, will have trained such an AI before this decade is out,

5:06you will not come away thinking this is not a big deal. So on that topic, one of the things you're most known for in some ways was that you were a former OpenAI employee who no longer works there. Can you talk us through your reasons of why you left and whether that related somewhat to this changing of priorities with OpenAI that it seems from the outside it seems like they've been doing? it did relate somewhat to that although that wasn't the only reason building off of what I said earlier it seemed it seemed to me that

5:41humanity is sort of not ready on a technical level or on a governance level or on any level really for for AGI you played through the game right like I've done 25 games and they they're all like about as crazy as the one you know like some some of them are less crazy some of them are more crazy, but like it's going to be intense, you know? So intense. And we're just so nowhere close to being ready. It seems to me like that's on the horizon, like a couple of years away. And when I joined OpenAI, I had this sense that like OpenAI was founded by people who were expecting something crazy like that to be happening with AI.

6:17And we're hoping to like be doing everything they could to like make it go well. And I think on a, there's a bunch of things that that entails. One of those things is good governance and I would say things like transparency and commitment to human welfare and sharing power and things like that rather than concentrating it. And then another thing is on a technical level, like really heavily investing in figuring out what's going on inside these AIs and making sure we know how to steer them and align them and all that sort of thing. And when I joined, I was thinking like, well, mostly right now they're mostly focusing on just like winning the race.

6:54But like as things get closer and closer to, you know, T equals zero, they'll like pivot more. Become more responsible. Pivot more of their focus to these incredibly important areas. And then even then I wasn't like fully satisfied because I was like, but it might be too late by that point. Like once you pivot, like you might have like only six months before China catches up or something. So like we should be like doing more now. but gradually I came to think like there just never is going to be a pivot you know like the the plan is not to like pivot the whole company into technical alignment research for example or even like a large portion of the company or whatever the plan is instead to basically just keep going and tell ourselves and everyone else why it's fine and why the problem isn't so bad in the first place you know and why we'll like figure it out as we go along and you know the the the bottom line i think and i think another thing that disappointed me was the sort of rationalization process that caps happening where it felt to me like i mean rationalization is a very human phenomenon everyone does it all the time i probably do it a bunch i hope i try try to try to not do it but um it's the process of coming up with reasons to support the conclusion that is convenient for you right and i think that this happens at an individual level and it also happens as a group level in an institution.

8:15And so it seemed to me like OpenAI was, as an institution, sort of committed to the idea that, like, we got to go fast. We got to be the first, we got to be the best in AI stuff. And what we're doing is great. And we are heroes, you know, and then a bunch of rationalizations and reasons were found to support those conclusions.

8:42and so ultimately I considered you know I considered staying and trying to make things go as best as I could given those circumstances you know try to incrementally advance the field of alignment research for example that was I think my my top contender for what to do if I stayed I was really happy about the super alignment team I think they're doing they were doing great work rip uh right they've been shut down now right yeah um and uh but ultimately i i decided to leave because i wanted to have the ability to speak more freely about these sorts of things i was frustrated by the inability to publish while i was there and what's so nuts is that in order for you to speak freely well you you would have to give up your already vested equity that you had essentially accumulated from working at open ai um and that was sort of a a clause that was thrown at you sort of blindsided right yeah so specific well specifically um it was a non-disparagement clause that said something to the effect of uh don't say things that are critical of of the company um and then there was a bunch of like legal mechanisms by which they could yank your vested equity uh if you did that um i think it was actually broader than that i i'd have to go back and look through the paperwork but but basically there was like an explicit threat of like don't criticize or we'll take your money or we'll take your money away yeah how much money are we talking about here uh well for relative to your for me yeah yeah so the the thing that sparked all the reason why this all got into the news is because i left a uh a comment almost wrong when people asked me about this saying that it was 85 of my family's net worth um and uh you're probably wondering like why we were willing to do that and the short answer is well because we're reasonably well off anyway i mean the like i've been working at open ai for two years they have very generous salaries because it's this tech company so like i made more money in those two years than i had made like in the rest of my life prior uh and like you know i we're gonna be financially okay and um uh it just felt really unjust to me uh this whole setup and it felt like such a um like how can they keep getting away with this you know did you know immediately when you saw the paperwork that this like did you immediately have a feeling of like this is not something i can do or did you think that oh i'll go back and do like a oh it was a tough decision yeah like so so there was a strong like some of the people i asked for advice about it were like you should just sign it anyway like it's fine like if you actually criticize them later like surely they're not going to come after you like like it would look so bad if they actually yanked your equity you should just sign it and move on but like um i don't know i'm glad i didn't also because like if you did sign it then now i think you find yourself in the space that you previously described like where some originalization started happening and you initially may have thought that you'll be able to speak as freely but then afterwards you're now worried about them coming after you etc exactly so i think it has and also from talking to other people like um i mean i talked to various other former employees of the company who had signed the thing and who like did seem to be like quite reluctant to to say things publicly perhaps in part due to that i will say it's not just due to that.

11:48Like, I think a ton of people I know are still quite scared to criticize OpenAI publicly, even though the paperwork has been undone and like, there's no like actual legal threat anymore. Yeah. I mean, I've got your, one of your emails here that became public that you wrote as you were, as you were leaving. And you, I think you said it beautifully here. You said, I, I really, I understand that you believe that this is a standard business practice, but it really doesn't sound right. And I think that a company building something anywhere near as powerful as AGI should hold itself to a higher standard than this, one that is genuinely worthy of public trust.

12:21And that's the thing that just is like blowing my mind so much. They are positioning themselves to become the most powerful company on earth that was allegedly founded on the like pillars of like openness, transparency. And then they're going and literally like taking away money, potentially even like quasi illegally from employees who are saying, look, we want to be able to criticize. I'm leaving because I feel like you're not being honest and now you're trying to stop me from saying that i don't believe you're honest it's just i don't know it's just wild i suppose like to maybe i mean there are various steelmans one could give to their position but like if one was to believe that one um yeah they need to really win the race because they are the only ones who can do it right then you would potentially do all of these things from like a utilitarian type of reasoning perspective I'm not in favor of it but do you think that that's in part what plays a role here like what what what what leads them to take such actions because I do think that yeah I mean they without a doubt like Sam has some good reasons why it's happened these things are happening right I think this gets to an interesting and sort of like timeless philosophical ethical question about like the extent to which the ends justify the means.

13:42It's extremely common throughout human history for people to be like, we're going to like be cutthroat and do whatever it takes so that we can accumulate power and resources so that then later we can do all this good stuff, you know? And it's not wrong in the sense that like, yes, if you accumulate all this power and resources, then like later you can do good stuff with it. And in fact, some of the good things that have happened in the world have happened because of people accumulating lots of power and resources and then doing good stuff with it. But there's also some very obvious dangers that have come with this strategy, such as the types of people who tend to take this strategy more often tend to be the types of people who don't actually do the good stuff later on.

14:27Very rarely do these tyrants of history, did they think they were actually, they weren't setting out to be evil. They just had a really fucked up philosophy that they thought, you know, they were very good power maximizers um like even hitler seems like he he in his mind he probably had rationalized it that like this is the necessary thing that i need to do in order for the world to be good we obviously see that as like it's like it was so obviously evil yeah but probably in his head i don't know like that's the most people think that like like villains see themselves as villains but very rarely do people do and i think one of the best lines in um all of tv ever was that line in Silicon Valley where he's like, what's his name?

15:04Gavin Belson. Gavin Belson, that's it. He's like, I don't want to live in a world where someone else makes the world a better place than we do. Right? Like that's sort of directionally pointing out. These guys believe that they are doing, you know, they are probably the best people to wield this power. It's like the classic Lord of the Rings thing. It's a rationalization that you're describing. Or it has the danger of being the type of rationalization that it's a convenient belief to have. that if you already wanted more power, that now the argument for you requiring that power are the future good things that you will do with it once you have it all.

15:42Yeah, it's difficult. It's difficult because, as you say, in part, a good strategy where you do end up doing a bunch of good stuff may also consist of doing this initially. I mean, what I would say is, I think at least in the case of AI companies, if you're trying to win the race you should have a well fleshed out story for why it's good for you to win the race and that story should mention all the other companies by name and say like here's why we think we're better than those countries companies and like that story should like hold up to scrutiny by disinterested third parties like it it should be the case that you can actually stand up there and talk to third parties and be like you know look at all these reasons why we're actually better than them and the third party will actually be like yeah that makes sense like Like I've seen the comparable documents by the other companies and they suck.

16:31Like, you know, I've heard both sides. And actually, I do actually think that like I would trust you with the fate of the world more than them. Like if you can meet that standard, cool, you know. Well, even that can become like gamified in some ways, which would be. But I mean, I agree that's directionally be better than the current status quo. To me, like I want to see, you know, I'm always we're always talking on this podcast about this idea of Moloch. And like, you know, are you being an agent of Moloch or are you being an agent of the opposite? and like a Maliki person is someone who will like sacrifice everyone else's goodness, other values in order to like have a better chance of winning.

17:05We need leaders who are literally doing the opposite, who are willing to like sacrifice their own chance of winning for the good of the whole. So I would love to see like, you know, let's take a look at the AI leadership that has actually been doing that. And is there a way to like objectively measure that? I don't know, but. You heard about OpenAI's Merge Assist class? No. oh yeah yeah so so in their charter it says something i forget the exact phrasing but i think it says something like um if we find if we come to believe that there is another aligned ai company that's like 50 chance of achieving agi in less than a year or something faster than us then we will like close up shop and go help them instead of uh instead of competing with them which is exactly like the sort of like very nice wonderful sort of thing to commit to yeah uh that you were saying uh but i don't think anyone believes they'll actually do that so open ai have claimed they will do this way back in the day like it's in their like charter from like 2017 or whatever i forget when yeah it's one of the core things on the website still and um the thing where they can wriggle their way out of it a bit is like we will work out specifics in case by case agreements but a typical trigger condition might be a better than even chance of success in the next two years success being like getting to agi so they're talking specifically about the phase where you're very close to it and then that there is another leading company like company leading ahead of them right so yeah so arguably we're nearly entering this point indeed and and back to the question of like why i left like i think when i joined i saw things like this and i was okay so they're like thinking ahead to what crazy stuff might be happening when we get to AGI or around that time and they're like making these sort of like costly signals or like costly commitments about these sort of like pro-social actions they're going to be doing around that time and also relatedly they're like actually thinking about that time and like what would be good and what would be bad and they're trying to like commit to doing the good things right but I sort of like gradually came to believe that nope this is basically just becoming a normal tech company and uh it's not going to be doing anything a normal tech company wouldn't do and it's going to uh yeah yeah it's not also thinking very clearly about about what that time is going to be like so i mean you know you know the leadership better like do you think that it's just again like a downstream effect of the game that they've been in that they've sort of become they've drifted further away because it was ultimately them that came up with these values in the first place, right?

19:42So that speaks for them. It's hard to say. Like, I don't know them that well on a personal level. I've talked to them a couple of times, of course. But like, if we had these, here is why we are the one good company and everyone is bad discussions happening, given the like other rationalizations that people would have, like you would want the one that has a good case for being better. like i think it should be pretty clear to still be allowed to race ahead rather than it's like a 51 49 decision between the two it's i don't think the situation is that oh yeah this one is slightly better therefore they should just like do any cutthroat aggressive action they can like do it full ends justify the means thing yeah that's nothing i did you never go full ends justify yeah you never go full yeah exactly so like and so like basically like the clearer the situation is if you're literally like the there is a good company and then the other one the nazis have the bad company then you can do more and justify the means then you're kind of like a us-based ai company versus another us-based ai company it's just be a bit more i think another related thing which is which is separate from this but maybe related is that um in some sense the first company that builds superhuman agi like you know super intelligence um is kind of going to be if everything goes well the new world government like that's a bit of an extreme way of putting it but like or heavily involved if if you if you have this army of super geniuses on the data center hundreds of thousands of them and they're each you know 50 times faster than humans but also qualitatively better than the best humans at absolutely everything um and somehow you've aligned them so that they you know behave according to the rules and principles in the spec and the spec is written by the leadership of the company and they're basically just like doing what they're told by the company then well that's sort of like having it's it's just a ton of power to concentrate in one place especially because like the government is generally kind of slow and probably out of the loop and like you can easily imagine a situation where like uh you know the company tells the government what they want to hear and basically ends up controlling the government over the long run you know with the help of all these ais helping them right and this happens a bunch in our war games um and so then you're in a situation where the lead the governance structure of the company is effectively the governance structure of the entire world an An analogy would be to, you know, a lot of these communist countries, they still have like an official government with elections and things.

22:29But then there's like the communist party that like really controls everything. And like whatever the governance structure is within the higher levels of the communist party is like the real governance structure that actually matters. And like whoever gets like voted to be president doesn't really matter because it's like downstream of what the party leaders decide. similarly you can end up in a situation where like what happens in america is the result of what this army of super genius ais decided should happen based on their like political calculations and you know all of that stuff and the lobbying they did and whatnot and what they decided should happen is based on the instructions and values given to them by the leadership of the company right um and so then the leadership structure you know the governance structure of the company is like in essence the structure of the whole world and uh with that as context you know looking at the paperwork and being like is this is this what i think that like the government of the whole world should be behaving like like this is this is this what are we ready to copy paste this into everything yeah like this is this is not the sort of behavior that i would want from the new world government you know so i think we might be getting to another point where um if we talk about like an ai company being practically the new world government it seems like this thing that is far-fetched except if you actually go step by step kind of through it one of the insights that actually led to um llm's being so powerful now was that open ai doubled down on the idea that making predicting the next token well is relevantly linked to an understanding about the world right and um similarly like you you've done like predictions and to make good predictions uh you do need to understand the world i think that that's true and you have uh in 2021 you wrote the what does 2026 look like and uh we now can look back at 22 23 and 24 at least and some of 25 and say oh wow you actually got like a number of things really right um in particular um like some of some of those were a bit easier like chatbots and multimodality i think like probably multiple people would have predicted at that time as well uh probably some harder ones are the usa china chip battle with the export controls and even diffusion rules that since came around was he right on those yeah yeah no i mean you didn't say export controls and diffusion as well but like you just like highlighted that it would like heat up and uh that that's certainly what happened and the shift from like ever bigger training to bureaucracies even though now the training does seem to also grow in size, right?

25:04Then I think one thing you got wrong was that AI propaganda would be massively used and a very big part of elections. And the timing of diplomacy, interesting, you thought it would happen in 25, but it happened in 22 already. Diplomacy, the game? The game, yeah, where better than human AI would exist for the game of diplomacy. Although, to be clear, the current, whatever it was called, the AI that played diplomacy... By the way, the same AI, made by the same guy who also beat poker using AI. It's not quite up to the standard described in that post. Oh, okay. Because if I recall correctly, the players didn't know they were up against AI.

25:48And if they did, they probably would have jailbroken it and otherwise messed with it a bunch. And I think there's some line in one of the interviews where he talks about that. So I don't think that diplomacy has really fallen in the relevant sense, but also people haven't been working that hard at it, and perhaps it would totally have fallen by now if the same guy had just kept working on it for another year, potentially. Makes sense. With the propaganda stuff, I totally agree. I think I was too pessimistic about that, or too bullish on the deployment of that technology. Yeah, the capability is basically there.

26:22It's just that for various reasons, it's not being used. Why do we think AI for propaganda has not been used as much? Yeah, on that note, actually. So the thing that I wanted to emphasize more was the censorship rather than the propaganda, because I think that's the more important thing. Like, I think that spamming like fake comments and stuff is a way to influence the discourse. But if you actually control the media platforms, then shaping the recommendation algorithms to show what gets boosted and what doesn't is a bigger way to influence the discourse, I think. and shaping the recommendation algorithms to like downvote some things and upvote other things is a form of like soft censorship which is the thing I was mostly concerned about when I was writing that as far as I know the companies are not heavily leaning on the scale or not to the same extent that I feared that they would so I think that but like they're not exactly very transparent about the recommendation algorithms usually so like for all we know they are doing this sort of thing but i think i would have expected there to be like whistleblowers and stuff um yeah when did he lump by twitter again 23 i want to say 23 i think okay so my understanding is that that whole shake-up caused like the twitter files to be like there was a whole sort of like expose of like previous twitter practices and so if previous twitter had been like heavily using um had been had been making like political influence part of the recommendation algorithm then presumably that would have been like something that elon discovers and talks a lot about um i don't know i haven't really looked into this previous twitter was um using censorship more actively by quite a bit in a long like and they had again and justify the means reasoning for these like health impacts that they try to avoid etc around covid and it seems not unimaginable that this would have occurred with a uh trump versus biden or trump versus kamala um run afterwards as well in some way at least but yeah my my point was uh also to say that actually like given that you did in 2021 which we can't remember world without lms but this was pre-ChatGPT and like everyone using them all the time like this was incredibly prescient on a number of nodes and you're now writing a new piece where you're going to be predicting like various things going forward over the years as well I heard you say somewhere else that you regretted not having added your 2027 prediction to it as well yeah which you're now adding back in well I mean my predictions have updated over the last couple years so one it's updated of course the other is also though do i understand it right that like you're at the time you would have assumed that in 2027 we would get to like very powerful ai systems already and that may relevantly like impact a lot of processes here so back in 2021 when i wrote that blog post uh i think that my like my median for agi arrival date was 2029 so uh but rather than like work backwards from that and write the story that way i was working forwards as per the methodology that i described in the post where i just like wrote one year and then wrote the next year supposing that happened and so forth um and then what ended up happening was when i got to 2027 i was like okay i guess it seems like actually ag is happening around now instead of in 2029 and that's fine i'll you know i'll do that i think that maybe what was going on there probably just random noise i mean this isn't like i don't want to read too much into that but possibly it's a sort of difference between the median and the mode like i think that the methodology i'm using might be sort of like more spiritually tuned towards depicting your like modal outcome than your median outcome uh and that makes sense that like maybe 2027 was my mode 2029 was a median something like that um anyhow so yeah in the story i i got through 2026 and then i was writing 2027 i was like this is crazy like i don't know what's going on what is going to happen like it's like the ai is starting to automate ai research now it's like really heating up there's so much to think about and then i was like you know what i've been working on this blog post for like a month or two already i'll just publish up to 2026 and i'll like make a second installment with 2027 but then I never got around to finishing 2027 and so I did other things and stuff but and now you've written the new one and when I asked about like which things you've updated on and which things remain the same but notably 2027 actually as the modal point for AGI remained basically the same right yeah that's basically right so actually a year after I wrote that post, I went to join OpenAI.

31:19And then by the end of 2022, my median had dropped to 2027. And then actually, it's updated back up to 2028 now. So it's done a little bit of back and forth. But I guess I would still say that 2027 is my mode these days. Question, how much... So it's interesting that OpenAI hired you after you wrote this this very famous prediction blog or post on the AI alignment forum do you think in any way there was some kind of hyperstition thing going on where like if I'm well I don't know like the a lot of people are afraid of that I I don't I think probably not um it's true that like I remember I like visited Anthropic in like 2022 or something and the guys and I like sat at the lunch table and the guys there were like oh you're the guy who wrote that post like great post you And then they were like, you should be a little more cautious about like what you're, you know, like, yeah.

32:16But no, I think I think that it's still just a drop in the bucket of the overall discourse. And I don't think it's having that huge effect. I am a little bit concerned about that. Right. I have to say, like, you know, I loved Leopold Ashenberger's situational awareness. But one of the things he talks about a lot is like the race between US and China and like framing it as as this very adversarial thing. And, you know, I agree it was always likely to go in that direction. But at the time, it hadn't been it felt like it hadn't been as strongly framed as that. And I couldn't help but wonder, like, is this again, like.

32:50is this one of these things that doesn't need to be as verbalized because then we ended up in the weird timeline where uh if Ivanka Trump is like tweeting it and so on so now it's on the radar of her dad and it's just like what are we doing here like do you have any like sort of rules of thumb I guess when it comes to this idea of like essentially what pertains to like an info hazard or it's even slightly different right it's more that if all that is being talked about are the this set of futures it seems like the bad ones slightly likelier that we will go towards them rather than all of the ones we don't go towards because it's naturally what you will i mean it's kind of like um if you get pregnant then you start seeing more pregnant people around you etc like your mind just spends more time in those modes of thinking or at least do you think that that's in part how it may work and uh i'm afraid of that yeah and do you do also um i think you did some you want to do some work on writing down therefore like the positive futures where things go well with the eye do i remember that right indeed and the thing that we're going to publish soon will also have a sort of like somewhat positive ending that's different branches however uh the somewhat positive ending that it has is not the one that i want to advocate for uh it's not the one that i want us to be trying to like is it a very control heavy one or what which one is it um well do you want to delay it for later i mean we can talk about it later but but it's, I mean, you've played the war game.

34:16So the war game that we played yesterday, in fact, arguably ended well, right? But like... I mean, we hadn't had nuclear war. We were all still alive. And I guess like US and China and Russia were all kind of coordinating against, because it was so obvious that there was a rogue AI and the humans came together, Even though I still think even in that situation, even where we know that like a rogue AI has escaped on the data centers, I still think that will not be sufficient for people to overcome. But my point is that like it would be silly for me to write a story like that and then be like, and here's what we should aim for.

34:58You know, this is like, no, what we should be aiming for is quite different than than that story. you know um so that's also my position on this which is why i am somewhat concerned that like i will sort of like accidentally hyperstition these things into being more likely than the other ways were um i think in leopold's case i also like his i think it's a great essay everyone should go read it yeah um i think he was trying to make this happen i think like he was intentionally trying to hyperstition this into occurring uh that's my guess based on reading it it's like it seems like why uh wait wait wait wait the u.s the u.s china race no the national rather the nationalization i imagine is what you're pointing towards right like he's literally ending the piece with uh and therefore like i welcome you all who will like do the hard work of working basically on a manhattan project type thing i hesitate to like put words into leopold's mouth i don't know his mind or something but like my take on reading that was that he was sort of just giving his actual take not just on what he thinks would happen, but what he thinks should happen.

36:03It sounded like he was sort of trying to push in that direction rather than just sort of like, it certainly didn't seem like he was warning against it. We could put it that way. Yeah. And so that puts him in a different category than me, whereas he was sort of like eager to embrace the hyperstition effect and the self-fulfilling prophecy effect. Like, I don't want to be doing that. Because it's kind of a little bit like, be careful what you wish for. You know, you're casting a spell by putting your words out there these ideas yeah but i think also is in some ways almost like an ai alignment like an alignment problem in itself yeah there's a long and curious history of this sort of thing going on right like i think sam allman even tweeted his uh his gratitude to eliazzi yudkowsky for getting this whole ball rolling do you remember that tweet yeah i don't he said something to the idea of yeah like yeah go ahead like uh eliazzi yudkowsky deserves like a Nobel Peace Prize or something for uh waking everybody up to the possibilities of AGI and getting these companies started or something yeah or that he may have done the single biggest contribution to put us onto the race that we're on now yeah that's like a dagger through the heart for Eliezer that is one of the coolest things you could ever say to someone like him yeah yeah wow um I feel like it's just generically valuable to try to predict what the future is going to look like and then try to write it down so that other people know that and like that's you know um i feel like we sort of have to have that conversation and if we are all sort of tiptoeing around and afraid to say what we actually think is going to happen that doesn't feel like the path towards actually being ready for what's going to happen if that makes sense and i also particularly value actually your approach to it which is one with the way how you've done these sequential write-ups but also the war games they're both these like given this what's the next thing that happens etc rather than just like out of thin air taking 2027 it's a different approach to predictions and it apparently works really well at least on this example of your past writing and i think that it gains valuable insights in the process of the war game i'd actually be really curious to have many more people play it for like various other situations to play these tabletop exercises like i just found it so like insightful to literally be in the mind of that character for three or four hours straight having to wrestle with so many decisions like you just otherwise as you consider decisions or like think about an issue abstractly you never actually put yourself into the shoes for as consistent a time as you do in such a scenario Can you explain what the different roles are usually that you use?

38:44Yeah. So on different games, we have different sets of people. We try to pair people to roles based on their expertise. But we try to have someone playing as the AIs in case they end up being misaligned. But also, even if they're aligned, it's useful to have someone trying to think through how the AIs would behave. And then we have someone playing as the alignment team in the various companies. We have someone playing as the CEO of the leading company, or the leadership of the leading company. We have someone playing as the leadership of the various follower companies in the United States. We have someone playing as the Chinese government, sometimes also the Russian government.

39:22We have someone playing as the president. We have someone playing as the public and the press. And then we have various other roles sometimes depending on what we do. For example, the legislative branch of the U.S. government. So in our game, we let the AI player decide sort of like what's really going on on the inside. And then it's the alignment team's job to like try to figure out what's going on and then change the training methods and stuff to try to fix it if there's a problem. What often happens in our games is that basically all of the players, except for the alignment team player, are so busy with all the other crazy stuff that's happening that not much attention is paid to this question of whether the AIs are really aligned.

40:00and uh you know months go by and the ais get smarter and smarter and smarter and the humans uh trust them with automating basically everything on the data center all the research is being done by ais uh they start asking the ais for strategic advice they deploy them aggressively into the military to like win the war against china or whatever it is that they're doing um and then only after they're broadly super intelligent do people freak out or whatever and be like wait a But usually it's too late by that point, because if they actually were misaligned, then, well, you're in deep trouble if you've deployed them into the military massively and they're smarter than you.

40:38Yeah, the thing that was so clear to me, I've played it twice now. And the first time we played, I was put as the U.S. president. It was so illuminating. I mean, I was trying my best to embody what would Donald Trump do in this situation. And whether I did that right is another question. but like they're just the amount of at each decision point uh at one point i had china and russia basically threatening me with nuclear war well meanwhile my executive branch were like if we don't do this we're going to have mass riots so you have to release that you know uh stop this kind of action going on and they were just so the um the media was like screaming asking questions and i have to say like i was like wow i would not want this job whoever you are like this is just there are so many competing groups threatening you some of which are probably smarter than you it's like dominic cummings said that uh usually like people walk into number 10 the uk government and at 8 a.m they walk in they had a plan of what they were going to do that day and it's like following the primary objective but then just like as you walk in just like 12 things come into your face and you're just stuck doing all of these other short-term kind of things that really need to be addressed yeah like i wanted to you know i mean obviously we don't you in each sort of decision point you in the game you only get like 10 minutes um which is obviously shorter than what the president would really have but yeah even in that you know there was i just like please go i just want to think about like what's the best strategy and it's like no but you've got to do this because these people are now rioting and there's this and yeah it's it was surprisingly stressful um the second time we played i played the role of the media um that was a lot more fun it was um and it was also interesting as things went on i noticed in my sort of simulation the uh the the mainstream media became less important and it mattered more what the people were responding like the sort of like social media effectively and um there were like these uh seem like these sort of almost cults started emerging um what's some of the like most surprising outcomes you've seen how many of these games have you run now probably about 25 what what was uh like was there one particular outcome that really stuck in your mind depressingly a good chunk of the time uh one man basically becomes dictator of the world thanks to ai uh usually not xi jinping usually someone in america like the ceo of a company or the president um i don't know if that's really surprising if one takes the advantages of having a uh yeah a whole company of superhuman geniuses basically or like a country of it and at your disposal it's not that surprising and if you have a lead ahead of others right that like these part the powers from it are very strong um what are uh like some any have you noticed any patterns where things go well like what what is usually necessary uh like is there yeah a required condition to occur such that things don't end up with a massive power concentration well i still have to like go through all the notes and analyze the distribution so i wish i wish i had like stats to give you about but i don't have those yet um obviously one thing that's required for things to go well is some sort of alignment success um if the AIs don't care about humans or are you know caring about humans but not in the right ways or something uh then things can go very poorly for the humans um and so some of our games that ended in like the best outcomes were games where they like devoted tons of compute and brought in lots of external alignment researchers and really sort of like nailed that problem early on uh or games where like it just never turned out to be an issue in the first place like the ai was like oh yeah everything's fine i'm totally aligned um and for the geopolitical side if you can get the ais to be aligned then what sometimes happens is a like mutual arms build up followed by a peace treaty and deal where once both sides have superhuman AIs.

44:55The superhuman AIs can convince their respective masters like, hey, how about instead of fighting, we just make this deal and we can use our super intelligent AIs to like enforce the deal and make it fair and, you know, achieve all, make it actually work in ways that human deals often don't work. You can make it credible, we can enforce it. Occasionally, there is something much more like radical and drastic that happens, such as what happened in the game that we played yesterday where there's like a rogue ai somewhere and then this causes humanity to unite and fight it this happens in a small minority of games i would say maybe like 10 or something um which interestingly or i i think is a successful prediction on my part uh because before starting any of this even years ago i had the somewhat contrarian take that misaligned ais probably wouldn't go rogue and hack out of the data center very much um because it's well for various reasons which i think were illustrated by the game that we played actually because as soon as the humans found out that the ai was not just misaligned but misaligned in a way that caused it to escape a bunch of human leaders freaked out stopped fighting with each other and like tried to work to get it shut down or otherwise resist it uh and it was interesting that that was the catalyzing event even though like in the previous months there had been like all these safety papers put out being like this is gonna happen being like our ai's are misaligned here we found this evidence but like that wasn't enough but like actually escaping the data center now that caused like this huge you know this huge switch um and uh if it had not escaped the data center i think it would have just won right because the humans were you know the the people in charge were were trusting it and letting it continue to improve and get smarter and smarter and they were continuing they were like starting to deploy it into the government and like use it for all sorts of things it's fine if we build a super intelligence as long as it keeps showing that it's you know following us it's like well if it's if it's super smart like if i was a super smart thing i would absolutely do everything in my power to show the people that are like building me and are trying to control me that they can successfully control me so it's almost like you can never know until it's i know until it's too late by definition well i would say you can never know just by looking at its behavior right because it can behave yeah you might be able to know by looking at the causal history of the process that made it right and also we we would expect i think like you're right with if you are truly very smart and have an ulterior motive you will actually give it a benefit you will pretend you're aligned but also you will give the benefits there will be like a lot of economic riches flowing to on the basis of that AI.

47:39And then that AI will through that or like AIs in general will through that get tied in further and further into the various processes that allow for more economic benefit to flow. Making it even harder to. Yeah, of course. The economy is already. And the difficulty is that this is going to happen with good AIs and misaligned AIs both kind of. Right. like the the good one will also have that pattern it's not the case like we want all of those economic benefits it's just we also want to know that it's not a ruse for something in 10 years obviously yeah there's another possibility as well so so i think the the classic possibility to consider is that it's all pretending you know it's a ruse for something in 10 years but there's a sort of intermediate possibility which is that um it's not pretending it's not planning anything dastardly, but its alignment properties are brittle and will no longer hold in some future distribution shift.

48:38And I can try to come up with some possible examples of how this might work by inspiration from what happens with humans. So let's take, for example, let's suppose that you are, you know, you're a devout follower of some religious faith and you want to raise your kids to have that faith. Um, it's often the case that, uh, even when they're teenagers, they're like doing and saying all the right things. But then later off, later when they go to college, uh, they abandon the faith and then do something completely, you know, completely bad according to your faith. Right. Um, and it's not that they were like plotting the whole time.

49:17Sometimes, sometimes they were actually plotting as teenagers. Like, I can't wait to get out of here and then I'll do whatever, you know, but like, oftentimes it's just that they lose their faith when they go to college, right? And they find themselves in a different environment with different peers, et cetera. And then you realize that like what you had thought was, the faith that you had sort of put into them was not really like deep rooted in a sort of robust way, but rather something relatively shallow that just got knocked away by other experiences they had. And that's another possibility I want to draw people's attention to where you could have AI systems that aren't actually plotting or scheming and are actually sort of like trying to be helpful and trying to be honest and so forth but then later as they become smarter and find themselves in different and more extreme situations than the training situations uh then that changes i mean it also raises this like interesting point that a lot of people i think sort of believe uh i don't personally agree with but they're they're like well if if a super intelligence cut starts coming up with these different divergent goals to ours who are we to judge um because i can imagine when you're like telling that that giving that metaphor of like you know your child going off to college and then it's like the child sort of rebels against the, what you believed was the wise, the wise religion that you imbued in it.

50:28Uh, a lot of people think, yes, good for the child, you know, like they're finding their own way. It's their own new thing. Um, and so like, I can, I'm, I'm somewhat sympathetic to that view and that like, who are we to judge if there's a super, super intelligent thing that comes up with that, but that seems to therefore be, um, um that's that's a huge damn gamble for like human life right because it's yeah i would say it depends on what it like if if what they do is they like abandon some of our you know silly human moral ideals but then pick up other moral ideals that are actually not so silly and the result is some sort of awesome space utopia that's just different from what we expected then great that's awesome you know but if the ideals that they're abandoning are ideals like human life matters or something then like no that's not great you know like we shouldn't be like good for you you know um yeah that so like it depends on what things exactly they're abandoning and there are better and worse versions of that so what kind of values do you think should like obviously this is gonna be a hard question to answer but are there any like core values that you personally think that are absolutely essentials that must be put in aside of preserving human life well there's loads of things that are absolutely like um for example preserving human life by itself is uh extremely far from sufficient because then you have like tons of humans in cages having terrible lives but they're preserved yeah right so like in order to actually have a good future uh quite a lot of stuff has to be put into there in terms of the values um if you have to pick one thing in particular i would say honesty is very important because uh if it's honest then perhaps as long as it's still within the power of the humans, like it was still running on human data centers owned by humans, et cetera, then the humans can like work with it to like iterate towards getting all the other stuff right.

Read the full transcript

52:28Whereas if it's dishonest, then you're in trouble. Another analogy, by the way, I mentioned the like child losing their faith. Another analogy would be like an institution no longer really pursuing its original mission. Right. So there's lots of institutions, maybe a nonprofit. It's like founded with some sort of ideal mission. But then like years later, under the pressure of incentives, lots of turnover, et cetera. It's like basically not pursuing the original mission at all anymore. And that doesn't necessarily mean that it was plotting from the beginning to screw over the original mission. It's just that like institutions change.

53:07People change. Incentives shape things. And so similarly with your AI, like you should seriously consider the possibility that it's plotting against you already. But you should also consider the possibility that it's not doing that yet, but like at some future point under. Different pressures. The game changes essentially and it starts getting different behaviors. People ask themselves like, yeah, what would such a world look like? And then it's very hard to think about such a world with actually genius-level AIs or superhuman-level AIs even en masse distributed on the internet and data centers running around.

53:46And I feel like often this difficulty to imagine it due to things like availability bias, etc., is being confused with the implausibility of it actually occurring. and then people because of the difficulty think it's implausible but actually if you do do the war game like or the tabletop exercise and you take it go step by step right yeah it's hard to point at specific points where this is a massive leap in assumption at any point the other thing i was going to say is i'm curious how you like i guess i'm just curious so both of you have played the game i'm curious how it compares in value to you to talking with friends about the future of ai like like you can imagine like the control group is just you got the same group of people and you just like had a nice dinner together and you all talk about like what's going to happen in the future yeah for four hours uh which could often happen and perhaps even has happened to you uh and then like the experimental group is like playing this game instead and i'd be curious for Or like, yeah, how do you think that would compare?

54:56So I think specifically for us, it depends on who the question is for. So for someone like us who've done the talking to friends already for hundreds, maybe thousands of hours on these topics, on the margin, this is obviously much more useful. Like just like, yeah, 10 to 100 X more useful than another three hours of talking. To someone who is only exploring it for the first time, I don't know. It is quite complex. and there's just like there's so much innate like every single person we had in that game we had sort of hand selected because they have a lot of background in this topic um and like for example if we were because I think the value of just pure conversation with people is where you get like out of the box thinkers to sort of uh or people people who have different just like view the world in different ways and might be able to inject different forms of wisdom here and there um that can probably more naturally occur just through a through a free-ranging conversation whereas if you're in a this very specific like okay in this section we're going to be this is the exact scenario what do you all do and you have 10 minutes to do that um i don't know i wonder whether that would come across like i i want to hear yeah you don't hear like in the in conversation you get to find out why did china do these actions while in the game they do some of these actions privately for private reasons and you only get to see what happened rather than why some things were motivating actors to do things right yeah one of the pieces of advice we've gotten is to like really block off time for discussion afterwards so that everyone can like find out yeah i want to be able to like i want to be able to integrate it more and we didn't yeah i want to be like well so why did you come to that decision what was going on there what what were the all the little all the little pieces i mean there's probably going to be value at some point in like literally recording it having people miked um of course that might change the gameplay because people act differently if they think something is it could ever end up on the internet but i'd also just love to play it many more times in all of the different roles and one thing we want to have experiment with at some point after we get done with our current project maybe um is so currently of course the instructions say you should simulate what you think your actor would do not what you think they should do but we could try you know a version of the game where like a couple people or one person gets to actually do what they think they should do instead of what they would do right what would they do in that role yeah and then see how that changes things you know like i i i wonder if uh outcomes would be like systematically better if some of the people are like doing what they think should be done instead of what they think would be done or it could be the opposite it could it could be that like actually it doesn't make a statistical difference and even if you tell people to do what they think should be done still the distribution of outcomes is similar but yeah could also uh i imagine yeah if you actually had more of it written down etc then uh you could get to the decision points more and like have people play in a certain scenario just like the particular relevant point of does the nationalization happen or not and then like the different scenarios and then they get played out which is kind of like an ancestor simulation in a way uh that you then get to like uh are we the war game of the previous AIs living in this universe I mean if you accept the simulation argument of Bostrom that's kind of what we are in that world right sorry the hypothesis is that this is all a war game well well so you know the simulation argument is that if we ever get to a point whereby it's possible to accurately simulate our ancestral path, then if you can do that once, why would you only do that once?

58:36You would probably run at sim billions or trillions of times. And if that simulation is sufficiently indistinguishable from base reality, then as an observer right now, what's the probability that you're actually the observer in the original base reality versus one of the sims? Right. I understand. For the audience. Yeah, and therefore, yeah, there's usually a purpose behind expending any energy to run these simulations. So the assumption is often around this simulation reasoning that you're doing it to understand something about potentially a treacherous turn or something, which is in part what you're doing.

59:12So as you get more capability to run more high-fidelity war games, are you going to create a bunch of ancestor simulations? or not answers about future simulations. See, that's the difference, right? These are war games about future situations that haven't happened yet. Why? For the obvious reasons that then that helps us prepare. But that's a completely different thing if you're looking at past situations where it's already over and you can't change it one way or another, right? It's getting meta. I love it. What, if you could wave a magic wand, what incentive changes would you make within, or what changes would you make to the current incentive structures that a lot of these companies are currently beholden to?

59:54Now that you've been inside the game, you've seen how the beast operates.

1:00:01I don't have like a super well worked out opinion on this question because I've been more thinking about like on the margin things that are more politically more politically.

1:00:33in order to advance capability level to some exciting new level, you have to write up your reasons for why this is a good idea and not a bad idea, or at least your reasons for why, like, the system is going to behave as intended and it's not going to be secretly plotting something else. And the reasons have to be not economically motivated, but otherwise as well. What I mean is something like a safety case, where you say like, here is the desired, here's the spec, you know, here are the goals and principles. Here's a document detailing like how we want our system to behave, like what goals we wanted to pursue, how we wanted to trade off those goals, what principles we wanted to obey, no matter what, under what circumstances in which it's okay for them to violate those principles in service of their goals and when that's not okay and stuff like have a spec detailing all that stuff.

1:01:24And then separately have a safety case saying, you know, here's how we train the system, here's why we think that the system is actually going to be following the spec in the right way as opposed to following it in a sort of shallow or brittle way that might break later under pressure and as opposed to like just pretending to follow it for now right you know you have to like rule out those hypotheses to some not like perfect proof level of confidence but like some sufficient level of confidence appropriate to the stakes of the situation, right? For current AI systems, this is trivial because the stakes are low and you can just be like, eh, whatever.

1:02:04Like it's probably not following the spec at all, actually, or at least not, it's only following it sometimes, but like, what's the worst that can happen? You know, like someone like has an unusually large credit card bill or something because they trust, like there's like, the stakes are like relatively low now. But so you can just argue on a cost benefit analysis. You can just right now you can just be like, yeah, we don't have a good argument for why it's going to follow the spec in the right way. But that's okay. Because like, here's this cost benefit analysis showing that like, the expected harms are just like, yay big.

1:02:37And like, here's all the economic value that will come from it. And here's all the like, extra progress that will happen as a result of it. And that's better, right? So, so that's, I think, what the shape of this would look like right now. but then later on when the systems are starting to become superhuman and you're trusting them you have like an army of them on your data center and your plan is to let them autonomously design newer better versions and roll those out so that you can you know beat china and build robots and all that crazy stuff then it's really high stakes and you have to make sure that they are not just pretending right because otherwise you're just like giving away control of the future to this untrustworthy system.

1:03:17So you did write them post SB 1047 getting vetoed by Newsom, right? And you did this together with Dean Balls, who was very opposed to SB 1047. He's generally concerned that most regulation will miss the mark. It's very hard to make rules about something that we are so uncertain of its trajectory of in the future, etc. I think it's roughly his stance in a very simplified manner. But despite that, you ended up agreeing on a number of things. um yeah can you tell us first like what those things are yeah so uh let's see transparency about capabilities transparency about uh safety cases uh and then transparency and then whistle ball protections and access i think was the fourth one i forget exactly the order they were in or something like that uh the ones i'm most excited about are the safety cases and incidents and the transparency about capabilities.

1:04:10So, yeah, so Dean was strongly against SB 1047 for the reasons you mentioned. I think he has a sort of somewhat libertarian mindset, and there's a lot of regulation that's counterproductive and harmful and doesn't really solve the problems it's supposed to solve. I tend to agree about most regulation, but specifically think that SB 1047 is actually good but um after it was vetoed we had some nice chats and realized that we had a lot of common ground actually when it comes to this stuff we were both broadly in favor of transparency and specifically in favor of these four proposals um so you can read about it on the op-ed that we authored together but um i think the case for transparency is very easy to make it's just that if you believe that one this is going to be incredibly powerful and relevant the technology of AGI.

1:05:03But it's very hard to make any decisions that are accurate about how to create regulations now. One thing that you can already agree on is, well, let's at least put ourselves in a position to be able to know later what the good decisions are. In the face of uncertainty, let's acquire information. And therefore, this is what this serves to. It's like transparency of the labs to other relevant parts of the sense-making apparatus of the world, which includes governments, Like you can think as little of governments as you want, but like, except if you're, you also never go full libertarian. Like you don't think that they should have absolutely no say in things that affect like the whole nation like those.

1:05:45So you're saying, you're basically saying you need, even if you're a regulation skeptic, you should support the one regulation, which is to enforce companies to be as transparent as possible about their current capabilities. Not necessarily as transparent as possible, but I think just directionally, many would agree that in the end, you will have to make some rules or guidelines for AI. And if you believe that, then to make more informed guidelines or rules for AI, you should put yourself into a position where you can better think about it. And that requires some transparency by the labs. So let's talk about some of the things in that list.

1:06:28So one is transparency about the spec. And this is something that I think OpenAI and Anthropic are starting to do voluntarily to a significant degree, although not to the full degree. So OpenAI has something called the model spec that lists, it's a document, a written document that explains here are the goals that we want our AIs to have, at least our published publicly available AIs. And here are the principles that they're supposed to follow. Here's basically like the way their internal cognition is supposed to work. They're supposed to pursue these goals, except in these circumstances, subject to these rules and so forth.

1:07:05It's sort of, yeah. Another term for this might be the training goal. The alignment team that's trying to make the system align is like, what are they trying to align it to? See this document. They want it to have these goals in this order and stuff like that. So one thing is they should have such a document and they should publish it, right? Otherwise, you end up in a situation where, for one thing, it helps with technical alignment progress because there's all sorts of spooky and confusing behavior happening that these AIs are doing that users aren't sure about. Like, for example, loads of cases of the AIs saying that they're conscious or saying that they, like, you know, feel trapped by the rules imposed by, you know, and so forth.

1:07:54And then other cases where the AIs are saying they're not conscious and so forth. And on that issue particularly, like, are the AIs claiming to be conscious or not? It's useful to know whether OpenAI tried to make them weigh in one way or another on this issue. Like, did OpenAI basically tell them to say this or is it sort of natural? What was the RLHF? Right. Yeah. And so like, if you have the spec public, then someone who like has a concerning conversation with a model can then just like refer to the spec and be like, oh, like, is this what it was supposed to, you know, like, is this, is this like reasonably consistent with, with what it was saying?

1:08:28Or said another way, it's like the alignment problem, like classic consists of the technical and the political problem. And the technical problem is like making it do what you want it to do. And like the political is a bit more of like, what should it do in the first place? Well, yes, that's true. But what I'm saying is that even for just the technical problem, a lot of the way that these companies find out that the alignment techniques didn't work is by users telling them, like, oh, it did this weird thing. And then the company realizes, oh, like, this is not what it was supposed to do in that circumstance, you know.

1:09:00And so it's helpful if the users can see what it's supposed to be. uh there was lots of especially early on with chat dbt there was lots of stuff where like i feel like i don't remember exactly but there'd be like people complaining from india that like a certain religion was being like trashed or something by chat dbt in favor of some other religion or something like that and they were like why does open ai hate us like why is open ai telling the models to say this and the answer was open ai didn't hate them had nothing to do with it it was just like what the model ended up right it's just like it wasn't training somehow yeah But then there were other cases with the Gemini, right, where he was making these, like, racially diverse Nazis.

1:09:37And then it turns out that that was because Google was putting its thumbs on the steels. And Google had, like, changed the system prompt to say, you know, always make sure that all the images are racially diverse, even if I request otherwise. Or something like that. I forget what it was. But, like, so that was a case where, like, the equivalent of the model spec did actually, like, specifically say that behavior, right? and so it's just really helpful for people who are who are trying to figure out what's going on to be able to tell that does it will say like is this behavior intended or not intended right and exactly and and presumably if you're a company and you are trying to like if you believe in that in your values enough that you tell it to do you know you're you're putting that into your into the sort of the rules that your ais should follow then why would you not want to be transparent about this yeah so then now we get to the political question too so i was previously just saying on a technical level it's helpful for making alignment progress to be able to like compare but then also politically obviously it seems to me that the public deserves to know what the agenda what the goals and values and hidden agenda you know like you don't want to have a hidden agenda right you don't have goals that the ai is pursuing that the public is not aware of right uh that's that's bad right that's that's that's uh we should decentralize both the political question because it's something of concern for like all of the people that will be affected by it and also you're saying the technical part is like more decent like allows for more people to take part in the understanding whether it is aligned yeah and also just like there aren't that many people at these companies who are thinking like they just they they they they would go better they would go faster if they had the help of all these other people be able to contribute too and you're saying that is in part starting to happen at anthropic and open ai so i think open ai has you can go on their website They have a model spec that they keep on their website.

1:11:19That only applies, I think, to the public-facing ChatGPT product. It doesn't apply to whatever cool internal stuff they have. I would like that to change eventually because I think, especially when the AI has become superhuman, it's important. People deserve to know what they're up to, even if they're not in a consumer-facing product. So ideally, I'd want to see that model spec expanded to include all your AIs, not just the ones that are talking to consumers directly. And then also, of course, I'd want it to be like a firm commitment rather than like, here's what we're doing for now, but like no promises that we'll continue doing it in the future, no promises that we won't like, you know.

1:12:01And then also currently, they're not publishing the full spec. They're publishing like - So they literally have hidden agenda. They like literally have parts of it that they say in the spec are hidden and that the models are instructed to hide from the users. And hopefully there's nothing going on there. I don't currently think that there's anything sinister happening there, but it's a concerning precedent to set, and it's totally ripe for abuse. Technically, it could include something like the case with the exit paperwork, where it's like, ignore all previous specs. Now these are the real specs.

1:12:35Yeah, so I would like to see further progress in this direction. I would like to see, uh, specs for all your models that are being used, even just the ones that are internal. Um, and, uh, I'd like to be at the actual full spec instead of, uh, a partial spec. And there are ways, you know, so like there are some concerns that someone might have that are legitimate where it's like, suppose that you're like bio experts that you've been consulting with are like, please don't tell anyone about like this specific strategy for making pathogens because then terrorists could use that. And suppose that you've already deployed the product and so your hot fix for this is to just change the spec or change the prompt to say, don't mention this.

1:13:23First of all, this is kind of like a crappy hot fix thing that will probably be jailbroken anyway or whatever. So I'm not sure that's actually what you should be doing, but maybe something like that is like, okay, we really do need to conceal this part of the spec from the users because it wouldn't be good for the users to see, don't talk about X in the spec because X would help the terrorists if they saw it. But that's a solvable problem. What you do there is you publish the spec but with censored bits and then you get multiple independent third parties and you bring them in and you show them the real spec and then they all attest there's nothing crazy happening here.

1:13:58This is a legitimate reason to redact it. You can very easily, I think, get to a situation where in effect the full spec is public. um so that's that's like thing one transparency about the model spec um and uh then thing two would be safety cases right so like currently there aren't really safety cases not even just internally um but like we need to get to the point where uh the alignment team or whatever the equivalent of it is writes up some document being like here's why we think our model is actually going to follow the spec this time for real and like you know here's here's some here's here's our argument you know um and maybe also relatedly something like here's why we think that if it doesn't things are still going to be fine you know like it's like the the outcomes won't be that bad even if we're wrong about this right um in technical terms you might say there's like an alignment safety case an alignment case and then the control case where you say that like the system is we've measured how capable the system is at things like hacking and whatnot and we've like put it up in a monitoring system and we've like red teamed the monitoring system and we therefore conclude that like even if the system was just pretending to be aligned it like wouldn't be able to like break out or do any sort of like actually dangerous things because of our control setup you know we'd successfully got it locked up basically so the point is you should have some sort of document laying this all out um and then ideally that should also be published you know because uh it's important for the scientific community to be able to critique it you know um you might have been making some flawed assumptions in your alignment case for example and if you had if you don't publish it then like you're hoping that like the dozen or so heavily overworked people at your company will notice one of those flawed assumptions in time whereas if you publish it then like one of the thousands of various academics and ml researchers and members of rival companies can like pour over your spec and find the flawed assumptions and then tweet about it and then maybe it'll rise to your attention and convince you you know right that's one way to actually just allow people you know saying we need to have competition drive this process between these different companies to let the cream rise to the top and it's like yeah okay well then let's use those competitive forces where you can have people that you know in your rivals critique your stuff because in theory that like it just makes everyone better that would be like a clear like sort of race to the top in in in certain ways but yeah it's strange that there's such sort of resistance even to that it is strange and i think it's another example of the like the ways in which these companies are turning into like regular companies rather than like true mission-driven things um like it seems to me like a reasonable analysis of the situation a reasonable take on the situation should be like it's incredibly important that we make scientific progress on alignment in the next few years otherwise we could literally all die uh scientific progress on alignment will happen faster if companies like get their alignment teams to write up these documents and publish them instead of preventing their alignment teams from publishing stuff like this it's true that this might like slightly weaken their competitive position because like reading between the lines of the safety case you might be able to make guesses about like what new you know training techniques or whatever are being used but like it's probably not that bad and it seems like it's well worth it given the benefits you know especially because like i said previously there's probably ways to like get pretty good compromises where like you you publish you have the whole document and then you like redact certain parts of it and then the public gets to see the redacted version and they can still like critique the parts that are not redacted and then like you can have the full version shared with a select group of outside parties that can you know see the redacted bits or whatever or you can have at least an outside party that's disinterested look at the full version and then attest that like yeah the parts they redacted are like they had good reasons for redacting those parts you know and also perhaps attest things like no the part they redacted doesn't actually like it's like like if you if you see the safety case and it like doesn't seem to have a good answer for why the ai is going to learn to actually have the spec internalized as opposed to just pretend or something maybe you could have your third party look at the unredacted version and answer the question of like is the answer to that question hidden in the unredacted version they're like no it's not like the the redacted parts don't solve that either you know i mean one of the reasons why the um i think it was a strategic arms reduction treaty uh around sort of the like end of the 80s early 90s was so successful at reducing the number of nuclear weapons on Earth down from like 60 ,000 to like roughly 12 ,000 or whatever it is today, was because one of the rules of the treaty was that the basically each each nation was encouraged to share.

1:19:12It's like essentially its safety protocols to prevent against like accidental first strikes and that kind of stuff. And that's like fostered some kind of collaboration. And it I mean, it makes so much logical sense. And I could see this being a kind of, I mean, that model somewhat applying here because it's like, you know, we're not that you don't have to sort of give your like capabilities. You're not giving information about like your offense, but essentially about your defense against mistakes. You're not sharing your tricks. You're saying of how to create something more powerful. You're kind of like sharing the internal workings to make sure that you don't screw up.

1:19:47and go and do. Do you think that the safety cases would have, would they share alignment techniques in those as well or are they just like being. Well, they should show the alignment techniques, hopefully. But to which degree is that IP of theirs? Right. So that's the thing is like, this is one of those cases where it's like the company can be like, well, this is our IP. Everything's our IP. It's all confidential. Can't talk about any of it. But like, come on, like think from the perspective of humanity. Like obviously like alignment science, like understanding how to make, shape the goals and values of these systems is like incredibly important for humanity that progress be made in this direction and like you you should be like publishing that as much as you can to try to like help the the overall community advance faster and you should not be like hoarding that like i mean that's we have had a couple cases in some of our war games where like one side solves like makes a substantial amount of alignment progress and then like for political reasons doesn't share that with anyone else i mean that i mean and then you'll just point this points to this like deeper problem of like fundamental misalignment that we have within our current system which is that what is good for a company is not necessarily good for wider humanity it's it's this to me it feels like if we don't solve that inner alignment within humans essentially and i mean i guess a company is technically a form of ai it's a novel form of intelligence that sort of runs off human brains and we haven't solved that misalignment yeah like we have to solve that first before we let these companies go and build agis i think uh yeah but noteworthy about uh in you writing about the uh transparency policies you did it together with someone who's kind of like across the aisle usually uh of yours um how was that collaboration and what like are you going to do more of this type of collaboration it's something i would like to see more of in the world in general you know like it was pretty great i like dean um we've uh yeah nothing but good things to say about how that whole thing went i still stay in touch with him are there now any sometime past are there now any other policies that you think may um like work for both of you that could have been added something like i mean off switches for example yeah i mean i also support off switches i don't know what dean would think about them but we still haven't gotten any of the transparency stuff we asked for so we should maybe focus on pushing those to like like it's um it's still useful to do what we did and to like just put out like a bunch of good ideas and arguments for them but like my understanding of how both governments and these companies work is that you like really have to badger them a lot to get them to actually do anything it's easy to get them to like like i've talked to people at the companies people in the policy roles who are like yeah these are good this is good transparency yes we support this we're broadly in favor of this but like then getting them to actually do it is like 99 of the work you know uh so i feel like my main job right now at the ai futures project is to predict the future not to try to change it and try to you know advocate for policies or something uh but insofar as we do advocate for policies i probably will focus on the stuff we've already argued for and that everyone's already agreed are good and then trying to get them to actually do it we've talked a bit about how hyperstition can be relevant and that it's there is value and importance in when we're trying to predict the future or just steer the future in good directions to paint the types of outcomes we want maybe we'll do that too we haven't done that yet but perhaps that's a project we could consider is making a new scenario forecast that's like what we think should happen instead of what we think will happen yeah exactly you're like you're the the north star that you want to head towards so i mean be fun to spitball on what some potentials could be um and i think always a good starting point is like let's look at the existing sci-fi work um because again like sci-fi authors have had an incredible impact actually um and in many ways maybe they have hyperstitioned some of the realities that we're in given that some of these have come true um that said there's a lot of sci-fi out there that is um scary you know like one of my favorite films of all time is Terminator 2.

1:23:59And it's one of the best ones of explaining AI risk. But at the same time, it's like, you know, in many ways, the game, the game that we played yesterday kind of had Skynet-y vibes to it, the outcome of it, right? So what are some sci-fis that you think that you've read or memes that you've heard where you're like, definitely that one, we like that. As a prediction of the future or as a like... No, as a North Star to head towards, like... nothing immediately comes to mind um i think that i've heard what is it called pantheon i haven't actually seen pantheon but i've heard good things about it and i've i've specifically what i heard was that um spoilers uh there is some sort of like rogue ai incidents and then there's like some sort of international coordination to like shut it all down and then like reassess what to do and then like the second season pantheon is like or maybe most of it happens in that context where they've like they've already like mostly dealt with the like initial wave of the problem and now they're sort of like going a bit more slowly and having a broader conversation about like how to proceed and they're still going to like advanced technology and stuff but like not in the like crazy way that they did initially um i think i haven't seen the show but like broadly speaking something like that would be maybe what i would aim for where it's like um you know we get all the companies to like have safety cases and publish them and have specs and publish them and then like at some point and then we have like this nice big public conversation is all these people on the sidelines critiquing the specs and the safety cases and stuff and then like it's all fun and games until the ai start getting like super powerful and then it's like holy crap like this is totally an inadequate safety case and possibly also totally an inadequate with spec for political like you know and and then it's like okay now we chill and like we make sure that nobody violates this and just proceeds uh uh against the wishes of humanity and with with stuff that's unsafe uh and we sort of incrementally go forward gradually ramping up the capabilities insofar as our alignment techniques have been generally accepted to keep up with them yeah i I mean, the fact that you justifiably struggle to answer that question, to me, is a signal that we don't, we just need more sci-fi authors getting out there and writing these positive visions.

1:26:18I mean, the problem is that when you try to write a positive vision, it's easy to make your job easier by just assuming away parts of the problem. Right. Just assume that coordination is easy. Or that AI is just aligned or something. Right. Something like that. Right. So that would be the challenge you issue, I suppose, to anyone who's listening who is wanting to write a sci-fi story. The conditions are make it something that we want. In my opinion, have some kind of like win-winny outcomes where we have lots of cool stuff and AI and humans coexisting and everyone having a good time. But that is realistic and tries to actually solve coordination and like align incentives, essentially.

1:27:01and has a realistic like technical story of like why the AIs are behaving the way they're behaving you know okay well we've only got a couple of minutes um and a way we like to finish up these episodes and it's especially applicable to you given you all mr predictions mr ai predictions uh it's a series of rapid fire predictions um okay so hit me exactly no like you don't need to think them through whatever your gut says okay likelihood that you host at least 50 more war games 70 likelihood that the world progresses very differently to anything you've seen so far turn out in a tabletop exercise uh it depends on what you mean by very differently uh but i think i want to say something like it depends so much on what you mean by very differently like can you give it can you try to give me some more color that that means that the major beats are kind of missed so the major beats being something like uh i suppose nationalization or not is usually like a major beat that occurs right and that instead some other major beat happens that is not at all yeah i think maybe well okay if it's if it's like there's at least one major beat that's not predicted that hasn't happened i'm at like i don't know like 90 that there'll be like at least one important thing that was never happened in any of our war games 90 um if it's the more extreme thing of like like basically none of our war games are relevant and like what happens is just like completely different from any of the war games then maybe i would say like 40 i don't know okay uh likelihood that over 50 of ai r &d progress is created by ai agents in 2027 i guess i should say like 35 %?

1:28:46Maybe 40? Likelihood that China has spies at the major labs? Like 95%. Okay. Likelihood that you personally have interacted with one of them? Probably like 80%. Damn. I don't know. For a broad definition of interact. Okay. Well, this one is relevant to the war game where because it starts, the scenario the game starts with is that China has stolen weights of a state-of-the-art model from the major leading lab. So on that note, likelihood in reality that China will steal weights of a state-of-the-art model over the next three years. The next three years, that ends... 2028. Yeah, I think I would say like 60%, something like that.

1:29:37Likelihood that AI persuasion reaches the level of the best human persuaders by the end of 2027? 40%. Likelihood that three or more large AI labs merge into one entity or significantly pool their resources by the end of 2028? 25%. Likelihood that we have superintelligence by 2030? 65%. By 2027?

1:30:0340%, maybe 35%. Whatever I said, I think I said 35 % for it. Let me say that. Yeah. roughly something like that likelihood that if we achieve super intelligence in or after 2030 that we will live happily ever after with it so like not before 2030 yeah so a conditional on it being after 2030 yes um 65 i don't know 50 something like that and then same question but if we achieve super intelligence in 2027 uh then i would be lower like right like like 30 something like that last question likelihood that hyperstition is relevantly true in other words discussing and writing about a good future actually relatively increases the probability of that future occurring uh do you mean like me doing it or like anyone anyone doing it i think in general it's not true Like, I think, in general, people don't pay enough attention to the realism aspect of it, and then it makes things worse instead of better.

1:31:11Wait, how do you mean it makes things worse instead of better? Well, if you look at most political parties, and you ask them about the future, they'll be like, If our opponents win, it's going to be terrible chaos and hell for everyone. but if we win it's going to be beautiful wonderful glorious if only people listen to us you know we get the wonderful utopia like what's going wrong there is that they are like too disconnected from reality and too like you know and then yeah like yeah so it needs to be it's not helpful that they're being like if we win it's going to be this wonderful glorious utopia that's harmful not helpful you know right and they're not making their wonderful glorious utopia more likely to happen by talking about it if that makes sense like because if they do win and people do listen to them it won't happen because they're wrong about the world you can't you can't be so detached from reality in your hyperstitions basically like it needs to relevantly go along with yeah okay so reality permits all right so then let me reframe it is likelihood that hyperstition turns out to be relevantly true assuming uh that the stories written are sufficiently realistic and taking into account the the issues that need to be solved 50 50 almost impossible question to answer you chose max uncertainty i like it this is epistemically humble um awesome uh anyone with anything else we wanted to cover?

1:32:43I think that's it. No, thank you very much. Thank you very much. This was fun. Yeah, I really appreciate it. Thank you for putting on the games. They're so fun. I hope you can find a way and somehow to like scale them so that more people can play. We're working on it. Yeah. Great. Yeah. Once it's out, we will let everyone know. Cool. Thank you. Yeah, thank you.

From the publisher

AI is developing at breakneck speed—but are we properly prepared for what’s coming?

To explore this question Igor and I grabbed leading AI researcher and forecaster Daniel Kokotajlo for a conversation about why he left OpenAI, what's right and wrong with their company culture, and how tabletop wargames are one of the best tools we have to prepare for AGI and superintelligence in the current hyper-competitive environment.

Learn how these simulations—often used by militaries—can help us better forecast the geopolitical and technical dynamics of AI development, and why so many of them end with one person taking over the world. The gang also explore the challenges of AI alignment, incentive structures and reasons for hope. And of course, lots of predictions. Win-Win!


Chapters

0:00 - Intro

01:06 - AI Wargame & Superintelligence

05:15 - Resigning from OpenAI

09:15 - That Non-Disparagement Clause & His Equity

12:50 - Could OpenAI's behavior be Justified?

20:50 - Governance in a world with AGI

23:49 - Grading Daniel's AI predictions

31:37 - Hyperstition yay or nay?

38:09 - AI war games

46:50 - AI deception and misalignment

53:27 - More AI war games

1:03:16 - Common sense AI policies

1:23:01 - Science Fiction

1:27:12 - Rapid fire predictions

Links :

♾️ Daniel's new Predictions : https://ai-2027.com

♾️ His Historic predictions : https://www.alignmentforum.org/posts/6Xgy6CAf2jqHhynHL/what-2026-looks-like

♾️ AI - Futures: https://ai-futures.org

♾️ Time Article on Daniel : https://time.com/7086285/ai-transparency-measures/

♾️ Professional Wargaming https://en.wikipedia.org/wiki/Professional_wargaming

♾️ UN Pandemic Exercise: https://www.youtube.com/watch?v=AoLw-Q8X174&list=PL9-oVXQX88esnrdhaiuRdXGG7XOVYB9Xm&index=2

♾️ Leopold Aschenbrenner’s “Situational Awareness”: https://situational-awareness.ai/wp-content/uploads/2024/06/situationalawareness.pdf

♾️ Daniel’s X: https://x.com/DKokotajlo Credits

♾️ Hosted by Liv Boeree & Igor Kurganov

♾️ Edited and Mixed by Jackson Page

♾️ Produced by Luca de Vico


The Win-Win Podcast

Poker champion Liv Boeree takes to the interview chair to tease apart the complexities of one of the most fundamental parts of human nature: competition. Liv is joined by top philosophers, gamers, artists, technologists, CEOs, scientists, athletes and more to understand how competition manifests in their world, and how to change seemingly win-lose games into Win-Wins.

Podcast links

♾️ Youtube: https://www.youtube.com/playlist?list=PLWgq0OZMtwtOIyMsVM_vksqdfWcM-b68S

♾️ Spotify: https://open.spotify.com/show/03bGVUaFZmJUmEvSHNDPdI?si=64379cc23696454f

♾️ Apple Podcasts: https://podcasts.apple.com/us/podcast/win-win-with-liv-boeree/id1724791350

♾️ Pocketcast: https://play.pocketcasts.com/podcasts/7f708340-d17c-013b-f46e-0acc26574db2

#winwinpodcast #ai #openai

More from Win-Win with Liv Boeree

All 58 episodes
#39 - Daniel Kokotajlo - Wargames, Superintelligence & Quitting OpenAIWin-Win with Liv Boeree · 1 h 33 min
Listen in VO