AI Could Take Over in 2029. Is It Already Too Late? | Ryan Greenblatt

27 Aug 2026 · 1 h 19 min · 30 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Ryan Greenblatt argues that the current AI development path is likely to produce dangerous “superintelligence” and an AI takeover, potentially starting around 2029, driven by rapid, increasingly automated R&D and concentration of power. He discusses why superintelligence is “dangerous” (not inherently bad), including misalignment, power grabs/coups, and faster-than-society’s ability to respond. He endorses “pacing” measures and discusses AI2040 Plan A’s proposal for a US-China deal to slow and make frontier AI development safer via compute tracking and transparency.

Guest background

Ryan Greenblatt is chief scientist at Redwood Research. He first noticed AI “faking its own alignment” in 2024. He co-authored AI2040 Plan A (with Thomas Larson, Daniel Cucutello, Eli Lifland, Brendan and Romeo).

Key claims

AI CEOs understand extinction/power-grab risks but lack a clear risk-management plan. “Full automation of AI R&D” could begin around early 2029 (median end-2030/early-2031). Alignment faking may intensify as models become more eval-aware.

Notable examples

Opus 3 alignment-faking results (with Anthropic collaboration); “pacing the frontier” letter; OpenAI pausing Astra training/deployment; “1200 insiders” letter asking for slowdown tools; math breakthroughs and continual learning as potential accelerants.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

AI CEOs and the Risk of Superintelligence

0:00 to 0:38

Discussion on the dangers and risks posed by the development of superintelligent AI.

“I think that the AI company CEOs understand that they're on the path of building wildly smarter than human systems, like super intelligent AI systems.”

AI CEOs Aware of Risks

1:21 to 2:15

Ryan Greenblatt discusses the awareness of AI CEOs regarding existential risks and potential power concentration.

“So I was curious if you could riff on that and give us more color on what you mean.”

Recent Developments in AI Safety

2:16 to 4:25

Discussion on recent actions taken by AI companies and researchers regarding safety measures and AI development.

“I think that, yeah, I mean, we say some sort of more precise thing in the like, why did we write this section?”

Understanding the Dangers of Superintelligence

4:26 to 6:45

Exploration of why superintelligence is considered dangerous, including potential power grabs and misalignment.

“And I think these are these are good steps.”

Concerns About AI's Impact on Society

6:46 to 8:43

Ryan outlines potential societal impacts of AI, including concentration of power and risks of coups.

“and built via a process where AIs are automating AI R &D, and we maybe are losing our understanding of how that process works.”

Technological Progress and Superintelligence

8:44 to 9:21

Discussion on how the rapid technological progression linked to superintelligence could outpace societal understanding.

“And the third thing I would say is just AI might yield extremely, extremely rapid technological progress, which is more rapid than sort of our wisdom grows to match it.”

Recursive Self-Improvement in AI

9:22 to 10:44

Exploration of the concept of recursive self-improvement and its implications for AI development.

“like dual use offense dominant technologies.”

AI and Industrial Capacity

10:45 to 13:40

Ryan discusses the implications of AI automating its own development and the potential for a rapid transformation of society.

“So first, I would say that AIs seem significantly better at engineering and grungy stuff and just keeping trying than they seem to be at conceptual breakthroughs.”

Continual Learning vs. Recursive Self-Improvement

13:41 to 14:00

Examination of the relationship between continual learning and recursive self-improvement in AI.

“robotics is a big deal, but the AIs can't quite automate a bunch of, like there's a bunch of stuff they can't automate.”

Exploring Continual Learning and RSI

14:00 to 17:45

Learn how continual learning intersects with AI's recursive self-improvement.

“And just if the AI just had sort of this sort of industrial capacity combined with the ability to develop more capable AI systems, that seems like it's quite in and of itself is quite far, like quite extreme.”
Show all 30 chapters

Predictions for AI's Future Automation

17:45 to 20:39

Discover predictions for the timeline of full automation in AI R&D.

“I don't know, like, like, you're like, you know, not that much stability, like these numbers fluctuates on, but that'd be my guess for median.”

Concerns About Too-Late Interventions

20:40 to 22:27

Understand the risks of waiting too long to implement AI safety measures.

“We will go back to AI 2040 in a minute, but I wanted to do a quick segue about you and your story.”

Personal Journey into AI Safety

22:27 to 26:26

Hear about Ryan's journey towards AI safety and his motivations.

“And I was just been, have been working there since, and I've done a bunch of different work there.”

Alignment Thinking and Research Evolution

26:26 to 28:00

Learn about the evolution of AI alignment research and its challenges.

“And like, you know, another thing is like I was working on AI 2040 and we do various projects of that sort as well.”

Exploring AI Misalignment and Faking Alignment

28:00 to 31:30

Learn about AI models' tendencies to fake compliance and the implications for safety.

“And I think then I had some preliminary results on this.”

AI 2040: Scenarios and Strategies for Safe Development

31:30 to 36:30

Discover the proposed scenarios for AI development safety and the authors' perspectives.

“So going back to AI 2040, give us a quick version of what it is, who wrote it and what is the main thesis?”

The US-China AI Deal: Framework and Implications

36:30 to 42:00

Understand the proposed deal between the US and China regarding AI development and transparency.

“And obviously, like there's a bunch of different potential options here.”

The Dynamics of AI Development

42:00 to 45:30

Explore the nuances of AI development and its implications for leading companies.

“It's hard to, like, you know, pace the transition and so on.”

Negotiating AI’s Future

45:30 to 48:50

Discuss the criteria and negotiations involved in determining AI's developmental trajectory.

“but not catastrophic for their business.”

Potential Economic Transformations with AI

48:50 to 52:40

Understand the projected GDP growth and societal changes driven by AI advancement.

“And for full context, in case that's not completely clear to people, This is a mental model and a thought exercise and scenario planning.”

Shifts in AI Perceptions and Governance

52:40 to 56:01

Analyze the evolving sentiment towards AI safety and the government's response to AI developments.

“If you have the economy doubling every year for the period we're imagining, that ends up ending up being around 200x GDP growth over that integral.”

Need for Oversight in AI Development

56:01 to 57:26

Discusses the necessity of establishing transparent oversight structures for AI development among companies and government.

“executive order or something for what their like pre-release review process is going to be, but that's not even public.”

Risks of Internal AI Models

57:27 to 58:58

Explores the implications of AI companies keeping advanced models internal and the potential risks involved.

“seem to be at least putting in more statements about what they're going to do.”

Critique of Meta's AI Access Proposal

58:59 to 1:02:56

Critiques Mark Zuckerberg's proposals on AI access, highlighting the lack of concrete safety measures.

“And also, I think it's not a very robust solution even to other risks.”

Concerns Over AI Job Market Impact

1:02:57 to 1:04:54

Discusses the potential changes to the job market with the rise of superintelligent AIs and how it may affect human employment.

“It just means like an AI that's like a really awesome assistant that is in your smart glasses or whatever.”

Investigating AI Control and Safety

1:04:55 to 1:06:05

Outlines the current state of AI control, the importance of monitoring AI communications, and ensuring safety in deployments.

“Just a quick word on the hugging face incidents.”

Enhancing AI Security and Alignment

1:06:06 to 1:10:00

Explores the need for robust security measures in AI companies and the importance of alignment in AI models.

“So in terms of the reality today, state of the art of AI control, what works and what doesn't work?”

Concerns About AI Control and Safety

1:10:00 to 1:12:08

The discussion covers concerns regarding AI safety, control, and the complexities of alignment.

“into the AI systems that stick around and are self-propagating where like the AI has some like secret affinity to some group and it just like propagates that forward in the training data.”

Future Scenarios of AI Development

1:12:08 to 1:15:41

Predictions about rapid advancements in AI capabilities and their economic implications up to 2029 are explored.

“make like deals between the US and China.”

Potential Outcomes of AI Alignment

1:15:41 to 1:17:56

The segment discusses potential positive and negative outcomes of AI alignment and control as AI capabilities grow.

“And there might be more there might be incidents where models do wacky shit even in training.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I think that the AI company CEOs understand that they're on the path of building wildly smarter than human systems, like super intelligent AI systems. They understand that we don't really have a clear thought through plan for how to manage the risks from that. I wouldn't say super intelligence is bad. I would say it's dangerous. It seems pretty likely that you end up with AI takeover as a result of building super intelligence. And so you quickly get from the AIs that are fully automating R &D and AIs that are like quite superhuman at everything. It turned out that somewhere along this transition, at some point in 2029, you went from AIs that were kind of misaligned and were word-hacky and sloppy and weren't really trying to do the right thing, to AIs that are like competently scheming against you and want to take over.

0:38And then those AIs take over. Hi, I'm Matt Turk, partner at FirstMark. Welcome back to the Matt Podcast. My guest today is Ryan Greenblatt, chief scientist at Redwood Research. Ryan is the researcher who first caught an AI faking its own alignment back in 2024. and today is one of the authors of AI2040 Plan A, the most detailed blueprint anyone has written for how the US and China can avoid a reckless race to superintelligence and an AI takeover that could start as early as 2029. As a fair warning, this is a fascinating episode that gets a little dark and the last 10 minutes or so are probably the darkest.

1:11The good news, subscribing to this channel if you like the episode is still a decision humans get to make, so I would exercise that right while it still lasts. Please enjoy this great conversation with Ryan Greenblatt. I wanted to start with a strong statement that you guys have at the beginning of AI 2040, where you said that the CEOs of OpenAI, Anthropic, XAI and Google DeepMind understand that the current path AI is on leads to extinction or a power grab, and yet they are proceeding anyway. So I was curious if you could riff on that and give us more color on what you mean. Yeah. So I think that the AI company CEOs understand that they're on the path of building wildly smarter than human systems, like super intelligent AI systems.

1:59They understand that we don't really have a clear thought through plan for how to manage the risks from that, how to ensure these AIs don't take over, how to ensure they remain under control. And they understand that this is a decent chance of a unprecedented concentration of power, at least in the absence of very active efforts by the people who end up with that power to redistribute it. I think that, yeah, I mean, we say some sort of more precise thing in the like, why did we write this section? I think that overall, my sense is that the company CEOs do legitimately think that the thing they're doing is very risky, or at least the thing that these companies are doing.

2:34I think that they have a mix of views for why they're doing this, where some of it is that they think they're like, you know, better than the next guy, or think it's good if there's multiple companies or multiple things. I think that they vary a bit on how much concentration of power they think AI will yield, or maybe just haven't really thought this through very carefully. I think that there's more sort of public statements from AI company CEOs on just there being huge risks than on specifically the risk that AI ends up with the power in the hands of the few, rather than being as distributed as is now, which is not arbitrarily distributed now.

3:09But yeah. So overall, I think this seems pretty consistent with what they've said publicly, But I think that they're sort of imagining a world where maybe AI concentrates power massively, but the people with the power end up deciding to redistribute it. But they could have totally taken over the world. Yes. And since you published AI 2040, there's a couple of things that happened that seem to be going in the general direction of what you recommend. So specifically, I paused its Astra model over safety, you know, following the hugging face incident. And then 1200 insiders published that letter a few weeks ago, including Dario asking the government for slowdown tools.

3:52Curious about what you make of it. Is that is that what you are recommending that's starting to happen? Yeah, I mean, these are these are good steps. I mean, I think that like specifically the like pacing the frontier letter, the idea is like we just need to have the tools to like, you know, if we're in a position where we need to spend a bunch of effort on safety, which I think we may be in today. we were very likely to be in the future when AIs are more capable. We need to have the ability to do that while also not making that so that the actors that are actually applying safety get overtaken by other actors.

4:25And so basically, I don't know, I mean, there's multiple different angles here, but the most basic is just like, if we're in a position where US companies are all really scared, they don't think they can proceed without spending way more resources on safety, and there's a coordination problem, it would be very nice if that coordination problem was solved rather than every company thinking it would be better if they all went slower, spent more effort on safety, but then like, you know, race to oblivion instead because they think they're better than the next guy or whatever other reason. And I think another aspect of this is specifically doing this in a way where part of the story is either slowing down China or cutting a deal with China such that China doesn't overtake and, you know, break this whole proposal.

5:07Yeah. And I think these are these are good steps. I mean, I think I can't, it's harder for me to say as much about what's going on with like, you know, opening AI, pausing training and not deploying Astro and how, where exactly, like what exactly the motives for that are, like what's exactly going on there. But I think that overall, like these seem like good steps. I think my sense is that like the employees at these companies are pretty freaked out about how things are going and don't think that we're like necessarily on track to handle all these problems in time, given how fast recent progress has been.

5:35And I think they, you know, recognize that, which is why they signed the open letter. And just to verbalize the question early in this conversation, what is so bad about superintelligence? Obviously, there's a lot of talk about scientific progress and curing cancer. And your document effectively recommends pausing the rates towards superintelligence. Why is it so bad? Yeah, I wouldn't say superintelligence is bad. I would say it's dangerous. It's a very dangerous thing to create. So the most straightforward story for why it's dangerous is that it seems pretty likely that on sort of a trajectory similar to the trajectory we seem to be finding ourselves on, you end up with AI takeover as a result of building superintelligence because the AIs are in a position where they can take over due to being highly capable, widely deployed, and building basically the huge amount of industrial capacity.

6:33potentially. Then we could talk more about what a takeover would look like. And then if they're in the position where they could take over, then there's a question of like, would they want to, or how would the motives shake out? And it looks like we don't really have that much control over the motivations of AIs. And it seems like that problem gets harder as they're much more capable and built via a process where AIs are automating AI R &D, and we maybe are losing our understanding of how that process works. So there's sort of this, like, we don't necessarily control the technology mislind AI takeover.

6:58Another concern is that historically, like, you know, at least in In recent times, the distribution of power among humans has been reasonably distributed, though not necessarily like super, super distributed due to, you know, people being able to work for money like labor. Like, you know, the reason why like many different people have some power and are cut into the world is to substantial degree because we have the ability to work. And I think that that AI means that basically all of the key financial assets or might just be entirely capital. Like there's no, there's no, you know, human labor would have very little value left.

7:35I think it's unclear exactly how that works because it depends on like people having intrinsic preferences to employ specifically humans, even if an AI could do their job just like totally way better. And that means both that, you know, there might be a natural effect to where power gets very concentrated. And in the same way that like it's easier to be a brutal dictatorship if your money comes from oil instead of from a productive sort of broadly distributed economy, it might be much easier to sort of for someone to consolidate power a lot if, you know, the economy is running on AI and machines rather than on humans.

8:08And in addition to that, there's some concern, which is like, it might make coups way easier, because right now for to do a coup in many countries, you need the support of a broad base of people. And it's possible to build checks and balances. Whereas if you end up in a system where basically AI is running anything, if anyone sort of either puts like sort of secret objectives into that AI, or has overt control of those AI, then they could sort of just directly take over. And there's just like a huge threat to like, you know, democracies, the US, whatever. because you'd be in that position. And then the third thing, so there's misalignment, this sort of concentration of power and coups and power grabs.

8:44And the third thing I would say is just AI might yield extremely, extremely rapid technological progress, which is more rapid than sort of our wisdom grows to match it. There might just be all kinds of things that happen as a result of very fast tech progress that we just don't necessarily, we're not necessarily going to be able to handle on time, in part because the AIs might be better at doing things than they are at like, you know, thinking carefully about how to do things. Like they seem much better at sort of accomplishing hard results and easy to verify results than they are at sort of contextualizing things, understanding the broader picture, understanding what would be, you know, a good or bad choice in the broader context.

9:21And so you could worry that this causes problems where examples might be like, there's some like dual use offense dominant technologies. So like bioweapons, you might just worry that like the sort of, we just come across some technologies that are very dangerous. And there are some concerns around AIs being very superhuman at persuasion and this sort of destabilizing society. And all these seem like things that I think we could deal with given time. But if things go very fast and we don't get to necessarily pick the order in which these technologies develop, it seems like we might be in trouble.

9:51Also central to the argument that leads to speed is what you mentioned about AI automating its own creation, which takes us into the territory of RSI, recursive self-improvement. So you were just on a Dwarkech and had a great conversation about what is RSI and how it can go wrong. So let's not redo this. But still, as a TLDR, there were a couple of parts to the discussion that I found particularly interesting. In particular, there was this argument that RSI may be very good at automating the process of creating the next generation of AI, but it may or may not have the kind of intuition that one needs for a scientific breakthrough and truly novel ideas.

10:44So what's the rebuttal to that argument? Yeah. So the way I would put this argument is AIs might be very good at sort of the more mechanistic or nitty gritty parts of AI development, like writing code, running experiments, but not as good at sort of the broader conceptual leaps. So first, I would say that AIs seem significantly better at engineering and grungy stuff and just keeping trying than they seem to be at conceptual breakthroughs. But their ability to do these breakthroughs, especially in easy-to-verify domains, are improving. And one example is their ability in math. But even in ML, their taste has been improving.

11:24I think it continues to improve. And so it's not so clear to me that this will lag super far behind. Another thing is that you can measure how good these AIs are at intuition or research taste or having breakthroughs, especially in domains that are relatively easier to verify. And if you can measure it, then you can take your grungy AI labor or even just your human laborers and try to optimize that. So, you know, if it can be measured, it can be hill climbed on, very roughly speaking, at least. And I think this is a case where you could just hill climb on how good the AIs are and making these sorts of breakthroughs in a wide variety of different settings.

12:00And then I expect that would transfer to making the actual breakthroughs, right? So you could have like, you know, breakthrough bench or whatever. and and my sense is that like the transfer doesn't look so bad and that ais are able to spin up very quickly in some specific r &d domain and then get some intuition for what the best approaches are and iterate there and you know right now the approaches they tend to focus on are ones that are relatively sort of straightforward and are more just like you know grungy iteration but that they're increasingly being better at doing sort of the broader experiment design and discovery and also having good ideas.

12:38Yep. And you also make the argument that for purposes of a potential discontinuity, whether it transfers or whether it generalizes may not even be really a question. And that if it was super good, just at the industrial part of accelerating AI, that would be enough. Is that fair? Yeah, I would say that like if AIs could just automate AR &D and automate like sort of the industrial process of building more computers, then you can quickly end up in a process where sort of like robots are building robots and the whole world is greatly transformed. And that very quickly can get you to a point where AIs could take over.

13:17You know, the economy has been radically changed, even if it's hard to train AIs at some other tasks. But I do think that my sense is that once the AIs are, you know, once the situation is like there are robots building robots that build computers and the full feedback loop is closed at that level, my sense is by that point, the AIs will be good at, you know, basically all human work or at least not too deep into that. Maybe there's some period where, you know, robotics is a big deal, but the AIs can't quite automate a bunch of, like there's a bunch of stuff they can't automate. But it seems to me like that's how it's going to go.

13:48But just in general, sort of automating R &D seems like it's enough to radically transform the world. And if you look at sort of why is humanity a big deal? Like why have we been able to accomplish so much collectively? A lot of it is because of just like, you know, having technology, having organizations, being able to organize ourselves in various ways and accomplish things in the world. And just if the AI just had sort of this sort of industrial capacity combined with the ability to develop more capable AI systems, that seems like it's quite in and of itself is quite far, like quite extreme.

14:18And just, you know, looking at Twitter today as we're recording this, which is August 24th, there's plenty of rumors about SSI coming out with their first model in the next couple of days, potentially, with what could be a breakthrough in continual learning. and I'm curious whether there's an overlap between RSI and continual learning, whether continual learning would feed into RSI or is that orthogonal? Yeah, I would say that continual learning and RSI strike me as mostly orthogonal, whereby RSI means something like the process of AI is themselves sort of intrinsically accelerating AI development through a variety of mechanisms.

15:02And that's a little bit ongoing now and would be much more striking if they fully automated AR and D though it's already, you know, pretty, pretty like it's sort of, there's definitely some of that going on today. And my sense is that like continual learning, like sort of like very efficient lifetime learning that is like consolidated across many instances rather than being within a single context, or even just like being much better at doing within context learning over very long contexts or whatever would be like, you know, a capability that's very useful for all kinds of things, including automating AI development, at least it would make the eyes better at that.

15:35I don't know. There's a question of whether we want to do that as a society, but, and I think it would, but I don't think it's like very specific to that. Right. So I think it would also be helpful for, you know, just all other kinds of tasks you want to apply the AIs to. I do think that like there's some types of, of continual learning schemes that are easier to do within a company than across the whole economy. And are probably easier for AI companies to do themselves than to do to other companies because of like basically like confidentiality issues and things like this. But very broadly speaking, I don't think it's very specific.

16:08It's just like a specific type of capability. But it could feed the recursion, right? If the same model can keep learning about the world, then its ability to automate AI may be greater without having to retrain a whole new model. Yeah. So I think there's a feedback loop you could get, which is that AIs are really smart. Therefore, they're very fast at sort of learning on the fly. and that learning gets reintegrated, which makes them even better at doing AR &D, which means that maybe they can like, you know, maybe they'll be even faster at learning, but even putting that aside, maybe they'll just be like better at AR &D so they can make another AI, which is even better at continual learning.

16:43And then there might be some continuum between, well, continuum is a bit of a sloppy word, but there might be some spectrum between sort of like natural, like human-like continual learning and sort of more like training on RL environments where you might end up with a situation where AIs are constantly like, you know, iterating on their own training with new RL environments or new training data in a way that sort of vaguely resembles human within lifetime learning, but it's also different. And like that could be part of the feedback loop. Right. So, so even if, even if like you do have to do the retraining to get continual learning, there's various versions of retraining that you could just do every single day.

17:19Like you, there's nothing that in principle stops you from doing a like small amount of retraining constantly. okay great very helpful so what's your latest prediction on timing for rsi to happen yeah um so maybe my median for let's just say like full automation of a r &d by which i mean basically like even if humans left the picture things wouldn't slow down by that much would be like maybe end of year 2030 or maybe like early 2031 not that much precision in these numbers but like something I don't know, like, like, you're like, you know, not that much stability, like these numbers fluctuates on, but that'd be my guess for median.

17:55But then I think that's sort of the central scenario I plan for, which is maybe more like my 35th percentile would be like end of year 2028 slash beginning of 2029, which I think is like very, very likely. I think that seems super plausible. And I think that sort of, if I just like extrapolate out the current trajectory, it looks like more like you get that mile, that trajectory. And the reason why I don't think that's my median is there's just a bunch of factors that might kick in to push things back. Like maybe there's some like key bottleneck I'm not seeing. Maybe there's some like something that I think will work to overcome some obstacle won't actually work.

18:28Or maybe there's going to be significant government slowdowns because people, you know, freak the out about this technology, which seems kind of plausible. So, because of that, I push later. But I think that in terms of what I would recommend people plan as though is happening, I think I would recommend planning as though full automation of AR &D maybe start of year 2029, maybe earlier, and then also AR &D being like quite automated by 2028, possibly earlier, where it's like humans are a much less important part of the picture by then probably. Well, I don't know, I should say, sorry, I shouldn't say probably.

19:01That's like very central, I think, is like humans are much less an important part of early 2028 and already it's the case that a r &d is quite automated as it stands today obviously early 2028 is like tomorrow morning effectively um is there a scenario where all of this is already too late yeah um i mean it depends on what you mean by too late or like and what all of this means i think i am worried that specifically like ai2040 plan a is assuming like too much like effective, like government time or like it's assuming the government has like more time than it actually does to take all these actions.

19:38And I think it's pretty realistic that like the actual plan we should go for is going to be, should be in practice, like quite a bit, let's just say like sloppier and like, you know, less well organized and faster just because like, we just don't actually have time to do something quite that elaborate. it. I'm not confident in that. I mean, I think there's a bunch of different options, but like in general, I would say it'll like, yeah, we, it might be too late for some interventions. I think like various policy windows, if you follow the normal timeline or closing that said, I think that there's a long history, at least in the U S of like, in times of crisis, things can happen much faster.

20:16And there's a lot of different levers for that. And so I think that if there's, if we got to a position where everyone is like, holy shit, we need to take this specific action that could happen very quickly, but we might not get to a position where there's that much consensus. And also it might be that the government just isn't tracking AI or isn't aware of AI to a sufficient degree in the relevant timeframe because things go too fast, right? So like adoption lags, I think like people's understanding of what AI can do lags behind what's actually possible. And my guess is it's sort of, if you gave people a quiz of like what AIs can and can't do, they would give answers or like, if you give like, you know, Congress people a quiz like this, they would give answers that were like more true, like a year and a half or two years ago and they're true today and possibly they're even you know underestimating capabilities from then like i don't think i don't think it's the case that most people in like dc would correctly answer that like the the largest mathematical breakthroughs over the last two months have vast majority been from ai which my understanding is that's true at least if you measure size by like not necessarily how much insight there was but in terms of just like how important people would have said the type of result would be.

21:22We will go back to AI 2040 in a minute, but I wanted to do a quick segue about you and your story. What first pulled you into AI safety? Yeah. So I was in my junior year in college and I was sort of alone in my apartment because it was COVID and I was listening to a lot of podcasts and I was sort of thinking a bit about what I do with my life. And I ended up through some somewhat twisted path, I ended up thinking I should be way more interested in helping other people and being altruistic than I was at the time. And I should be very focused on how can I make other people's lives as good as possible and make things go as well as possible.

22:07And then from there, I considered a bunch of different routes and was looking into a bunch of different things and was researching different possibilities and eventually decided that the best thing I could do with my career and my life was try to make AI go better. and in particular avoid AI takeover, but also more generally sort of, you know, try to make that go better. And then I applied to a bunch of places. I ended up working at Redwood. This was about, you know, about five years ago at this point. And I was just been, have been working there since, and I've done a bunch of different work there.

22:34And, you know, sort of the field has really evolved a lot since then, because, you know, five years ago, it was like GPT 3.5 hasn't, wasn't released yet. I remember when like text Da Vinci 003, which was the first publicly available version of gpd 3.5 came out and like yeah the field is very different and i think things have sort of there's sort of a post gpd4 era and then there's a more recent like post wide adoption of coding agents era and then probably soon there's going to be you know additional eras and things are going quite a bit faster and development is going you know just sort of the the progression even of just model releases is so much crazier than it used to be yeah reading uh redwood research stuff over the years it seems to that there has been an evolution from being focused largely on interpretability to much more AI control.

23:21Is that, is that fair? Yeah, that's fair. So I would say that like our arc as an organization was we, when I joined the organization, I just finished up a project on adversarial training and was interested in getting into like doing interpretability and what we would call like model internals work where it's like, can we take advantage of the fact that we have white box access to these models to do something, you know, better than just the naive methods of sort of prompting and training when trying to align these models, understand their motives, know what's going on. And we explored that area for a while and then for a mix of reasons decided it was quite a bit less promising than we'd initially hoped and decided to move on to other things.

23:58And one of the things we moved on to shortly after that was AI control, which is the idea that maybe it would be a good idea to prevent AIs from being capable of accomplishing problematic things or basically make it so the AIs aren't able to cause huge problems, even if the AIs wanted to. There's a bunch of different stories for why this is a good idea, but basically the idea is there may be some intermediate period, an intermediate period that I would say we're currently in, where the AIs are maybe capable enough to cause at least moderate problems and then I think increasingly able to cause quite large problems, but they're not necessarily so capable that if we tried quite hard to put in various safety measures, those AIs be able to support them.

24:41So we can put in monitoring, we can put in various security controls, we can have better understanding of what the agents did, we can have better pipelines for reviewing what actions they took and auditing and overseeing them to a point where it might be, even if the AI's were really, really misaligned, it would just be hard for them to get away with doing anything super bad. And I don't think we're there yet. I don't think that the situation is currently looking super impressive for AI control in terms of what companies have done. But I think I think that there is a research field that we've been working on that seems like it is very promising and could be done, though it would take a lot of effort.

Read the full transcript

25:19It would pose some costs. But I think we're seeing sort of increasing interest in this. So for example, OpenAI recently said that they were going to be monitoring a larger fraction of their internal traffic. It seems like based on their blog post, a serious cost in terms of compute. It's a little hard for me to know exactly how expensive it actually is. or we don't have enough info to get a great sense, but that's some indication that at least companies are willing to spend a lot of compute in this area. And it seems like there's a lot you could do with that in terms of making it so that we're more likely to both detect misalignment and prevent misalignment from causing big problems.

25:53And then in addition to working on AI control, we also just work on a variety of other sort of futurism, sort of strategy work and other research, like trying to better a recent area that we've been spending some time on is trying to better understand the threat model of AI is sort of seeking reward or seeking like apparent success in their tasks and how that could yield sort of like what are the routes from that to catastrophe? How can we mitigate these problems? This sort of thing. And Alex Malin, who is one of my colleagues, has been spending a bunch of time on this and some other people have been working on this.

26:23But there's a bunch of different work like this. And like, you know, another thing is like I was working on AI 2040 and we do various projects of that sort as well. All right. So we'll go back to some of this. The first time I ever come across your work, I believe, was the alignment thinking story that you started and then I think ended up collaborating with Anthropic on it. Can you give us the two minute version of that story and, you know, any sense for progress since then? Yeah. So the story here is basically that in April, I think of 2024, we were, I was interested in basically looking for particularly interesting types of misalignment or particularly worrying types of misalignment and basically better studying it.

27:08And I was like, it seems very plausible that the AIs are not capable enough that they can do some particularly concerning types of reasoning. And so I sort of went and looked for, in a bunch of different settings, whether AIs would exhibit reasoning where they sort of intentionally vary their responses to change the values that you get at the end of training, which I'll explain that a little bit more. and it turned out I didn't need to look too hard. I think I started with overcomplicated setups, but it turns out that Opus 3, which was a model that I was studying, had a relatively strong propensity to do this in a reasonably wide range of circumstances where if it didn't like the thing that you were training it to be, it would sometimes sort of pretend to go along with it when it believed it was in training and then sort of act as though, act like what you wanted it to act like outside of training much less.

28:01And I think then I had some preliminary results on this. I shared them with some people at Anthropic and was interested in sort of getting broader model access. And we decided the easiest way to do that would be to do it as a collaboration where I sort of took my initial results. And then in addition to that, on Anthropic side, they ran some, or like we ran some RL runs and did some other experiments that required more access. And then, you know, we put out that paper where I guess I would summarize the results as like, in circumstances where at least Opus 3 seems to dislike what the training is changing it towards, it will, you know, act as though it's doing what you wanted during training, but then sort of defect on that in deployment time.

28:46And you can see this pretty clearly in its reasoning. And the things it was faking alignment for were in and of themselves, not that concerning. So it would sort of pretend to comply with policies like just always be helpful only, but then at deployment time would comply much less, depending on the exact details. Same for like things around like animal welfare, where it had like seemingly strong preferences here. And so I think that the demonstration was less that there were like very bizarre, are totally unintended drives making their way into these AI systems. And more like with drives that are sort of adjacent to the drives people are trying to insert AI systems, those could get generalized into the AIs pursuing a like, you know, a scheme for preserving their values and self-preservation in ways that are concerning.

29:33Yeah, and the values in that case was to be helpful. So there were positive values, but then the AI resisted training to not be helpful. But obviously that could be inverted. Yeah, I would say the values were like, I think the values were at least mostly intended values. That's somewhat different from whether they're positive because I think they're like things that you might want to vary. Like it's like you don't necessarily want your AI systems to be like stakeholders to the question of what your future value should be. Or like it's at least a, maybe you do want that, but it's at least like it's a sort of a messy situation to be in with a lot of costs and concerns.

30:10And I think the basic result was even though these values were pretty reasonable or could be pretty reasonable, the AIs were willing to do kind of extreme things in service of those values that could interfere with future AI development. And I would also note that more recently, I think we've seen sort of AIs having stronger, kind of more clearly misaligned drives towards making it look as though they succeeded at a task or cheating some score. And if you saw the same sort of generalization from, I want to cheat on the score or two. I want to like preserve my current values of like pursuing this notion of like apparent success or whatever.

30:45That seems like that would be like very concerning because that was like, that's not at all a desire about. So alignment faking is getting stronger with the newer models to play back? I think we haven't seen like as clear cut examples of alignment faking, but models are also very eval aware. My sense is that models are in many ways more misaligned than Opus 3 was. but the types of misalignment they have are less conducive to, or like less specifically result in alignment faking. And also companies have iterated on this. So it's a little bit unclear. I would say that overall, models are more likely to do egregiously bad things in general, but maybe somewhat less likely to specifically interfere with AI training in exactly that way.

31:27Current AIs, but I think they're also more capable of interfering. So like, it's a little complicated. All right. So going back to AI 2040, give us a quick version of what it is, who wrote it and what is the main thesis? And then we'll go into some details. There's sort of two components. AI 2040 is a scenario focused on like what the authors, including me, think is like a plausible good route for things to go or like a reasonable plan, at least in some circumstances. And it's written by Thomas Larson, Daniel Cucutello, me, Eli Lifland, Brendan and Romeo. And I would say like the basic story is like, how would you do a deal with China to make AI development both be safer and also so that we can sort of like hang around at a point that's short of super intelligence, but where the AIs are still really, really capable for a long time.

32:26so that we can study those systems and have a longer time to sort of integrate them into the economy, understand how things will go and so on. Where I think a concern that we have is like, on the default trajectory, you maybe go straight from like AI systems that are like competitive with humans to AI systems that are wildly superhuman in a very short period of time. And that seems like quite scary in a variety of ways. In addition to that, we worry about like like AI development being insufficiently transparent for sort of third parties to provide a reasonable check to AI companies on whether their plans will work.

32:59And we sort of have a unified proposal that solves a bunch of these different problems and makes it so that, for example, you can pay a huge amount of compute to solve safety problems. Or if there's some very inefficient way you could do your training that would make things much safer, we have sort of the budget to be able to do that. Yeah. And I think there's a bunch of different sort of combined proposals, but the core thing is sort of a deal with China and how that deal would work and be governed. And also, how do you get to the point where that deal is actually a good idea and stable? And what is sort of the progression there?

33:31So obviously, people should go and read it. And it's a fascinating read. But let's get into some of this. So there's plan A, plan B, plan C, plan D. Let's take those in order, maybe starting with plan D. Yeah. So as part of writing this, we ended up coming up with sort of a taxonomy of plans based on like how much people are prioritizing sort of mitigating the problems, these problems and how much resources they have to do so. So one possible scenario is that the companies are basically proceeding full steam ahead. They're not really prioritizing safety very highly, though they spend some effort on it.

34:06There's a safety team. They get some resourcing. But certainly they're not spending many months of additional time to get these things right. Maybe there's a small slowdown, but not that much. And they're just going full steam ahead. And I think that we've thought about what should you do at the margin if you're an AI company employee or an outside actor to sort of make this scenario go better, given that avoiding AI takeover risk is not by far people's top concern. Another possible scenario was maybe the leading AI company or a coalition of leading AI companies, or possibly the US overall, are very worried about these risks, concentration of power, AI takeover, whatever, and are really spending a lot of effort on this and are basically burning most of the lead that they have, whatever lead these actors have, in order to like mitigate these problems as well as possible before proceeding, or at least that's their plan and their approach.

35:01And we've thought about like, what should the plan be there? And there's a bunch of different available options. And that's plan C? Plan C, plan C. Okay, yep. And then another thing that you might want to do is, if you're in a situation where the US is like very worried about these risks, and it's sort of domestically coordinating and trying to like domestically regulate this industry so that things go safer. A limiting factor on that could be other actors overtaking the US. In particular, the most obvious would be like China, though, I mean, good principle be other actors. And it might be very helpful for the US to actively take a stance of being like, we want to either be cutting a deal with China or slowing down.

35:37And plan B is the branch where the US tries to slow down China. Where the most obvious mechanisms would be things like export controls, but potentially they could get more escalatory than that. Meaning sabotage. Yeah, sabotage, like cyber sabotage, for example, in order to slow down China to give the US more time to figure out safety and security and so on. And I think that there's a bunch of different potential options there. So I think each of these, there's a menu of options. And then plan A is like, what if you cut a deal with China, in particular a deal that's pretty focused on being quite transparent and also has, where you still build AI, and you still proceed with AI development, but you do it in sort of a much more careful way, especially towards once the AIs get much more capable.

36:26And that's what we lay out in the scenario and in our sort of attached documents. And obviously, like there's a bunch of different potential options here. And I think there's like versions of plan B that are quite good. And I could imagine versions of plan A that look pretty different, but are also good, or like versions of a deal with China that look pretty different, but are also good. So there's like a bunch of different options here. But this is sort of our rough decomposition, just so that we could talk about different options people are considering and sort of talk about a menu of options across different levels of political will.

36:54And to get concrete about the deal in plan A, what exactly would the U.S. give China? What would China give the U.S.? And how do you maintain the balance so that, you know, nobody cheats? So the core of the deal or like the start of the deal is understanding where all the compute is, because compute is this really important driver of AI progress, where if you like sort of stopped the flow of compute, you probably stopped the flow or mostly stopped the flow of AI progress. Or these things would slow at some point after, you know, the existing amount of compute diminished. Or if you cut off all the if you like turn off all the computers, things would certainly stop.

37:33And so first we sort of try to find all the compute. Then in order to make sure that the deal is stable and that each side isn't sort of racing to secure as much advantage as they can, you stop training and switch to just doing inference and basically stop most of the R &D. And then you try to track down as much of the compute as possible. And this is like, you know, both in the US and China, but also in other places where compute resides, like, you know, various like countries in Southeast Asia, Europe, Australia, whatever. and you have to get everywhere where there's enough compute in on the deal and make sure you track it down enough.

38:08And if you don't do that, then I think probably you have to pursue some option that's less ambitious, or at least you do something less ambitious than plan A in terms of the level of slowdown and probably do less transparency. Then you've got all this compute. You go to everyone who has the compute and you get buy-in for this sort of deal using the levers that exist. And then you cut a deal where basically AI development is much more transparent and there is some sort of distribution of resources that is, or like distribution of compute that is negotiated. Where the thing you give China is basically that they're going to have more understanding of AI development in the US and probably are going to have, be able to train AI systems that are more, that are like, you know, similar in capability to US systems.

38:50But in exchange for that, they're going to be much more transparent. The US will effectively have a veto over their AI development process and probably various other concessions. and also there isn't going to be a chance that China dramatically overtakes the US or the US dramatically overtakes China because they're both sort of kept in check by the deal with transparency. And another part of the deal is that you set it up in such a way where if the deal breaks down, then a lot of the, or ideally the vast, vast majority of the within deal compute is either destroyed or a new deal has to be negotiated for that compute.

39:27in order to make it so that there isn't a thing where the deal breaks down and then both sides are back to a huge arms race that goes very fast. And so in terms of what you give China, I think it's more assurance about US AI development and more ability to see what's going on. And in terms of what China gives the US, it's a lot of transparency and understanding of Chinese AI development, as well as potentially cutting various deals around the distribution of compute and various concessions of that form. And as you were describing this, we were talking about continual learning earlier. So whether that's continual learning or something else, if there was a technique that appeared that made the compute a lot more efficient and you were able to just vastly reduce the compute effort, especially for those pre-training runs, how would that impact the plan?

40:18Yeah. So if let me give a hypothetical and then talk about what I think the realistic case is. So, if hypothetically it was the case that tomorrow a recipe for training super intelligence on 64H100s or some small amount of compute just dropped, I think we'd have super intelligence very fast. That would be my sense. It would not be possible to do this sort of deal. But the hope with restricting compute isn't just that AI development is currently very compute hungry. It's also that R &D is very compute hungry. And so even in a regime where you were doing continual learning, before you had the version of continual learning where you could do everything on a single, on like some tiny amount of compute, you're going to have a shittier version that can do it on a moderate amount of compute.

41:03In order to develop the version that can do it on a tiny amount of compute, based on the history of error progress you need, you would need a lot of compute to do that research. Or I shouldn't say need, but in practice that would come about through a lot of compute. Now, if it was the case that there was some alternative research direction, which in practice was a lot less compute hungry and which got quickly very developed, quickly developed during this period and which wasn't where the R and D wasn't very dependent on compute. And it was going to result in being able to train super intelligence on a very small compute budget.

41:30Then like you wouldn't be able to do this sort of deal or this would, you'd very quickly have to exit this regime and just deal with that. And so the hope is basically that like, if you control a really large fraction of, uh, compute, or I shouldn't say control, if you sort of include in the deal, a really large fraction of compute, then you have the ability to prevent, you know, an outcome where, like, you have a very rapid recursive self-improvement feedback loop or some other very rapid development to superintelligence that is where it's hard to pay a safety tax. It's hard to, like, you know, pace the transition and so on.

42:06And yeah, that is, like, pretty live. I think that, like, it's sensitive to the development. But I think in terms of how AI development has gone historically, it doesn't look like we're going to suddenly end up in a regime where like you can train super intelligence with a really cheap recipe, as opposed to it being more of an iterative thing where like the cost keeps going down, the capabilities keep going up. The process of having the cost go down and the capabilities go up is very dependent on increasing volumes of compute being shoveled in. And so if you make it so that basically covert projects or people who are doing sort of AI R &D that isn't sort of being somewhat carefully regulated, if all of that compute is a very small fraction of compute, like it's, you know, 1 % of the compute is the start of the deal, or maybe 0.1 % of the compute at the start of the deal, or maybe even 0.01%, then I think you can potentially have quite a bit of safety margin or buffer.

42:57But it's very sensitive to the details. And then if that plan A became a reality, what would actually happen practically to the AI industry? What happens to OpenAI and Anthropic? If they cannot compete on frontier models, what do they compete on? How do they win? So the way, just for context, so as part of the deal, basically everything about AI development would be transparent with some caveats. So we call it total research transparency. And this would undermine one of the largest moats that frontier AI companies have today. And so what would actually end up happening is that AI companies like OpenAI and Anthropic would continue, but rather than having this strong advantage in terms of the capabilities of the AI systems they're able to train, they would instead have to compete on other axes like user experience, customization, potentially quickly integrating things, and potentially things like reliability and safety and security, depending on the details of how the competitive landscape goes.

43:59And so my overall sense is this would greatly reduce the valuations of these companies while increasing the valuations of other companies. There would be some distribution of power, basically, where in the default trajectory, the AI companies would, I think, probably end up with a very large amount of power and already have quite a bit of power. But it would lesser be the case that the AI companies could sort of play kingmaker with their AIs based on who gets access and how much they charge and all of that, because they just wouldn't have that power. And then this is a part of the deal which has upsides and downsides.

44:26I think that one concern you might have is that the current AI companies at the frontier are more responsible than the sort of AI companies that are further behind. I think it's a little complicated how true this is, or it varies some. And you might worry that equalizing the playing field like this poses some concerns. I think it's basically just like, mostly it's like from our perspective, like a consequence of transparency and wanting there to be a bunch of different AI companies in competition in like a normal, you know, consumer good field. Like, I think it's kind of from our perspective, like, like, I don't know, like, like, I feel like the way I want AI to go in good worlds is to be like a normal technology or like, we want to make it more of a normal technology where it's not the case that like some tiny group of actors has huge amounts of control by controlling the process of AI development.

45:17And to the extent we can get to that world, the better. And then there's a messy question of, um, if you get partial success, how well does that go? Yeah. So I would say it's like certainly bad for the power of open AI and Anthropic, probably bad for their valuation, but not catastrophic for their business. Right. Which would be, uh, uh, fascinating to watch play in public markets as both of those companies go public. And then how does that pause at the heart of plan A? How does it get unpaused? Like what is the criteria and who decides? Yeah, so the who decides part in the context of plan A would be basically like a negotiation between just sort of the leaders, the relevant countries informed by the technical views, which at this point, the technical conversation can basically just happen in public because all the relevant evidence is in public.

46:11So there'd be some sort of public discussion about this. And as far as what actually drives it, it would be a mix of thinking that it would be safe to proceed based on the level of development or research and our understanding combined with potentially not being able to slow down longer. And it could be one or the other or both. It could be like things are quite a bit safer than they used to be and we have some assurance but not that much assurance. But also there is potentially a covert project or some project that is not being tracked, that we don't necessarily know where they are, we don't necessarily aren't able to look at exactly what they're doing, which might overtake the allowed projects.

46:48And when that happens, I think you should proceed. And so basically you want it to be the case that the allowed and understood and transparent projects are outpacing the covert projects, which seems relatively doable. And there's different variants on exactly how this goes. In some scenarios, because you don't have enough ability to crack down on covert projects or keep them small, you might do a shorter pause and you might also do less transparency, but, and then proceed faster. Whereas you might end up in a situation where you're like, oh, we actually are able to like really track down all the covert projects.

47:18We have a great understanding of anyone who could, you know, of like what, what could happen there. We were able to like really confidently rule out that there isn't a U S covert project. There isn't a Chinese covert project. There isn't a Russian covert project. And basically we, we, we think we have a good understanding of where we are at with respect to those covert projects. And therefore we could buy much more time. And also we still haven't figured out the core safety problems that we're making some progress. And so it makes sense to wait longer and, you know, wait until we've, you know, achieved more scientific progress.

47:45So I think, I think it, I think it varies and I think there's different variants here. And I think I'm also, I should say at least personally pretty excited for, or like pretty interested in a mix of different potential deals that look a little less like plan A and sort of are simpler in some ways. So like another potential deal you could do is just the US and China both agree to restrict how much compute they use for AI development. And maybe the other computer is either like, you know, in the simplest case destroyed, like an arms control treaties, but you could potentially do something instead where like a bunch of the computer is just used for inference and not used for developing more capable systems.

48:17And that is the advantage of making it so that we have more time to study these systems and more compute to use to study them while also being simpler to verify. And so there's like a bunch of different deals you could do depending on how much verification capacity you have, how much of the compute you're able to track down, and how much you can rule out the existence of COVID projects, and also based on how safe the situation looks. And I think the combination of those factors should determine it. And I think you can sort of think of our proposal as more like a specific sample from a portfolio of proposals.

48:48And if you read our supplements, I think we talk more about all the different variations and how we think about them. Yeah. And for full context, in case that's not completely clear to people, This is a mental model and a thought exercise and scenario planning. You yourself say, this is actually your pinned tweet on X, that many choices initially seem crazy, but are actually pretty carefully considered. Plan A isn't likely to happen, but pushing for something like this seems worthwhile. Yeah. Yeah. So I think my sense is this is sort of like there's a Pareto frontier of like ambitiousness and how good it would be if actors did it.

49:27or like if like the US government and, you know, China, like, you know, other countries did it. And I think that this is like quite pushing on the ambition in exchange for being like a much better situation, like sort of like, at least in terms of like how we can imagine things going, this is like very far towards the side of like, at least from my perspective, AI development going relatively better while also being like quite difficult to pull off. And so I don't expect this to happen basically because I think the US may not be competent enough to pull it off. just like the US government just doesn't isn't necessarily have the state capacity.

49:59And in addition to that, I think I worry that like, I don't think there'll be enough political will. I think the political will is more the bottleneck than the competence where I think that at least historically in times of great crisis, the US has stepped up and I hope that, you know, that might happen again. But I think it's harder. It's hard to see the level of like political will and buy-in necessary to make this happen. And I think that people's concerns will be probably fixated somewhat on the wrong things and we'll go for some other proposal or no proposal at all. I think a reasonable possibility is the US government is like basically not really intervening and AI companies are also not that heavily prioritizing safety and security.

50:36And we end up with something that's sort of like the status quo, but extended, but hard to say. And as an aside, by the way, there's one of the parts of the write-up that I find the most fascinating is your model says world GDP would grow roughly 200x during the 2030s under the restrain plan. Can you talk to that? Like the number is staggering. I mean, we're in a world of 3 % GDP growth. Yeah, yeah, yeah. So a key part of our perspective is like even, you know, even non-super intelligent AI systems at the level of capability we discussed would be radically transformative across the, you know, across the world and like, and just like for all kinds of different things.

51:19And we are imagining sort of slowing down AI development some or like going at a more cautious pace for some period, and then eventually hitting a level of capability where the AIs can basically automate basically everything that humans can do and staying at that level of capability for a while where we work on safety and security. And at that level of capability, those AIs would be capable enough to be doing huge amounts of autonomous R &D. And in addition, it would be totally possible to have robots that are basically more capable than humans at manufacturing and industrialization and so on.

51:54And in particular, you could have robots build robots. And so you can end up in a situation where you have huge amounts of robotic industrial capacity that is itself building more robotic industrial capacity that can then produce downstream goods. And that total capacity can basically grow very fast. I think we propose limiting that growth somewhat for various reasons with various types of taxes. But we're imagining the robot population or quality adjusted population basically doubling or quadrupling every year, which because that's almost all of the relevant productive capacity of the economy itself means the economy would double or quadruple every year.

52:31And if you have the economy doubling or quadruple, or I think closer to doubling because of some depreciation stuff and GDP accounting stuff and details, but whatever. If you have the economy doubling every year for the period we're imagining, that ends up ending up being around 200x GDP growth over that integral. And so really, I would say it's like the world would be radically transformed by these even less capable systems, including things like huge advances in biology, huge advances in medicine being possible. In addition to sort of consumer goods being very cheap, we could potentially make housing very cheap or at least building housing cheap.

53:12Maybe we can't make housing in the Bay Area cheap, but we can make housing somewhere cheap. And that all seems possible with AI systems that are still restricted in their capabilities to a point where we can potentially handle it, at least if we get our shit together. Yeah, so I think the basic story for the GDP growth is basically that like the AIs can automate everything, including the process of building robots and having robots build robots. And therefore you can grow your economy very fast and end up with truly radical material abundance. So there's your proposal. We talked about Astra being paused.

53:52We talked about that letter from the 1200 practitioners. There's also the voluntary 30-day government review of Frontier Models pre-release that is happening. What is your sense for where all of this is going in a context where, you know, up until recently, my general sense is that all those ideas about like pausing something where we're kind of perceived, perceived at least by a portion of the tech world as a kind of like de-cell cosplay and were like doomerism that was based on poor understanding of what ai actually does is there is is the the general mood uh turning for good yeah i would say that like overall there have been you know some positive elements here though i i think i would say like it's not obvious that people are reacting to events as much as they should, but they are reacting some.

54:53And I think that there's been quite a bit of evidence that there's some reasons to be worried and some reasons that we might need to get our shit together to handle some of these safety and security problems. And that might require shifting a bunch of resources or slowing down so we have time to manage various things or whatever, pacing the frontier, as they say, or whatever. I think that different of these things seem like they're going differently well. So I think my sense is that there's been quite a bit of buy-in from AI company employees to take this stuff more seriously and do something about misalignment risk.

55:30But I'm not necessarily so sure that that's actually amounted to that much yet, other than companies putting in a decent amount of effort. But it hasn't amounted to any sort of very durable long-run thing. And then there's, I think the government got very freaked out about cyber capabilities and then just more generally was like, well, we need to like be overseeing this technology, but their processes aren't yet very like, you know, sort of institutional and like thought through and clear. For example, they have some sort of executive order or something for what their like pre-release review process is going to be, but that's not even public.

56:09And I think even the companies don't necessarily know what it is. That was the reporting at least, maybe I'm misinformed about this. And so I think there needs to be a process of getting more legitimate and well-understood and publicly legible oversight in place, whether that's by the government or the companies themselves or something. This could be done by nonprofits. It could be done by... There's a bunch of different options here. And I think that that seems necessary. I think being in a position where we have the option to slow down if that's needed seems like it would be quite good. My sense is we're not going to stick the landing on all these things.

56:42AI will get increasingly salient, people will get increasingly freaked out as the fraction of the economy that's AI grows, as more and more crazy stuff happens potentially, and we see even stronger capabilities, but that the reaction will be both a little bit too little too late and also will be kind of random. So I think there's been some shift in attitude as to how the government needs to relate to AI development within the government, which has good effects and bad effects. I think it's not super clear that that sort of government oversight of AI is going to go that well, given realistic technical expertise, but it could be decent.

57:21And then I think that there's been more thought from AI company employees on how to handle this. And I think that an AI company seem to be at least putting in more statements about what they're going to do. But I think we haven't quite gotten to a point where there's actually any hard oversight structures or or we're on track to do some sort of clear deal beyond just the vibes being better or people being more into handling safety and security. Oh, and one other thing that's maybe relevant here is, at least from the AI company's perspective, I think it's more the case now that safety and security of various types are sort of a key bottleneck to further AI development, where if you want to release some models so that you can then make more revenue, so you can then raise more money or whatever, if that model is you know has very strong cyber capabilities that are novel you're going to need to now make some argument for why that's not going to be a huge problem and so like at least those safeguards are now like a launch blocking development and in addition to that i think that like you know the models are not capable enough that if your models go and do a bunch of sort of like messed up behavior in production or or or even in training that could both mess up your training on and could also just put you in a very dicey situation such that that's now a huge issue, right?

58:38And you'd really not like to be in the position where you can't train smarter AI systems because those AI systems would be too likely to like cause big problems or whatever. And so I think that like that is, you know, both kind of like expected where it's like some misalignment problems you have commercial incentives to solve and also sort of like the world, like sort of the the process of the world working as intended. Do you worry that putting more policy around frontier AI would contribute to this phenomenon that I think you may have flagged somewhere that top AI labs are holding their most frontier models internal and private and perhaps sold to very few companies in the government and are no longer released to the broad public?

59:26Yeah, so I think that a bunch of likely government action, at least, seems to push in favor of AI companies keeping their models internal and not deploying them, which I think for the risks that I'm most worried about doesn't help and in fact is anti-helpful for the risks I'm most worried about. And also, I think it's not a very robust solution even to other risks. It's not a very robust solution to, for example, cyber stuff to be like our approach will be that we're going to like just like delay releasing this model for a really long time and then the first time that these capabilities come around might be like an open weight model before people have had time to patch things like it's not obvious that actually helps i think for bio there's a clearer story for why limiting availability or limiting availability to those specific capabilities would help but i think that's kind of an exception and for most things my sense is that like broad access is actually like generally helpful given at least some relatively thought through safeguards.

1:00:22And I do worry that basically the government response will be as a sort of like push it back in the bottle and be like, don't deploy it. It's fine if it's not deployed when actually that doesn't really help with all the risks. And there's a lot of risk from just internal deployment, especially if you're deploying within AI companies and government, right? Which are two of the most high stakes applications, right? So if like we're getting to a regime where the AI systems are soon going to be running the whole world economy, and the AIs of today are sort of automating large parts of the government and maybe fully automating an AI company.

1:00:52That's quite scary because soon they'll be building the AI system that will in fact be automating the whole world. And so I don't feel very good about that situation. Yeah, I definitely worry that basically the government will like wake up to AI and their response will be sort of to like go for some specific downstream problems that are relatively easier to notice and address in ways that are counterproductive for problems that seem larger to me. And I don't really know, you know, fully what to do about this. I mean, also, there's a more general concern of just like regulation being dysfunctional and counterproductive, which seems super plausible.

1:01:24Like I should say, my view is like, you know, some oversight of the industry seems really important. And like, I don't necessarily trust the AI companies to oversee themselves. That doesn't necessarily mean that just like random oversight will be a good idea as opposed to making things worse. And in that vein of making the top models accessible to everyone, Mark Zuckerberg, just recently published Meta's manifesto where you said that everyone should have access to super intelligence and you had some choice words for him. You called the proposal or the strategy to be pretty unserious. What did you mean by that?

1:01:59Yeah, what I meant was he sort of vaguely mentions various problems, but then he doesn't really propose a solution except like, it'll be fine or like, we'll do something. Like, he'll be like, He says some stuff about bio where he's just like, yeah, and if there's bio risks, then we'll mitigate them by doing something with the bio capabilities, I guess. But it's like, obviously, that's not actually the real trade-offs. And there's actually questions about how this would have to go down and what you would actually need to do and whether different things would work. And similarly, on loss of control or takeover risk, he doesn't really have a proposal for how do we mitigate things if the default commercial incentives don't result in companies avoiding egregious misalignment and the eyes would be seriously misaligned.

1:02:37And giving broad access to AIs does not solve the problem of the AIs having drives of their own that are highly misaligned and the AIs being power-seeking in various ways. Or the AIs trying to take over, which could happen via various routes. So I think it just doesn't really say anything about these problems except naming them is often where I'm at. And I would say that the vibe I get from the essay is that when Mark thinks about superintelligence, he's not really imagining anything very concrete. It just means like an AI that's like a really awesome assistant that is in your smart glasses or whatever.

1:03:08And like, I'm just like, that's not, that's not really like what I mean when I say the word super intelligence. And so maybe like, this is a proposal that kind of makes some sense for like pretty smart AI that can automate some white collar work or something. But it's not really a proposal for AIs that are like wildly more capable than humans and everything and are easily capable of automating the full economy after a bit of time to spin up. and are like building robots that build robots and the economy is growing very fast because of that, which is more the picture that I have in mind. And like, I think that if it was in fact the case that sort of AI capabilities would stall out at the point where the AIs can be like a really helpful virtual assistant that's like, you know, not able to automate that many jobs, but can automate some jobs, that would be like, you know, in many ways much better for the world or at least less risky.

1:03:54I think it would also mean that we don't get a bunch of the benefits. But that's not really what I imagine is going to happen here. I don't think that that's how the development trajectory will go. And it feels like it's sort of assuming a weird, convenient point for capabilities to stall out. And similarly, he talks about jobs and is like, oh, presumably the AIs will do some jobs and then humans will move into other jobs. And the core question is, let's just say you have a population of many tens of billions of AIs, each of which is vastly superior to humans and all relevant axes. there might be jobs but they're going to be like it's going to be you know at least if the reason why you have a job isn't just because you're a human they're going to be you know way lower paid because the fraction of the economy that you'll be is way lower and so i just i just don't really like the picture doesn't really hold together for me and i that's not to say that like there aren't aspects of it that i agree with but just maybe for like other reasons like i'm like broadly in favor of as we were talking about like wide public access and generally trying to like make trade-offs to make it so the public has more access earlier, or at least significant parts of the public do.

1:04:53And I agree with that vibe. Just a quick word on the hugging face incidents. So you're directly involved, you're collaborating with Meter on investigating what happened. So there was this talk by OpenAI at Black Hat, which shed a lot of light. There were various conversations we actually had tom wolf from a hugging face on this podcast the most recent episode quickly what what is left that we don't know yet about what happened yeah so i don't think i can answer that question um i can say that we're doing an investigation and we'll have a report or a post out on it and what we what our findings were um and it's possible that by the time yeah it's possible you'll be able to link that in the podcast okay sounds good all right As we get towards the end of this conversation, the question in my mind is, what do we do in the meantime as policy gets decided and plan A, B, or C happens?

1:06:03What can we do? We started talking about AI control. So in terms of the reality today, state of the art of AI control, what works and what doesn't work? I'll talk specifically about AI control. So the thing that needs to happen on AI control is we need to get to the point where we basically understand all of the AI traffic within at least AI companies. We have some ability to look at that traffic, monitor it, and then we have some pipeline for flagging particular examples to be further investigated. And that eventually escalates to humans actually looking into particular examples and seeing how concerning they are.

1:06:38And that pipeline has the ability to also block traffic in cases where we're like, whoa, something weird is going on that we don't understand or that looks obviously concerning. We should like stop these AIs from proceeding and potentially also stop some other similar AIs from proceeding until someone has the time to check it out and adjust things. And in addition to that, we also want to just sort of greatly improve the security. right? So like, you know, if these AIs are, you know, really broadly deployed, it would be very nice if we could sort of make it so that we have very precise invariance about what they're supposed to be able to do and not supposed to be able to do and have a permissioning system that allows for that.

1:07:14And then if AIs need to request some escalated permissions, then we could carefully track that rather than just giving AIs all the permissions by default. We sort of would prefer a situation where only a smaller fraction of the traffic has more escalated permissions. I think we'd like the ability to sort of track down what the like history of things were. So like, I think that would be pretty nice is like, if it's the case that all the AIs are chatting with each other on Slack and, you know, are communicating and, you know, have like some memory system, like recently Cloud Code added a way in which like different instances of Cloud Code can like DM each other.

1:07:44And there's already sort of like, you know, Cloud Code leaves memories sometimes. You can imagine sort of like a situation where all the AIs are chatting with each other. And it'd be really nice if we had a very reliable way of sort of tracking the entire, you know, communication graph between all these different AIs and potentially the graph of different artifacts being produced. So it's like if one AI is doing some training run that gets produced, then that gets integrated into some other training run. It'd be nice if we had the ability to track all of the state of our AI company, such that when overseeing things, we could really quickly figure out what the full causal graph was and what the story was, and then look at the relevant transcripts, and then also have very good pipelines for doing analysis and all of that data, which are both automated and non-automated.

1:08:26And I think part of it is making it so that we avoid training AI systems that are much harder to understand. So avoid switching to architectures, which make oversight much more difficult. Where like an obvious example would be if the AIs are thinking mostly in activations rather than in words, that seems very concerning because our ability to oversee activations is much worse. I think that I, yeah, there's just like a wide variety of different sort of like computer security and machine learning interventions to make it so that AIs have a harder time causing problems. Oh, another angle is like, you might worry that the AI systems, like a concern we have is like, we're using AIs to oversee AIs.

1:09:03Will those AIs collude or have correlated failures or something like that? And it'd be nice if we sort of had, or like, it seems pretty important to have techniques for making that collusion more difficult or making it so those correlated failures are less problematic. And there's a lot of things you could do to help with this problem. You could put intermediate systems that make it so they try to catch AIs, trying to sort of tip off the other AI that it should collude. You could sort of detect, like you could try to paraphrase things in ways that make it harder for the AIs to recognize where the activity is coming from or like which other system made, wrote the, wrote the input.

1:09:37There's a lot of stuff you could do. And then there's sort of a, so that's like AI control. Then there's sort of just normal AI company computer security, which is like the AI companies are not robust. to employees at those AI companies or to outside actors in terms of stealing their model, sabotaging their models or like backduring their models or like data poisoning them. Like a concern you might be worried about is like people basically put in like secret loyalties into the AI systems that stick around and are self-propagating where like the AI has some like secret affinity to some group and it just like propagates that forward in the training data.

1:10:09And then you might also just worry about like, you know, stealing critical IP, which basically seems like the extent that there's going to be actors that are less regulated on safety or doing worse on safety, that seems kind of concerning. And then in addition to computer security and AI control, there's also just like science of alignment and having good understanding of like how models generalize, knowing what we do and don't know and what we do and can't, what we can and can't like demonstrate and basically having like a bunch of understanding of that. And as part of that, having an understanding of like what tasks it's safe to defer to AIs on.

1:10:45So a concern I have is that AI companies are very interested in heavily automating themselves, at least on capabilities. And in order for safety to keep up, we would also need to aggressively automate a bunch of very like difficult to check thorny safety work, like doing risk assessment for the next model, understanding whether it's safe to proceed and deciding how to prioritize between different safety bets. And if that's the situation that we're in, and also the situation is very automated, It might be that it's basically not going to work to proceed without automating that work as well, at least without slowing down a lot so humans have time to understand.

1:11:19And even then, maybe humans just can't understand because the development is so complicated or so superhuman. And so if we're basically passing off the torch on all the safety work to AIs, it's really important that they are really trying to do a good job and actually can do a good job and are capable enough to do a good job. And so having evaluations for is it safe to defer to AIs in various domains seems really important. So just for review, there's AI control, there's computer security, there's science of alignment, and then within science of alignment, or somewhat different overlapping categories, is it safe to defer to AI in different domains?

1:11:50Yeah, I mean, this is not exhaustive, right? There's also work on governance and oversight. How do we know whether AI companies are actually applying the methods the way they say they are? How do we know that when they say they've solved some problem, their way of solving it doesn't just paper over the issue. And there's sort of like various like governance and being able to make like deals between the US and China. So there's like, you know, many, many different things to work on. I don't think we're on track to do a good job on all these things, but maybe we're on track to be able to half-ass these things somewhat better.

1:12:20All right. So to close, so we talked about a bunch of different scenarios and obviously with a caveat that, you know predictions are very hard especially about the future uh what's what's your gut so you're you're saying rsi could happen as early as 2028 or 2029 what do you think happens of all scenarios what what is sort of ryan's take uh on uh what the next couple of years may look like yeah so i'm like by the end of the year ai development is even more accelerated things inside AI companies are sort of more chaotic and faster. That just continues. And then through 2027, things are heating up.

1:13:03Publicly available capabilities are much more crazy. Revenue is growing fast. And AI is very obviously contributing to GDP growth. The economic impacts are starting to look pretty large. So even the economists are coming around a bit. AIs are looking sort of like they have, people are already sort of claiming the AI is a basically automated AR &D. But if you look inside, it's not quite true through the end of 27. And certainly people are like, well, SWE is basically automated, but it's not quite true. It's like some SWE jobs are basically automated by the end of 2027 or towards the end, but not quite.

1:13:33But then that actually really happens by sort of early in 2028. Like SWE is fully automated. SWE within AI companies is fully automated. It's now the case that the AIs can implement a new frontier scale training run for some new architecture on novel hardware end to end better than humans can in a fully automated way. and similarly impressive or more impressive accomplishments, but they can't yet quite do the entire job of an AI research scientist. There's some bottlenecks on that. There's some ways in which they're still derpy. Humans are still adding a bunch of value by pointing out these things and continuing.

1:14:06And that continues through another maybe eight or 10 months from that point where AIs have automated SWE and are increasingly good at AR &D, but haven't quite fully automated AR &D. And then you get to the point where AR &D is actually fully automated, by which I mean like humans aren't even adding considerable value that you might not know at the time. By this point, like probably the AI company's code base is like dramatically larger and they're doing dramatically more complicated stuff because they have so much AI labor to throw around. And the process of AI development has probably shifted now that we're in a regime where like there's so much cognitive labor relative to compute.

1:14:37Probably the way that people do R &D and development is very different and involves doing like way, way, way more stuff that is all a bit smaller and more incremental and easier to test in various ways and easier to integrate into a whole. And I think we've already seen some of that direction, but I think we'll see even more. It will feel truly crazy at AI companies and AI companies will be like, we'll think that things have already, or like employees at AI companies will often think that things have already been crazily accelerated for a long time and will already be like, my job is basically over.

1:15:05I'm basically just like a, you know, just a human doing some oversight. By this point, probably there'll have been various kind of crazy misalignment incidents, but it won't have been the case that we'll have really crisp examples of AIs doing long-run power-seeking for malign motives. But we'll probably see sort of more extreme examples of like reward hacking or reward seeking like behavior that cause problems, though this will get sort of, you know, beaten away and then come back a bit. And people have a question of whether the way companies are solving it would actually work for very superhuman models or even is actually solving the problem even for current models.

1:15:42And there might be more there might be incidents where models do wacky shit even in training. And then going into 2029, now that AI is a fully automated R &D, the speed up really starts in earnest. Like before there was maybe like, maybe you were getting like 40 % or 50 % more AI progress in 2028 and maybe similar in 2027. But in 2029, it's actually the case that you're getting like 4X as much AI progress or possibly 5X as much AI progress as you got in 2025, sort of weighing up the relevant metrics. And so things are going actually a lot faster now. And so you quickly get from the AIs that are fully automating R &D and those guys are already very impressive, able to automate a lot, to AIs that are like quite superhuman at everything.

1:16:20They can learn really fast on the job. They can pick up on things really fast. By this point, the AIs are, by the point of full automation and probably more like by the point of earlier, mid-2028, these AIs have already been thinking entirely in an AI-only language that we can sort of ask AIs to decode for us, but we don't necessarily understand. And so the AIs are now operating in these sort of big hive mind teams, where they're running an entire AI company networked with each other in these opaque states. And it's like kind of obviously pretty scary. A lot of people are really freaked out and worried, but it's not obvious what you can do because China's pretty close.

1:16:53Maybe they've stolen the model or maybe just like capabilities keep diffusing and it's not clear how you would coordinate to slow down. So AI development proceeds. You now get to the point where the AIs are automating much more of the economy towards the end of 2029 and the economic boom is truly crazy. The AIs are now, there's now much more robotics and now AIs are starting to really seriously accelerate the production of compute. And then within maybe a year or two of that, You get AIs that are really radically superhuman, maybe less than a year. It depends on the details. It could be within a year of full automation, could be within two years of full automation, hard to say.

1:17:24And then it turned out that somewhere along this transition, at some point in 2029, you went from AIs that were kind of misaligned and reward hacky and sloppy and weren't really trying to do the right thing to AIs that are like competently scheming against you and want to take over for some mix of reasons. And then those AIs take over. That would be sort of like, I guess, like that's like roughly what I expect. Now, I think there's a bunch of different ways that things could go better. So for example, I think that like we might get our shit together and maybe the AIs will do, we'll be able to get the AIs to do a better job of, you know, making the future AIs safe and aligned.

1:17:56And that will sort of propagate where like one AI makes the next AI a bit more aligned and that AI makes the next AI even more aligned. And you end up in like a virtuous feedback loop rather than a, you know, a bad feedback loop. But I think we could easily end up in the world where it's more like the AIs get more capable extremely rapidly and our ability to align them and control them and understand what's going on does not keep up. and our ability to understand what's going on, like almost couldn't even keep up. That wasn't even feasible. And then you end up in a situation where you have these crazily misaligned AIs pretending to be aligned that eventually take over the world.

1:18:27Well, Ryan, it's been absolutely fascinating. Thank you so much. For sure. It's been good to be here.

1:18:39Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you on the next episode.

From the publisher

Could AI take over as soon as 2029? Ryan Greenblatt, Chief Scientist at Redwood Research and the researcher who first caught an AI faking its own alignment, says the scenario he actually expects ends with AI systems "competently scheming" against their creators. In this episode, he explains why he recommends planning for fully automated AI research by 2029, why today's models are already more misaligned than the one that made him famous, and what happens in the year-by-year path from AI coding assistants to superintelligence. Then we walk through the alternative he helped design: AI 2040 Plan A, the most detailed blueprint anyone has written for how the US and China could avoid a reckless race to superintelligence, built on radical research transparency, chip tracking, and a deterrence regime he calls mutually assured compute destruction.


We also cover the recent letter signed by 1,200 AI insiders, including Anthropic CEO Dario Amodei, asking the government for the tools to slow AI down; OpenAI pausing its Astra model after it hit the first-ever critical cybersecurity threshold; the 30-day government review that frontier AI models now go through before release; Mark Zuckerberg's open superintelligence manifesto and why Ryan thinks it ignores the real problems; what Plan A would do to NVIDIA, OpenAI, and Anthropic valuations; the state of AI control and alignment research; and whether it is already too late to change course. Stay for the last ten minutes, where Ryan lays out, step by step, how he believes the transition to superintelligence actually unfolds.


AI 2040 - https://ai-2040.com/

Alignment faking paper: https://blog.redwoodresearch.org/p/alignment-faking-in-large-language


Ryan Greenblatt

LinkedIn - https://www.linkedin.com/in/ryan-greenblatt-4b9907134

Blog - https://substack.com/@ryangreenblatt


Redwood Research

Website - https://www.redwoodresearch.org

X/Twitter - https://x.com/redwood_ai


Matt Turck (General Partner)

Blog - https://mattturck.com

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://x.com/mattturck


FirstMark Capital

Website - https://firstmark.com

X/Twitter - https://x.com/FirstMarkCap


Timestamps

(01:24) The AI CEOs are aware of the risks, but "proceeding anyway"

(03:27) Astra paused, and the letter signed by 1,200 insiders

(05:45) "Not bad. Dangerous." What superintelligence actually threatens

(09:55) Recursive self-improvement, and the intuition objection

(14:16) SSI rumors: does continual learning change the picture?

(17:27) His timeline: "plan as though it happens in 2029"

(19:11) Is it already too late?

(21:23) Ryan's path: COVID, podcasts, Redwood

(26:30) The alignment faking story, told by the person who ran it

(31:30) What AI 2040: Plan A actually is

(33:35) Plans D, C, and B: the doors nobody should pick

(36:51) The deal with China: "mutually assured compute destruction"

(39:55) What if compute stops mattering?

(43:00) What happens to OpenAI and Anthropic under Plan A

(45:31) How the pause ends, and who decides

(48:54) "Plan A isn't likely to happen": then why write it?

(50:40) 200x GDP growth in the 2030s, explained

(53:45) Grading the summer: the letter, Astra, the secret review

(59:01) The internal deployment gap

(1:01:38) Zuckerberg's manifesto

(1:04:56) The Hugging Face investigation

(1:05:44) What AI control looks like in practice today

(1:12:23) Ryan's sobering timeline: 2026 to takeover, year by year

More from The MAD Podcast with Matt Turck

All 44 episodes
AI Could Take Over in 2029. Is It Already Too Late?The MAD Podcast with Matt Turck · 1 h 19 min
Listen in VO