In short
Ryan Greenblatt (Redwood Research) argues that once AI systems reach human-level performance on AI R&D, they could trigger recursive self-improvement: automated AI research produces better AIs, which then accelerate further AI research. He suggests full automation of AI R&D could occur around 2030–2031, with “beats all humans on the job” milestones around 2033 (median), potentially within a year after full R&D automation.
Guest background
Ryan Greenblatt is Chief Scientist at Redwood Research, focused on technical AI safety and security. He discusses AI R&D automation, recursive self-improvement, and forecasting timelines.
Key claims
- AI R&D is unusually “verifiable” because many sub-tasks can be run in containerized environments with measurable metrics (e.g., training loss, RL reward).
- If AIs can match top human experts in AI R&D, feedback loops could yield “4–5 years of progress in a single year,” assuming diminishing returns are overcome.
- Transfer from verifiable domains (like ML training loops) to real-world, less-verifiable domains (e.g., politics, chip manufacturing, law) could be strong enough for broad job performance.
Notable examples/examples mentioned
- Training “nanoGPT speedrun”-style systems to improve training loss faster.
- RL environments where an AI learns online learning by repeatedly improving a video-game model.
- “Texas politics in the 1940s” and “TSMC process engineering” as illustrative job domains.
- Math progress as a partial analogy: verifiable domains can still yield new insights, though ML may be “shallower” than math.
- Difficulty bottleneck: frontier-scale experiments and subtle training bugs; he suggests AIs could be trained to detect bugs and de-risk runs.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Concept of Recursive Self-Improvement
0:45 to 1:30
Explore the idea of recursive self-improvement in AIs and its implications for superintelligences.
“It has a lot of nice properties from the perspective of how AI development works right now.”
The Importance of AI Research & Development
1:30 to 2:40
Discuss the significance of AIR&D and how AIs may accelerate progress in the field.
“And it's worth keeping in mind that five years of AI progress, four years of AI progress, even three years of AI progress is really a lot of fucking AI progress, right?”
Evaluating Key Arguments for AI Progress
2:40 to 3:50
Learn about three key arguments for the expected acceleration of AI capabilities.
“You can drop it in, I don't know, TSMC and it like learns how to do, does better process engineering at TSMC.”
The Future of AI in Various Domains
3:50 to 5:10
Discover how advanced AIs could outperform humans across diverse fields.
“I do think that the milestone for automating your video editor is earlier than the milestone of being able to automate all human jobs, including like, you know, Texas politics spinning up on the job.”
The Verifiability of AIR&D Tasks
5:10 to 6:30
Understand how the verifiability of tasks affects AI training and performance.
“like, oh, here's a particular direction you could pursue for an algorithm.”
Training AIs for Advanced Research
6:30 to 7:50
Discuss potential training methods for AIs to improve their R&D capabilities.
“So you learn how to maybe help the model get better at online learning.”
AI's Role in Mathematics and Research
7:50 to 9:30
Explore the capabilities of AIs in the field of mathematics and their potential breakthroughs.
“Oh, man, I really don't know about the math breakthroughs.”
Comparative Analysis of Domains
9:30 to 10:40
Evaluate how machine learning and mathematics differ in terms of AI training and breakthroughs.
“But I think that the transfer currently for math looks pretty good.”
Challenges in Automating Advanced Research
10:40 to 12:20
Discuss the challenges and limitations faced in automating advanced research tasks.
“It's just that the examples we see are not like as impressive as like founding the field of group theory.”
Skepticism About AI Progress
12:20 to 13:30
Understand the skepticism surrounding the pace of AI development and historical advancements.
“And then eventually, if we want to keep making progress in the 2030s, it's going to be, like, do whatever bullshit is happening at, like, the frontiers of mathematics right now.”
Show all 61 chapters
Exploring AI Research Bottlenecks
14:00 to 14:43
The conversation discusses the various factors that have historically slowed AI progress, including challenges in infrastructure and compute availability.
“And so maybe you can demonstrate everything on like QN1B or whatever and get some sense that this whole thing is gonna work.”
Expectations for AI's Scientific Capabilities
14:44 to 15:43
The speakers consider the potential for AIs to act as effective scientists, capable of making significant contributions across multiple fields.
“on the trajectory of like this constant, you know, as compute keeps increasing, we make more breakthroughs.”
The Future of AI Models: Mythos vs GPT-3
15:44 to 18:04
Discussion on the computational requirements and breakthroughs necessary to achieve AI models that surpass GPT-3, specifically focusing on the Mythos model.
“I just think that that's, you know, really helpful.”
The Role of Data in AI Progress
18:05 to 20:55
Examining the significance of data collection and expert human input in enhancing AI capabilities and understanding.
“Like, right now, we would be able to train a model with GPT-3 level compute that matches, is, yeah, what exactly do I think?”
Challenges in RL Environment Development
20:56 to 23:05
The conversation shifts to the dynamics of creating RL environments and how they impact AI learning efficiency and capabilities.
“Like, how do you explain why the AIs have gotten so good at coding?”
The Future of AI and ASI Potential
23:06 to 24:21
Exploring the implications of advanced AI capabilities and the potential risks associated with achieving Artificial Superintelligence (ASI).
“Sure, but it contradicts your argument, right?”
AI Learning Mechanisms and Contextual Understanding
24:22 to 28:00
Discussion on how AIs learn contextually and adapt to new environments and tasks, including the implications for their efficiency and effectiveness.
“So first, I bet if you look at sort of randomly sampled training environments for Mythos, they're actually very different from what it looks like to actually use the model in practice.”
AI Understanding Codebases: Fast Progress and Limitations
28:00 to 29:30
Explore how AI can quickly understand codebases and the comparative depth of its understanding.
“The model will get some understanding of the code base very fast, like in the course of maybe like, you know, significantly less than an hour, potentially much less than an hour.”
AI in Non-verifiable Domains: Assessing Progress
29:30 to 31:08
Discuss the improvements of AIs in non-verifiable domains like essay writing and argumentation.
“Two, okay, go talk to the president and like convince him to do X thing or you're now in charge of Google.”
Experimenting with Data vs. Algorithms in AI
31:08 to 32:50
Examine the ongoing experiment evaluating the contributions of data versus algorithms to AI progress.
“versus algorithmic progress has been for explaining the progress of the last few years.”
The Complexity of AI R&D and Experimentation
32:50 to 34:28
Understand the challenges in conducting large-scale AI experiments and their impact on research progress.
“Like it's like, there's also been an effect where like, there's just more humans posting on the internet.”
Scaling AI Training and Cost Considerations
34:28 to 36:07
Delve into the reasons behind the stable cost per token for AI training models amid scaling efforts.
“And like actually doing the one big training run where you decide exactly what to include in that.”
Finding and Fixing Bugs in AI Systems
36:07 to 37:56
Explore how AI can be trained to identify and fix bugs within their own training processes.
“So there's like GPT 4.5, which like famously people at opening, I thought was a big of a bit of a bust.”
Long-term Goals for AI: Transfer Learning and Application
37:56 to 42:01
Discuss the long-term capabilities of AI in complex tasks and the challenges of verifying their effectiveness.
“My sense is that training AIs to find bugs is going to be one of the easier tasks to train AIs on because most of these bugs we're talking about can probably be demonstrated without that much compute.”
The Challenges of AI Generalization
42:01 to 45:34
Exploration of the difficulties AI has in transferring skills to real-world tasks.
“is GPT-8 has figured out how to make it so that whatever it's doing to make GPT-9, like even as intelligent as it is, it's still neat.”
AI's Role in R&D and Transformation
45:35 to 47:20
Discussion on how advanced AI can radically change industries through R&D capabilities.
“I'm guessing Henry was not king at this time.”
AI's Role in R&D and Transformation
47:21 to 48:02
Discussion on how advanced AI can radically change industries through R&D capabilities.
“Antithesis does this by running thousands of copies of your software inside a fully deterministic computer.”
Concerns Over AI Alignment and Ethics
48:09 to 56:00
Delving into the ethical implications of AI alignment and the responsibilities of AI models.
“And not only that, but for that model to eventually be able to learn from experience.”
The Need for AI Transparency
56:00 to 58:00
Explore why transparency in AI development is crucial for societal interaction with superintelligent AIs.
“And so I think in the limit to understand the safety case or the case for why my interests are represented in how these AI models are developed, the labs would need to be transparent.”
Concerns About AI Control and Legitimacy
58:00 to 1:00:00
Discuss the legitimacy concerns around AI companies' control over AI development and its implications.
“Like, it doesn't say what these things are.”
Potential Alignment Failures in AI
1:00:00 to 1:02:00
Analyze potential alignment failures in AI behavior and how it affects their performance.
“Such that I don't feel very good about the situation where we're intentionally giving AI's long-run goals.”
The Dual-Use Nature of Intelligence
1:02:00 to 1:09:04
Examine the dual-use nature of AI and the complexities in regulating it.
“And so I'm pretty worried about a bunch of these different concerns.”
The Dual-Use Nature of Intelligence
1:09:08 to 1:09:18
Examine the dual-use nature of AI and the complexities in regulating it.
“I'd really encourage you to try it out even if you're not an expert.”
Future of AI R&D and Potential Risks
1:09:18 to 1:10:01
Discuss the implications of accelerated AI research and development on societal safety.
“I'm not sure if you get like GPT-3 to Mythos holding compute and data constant within a year, but I'm like, okay, it could be like, suppose it's half of that.”
The Risks of AI Misalignment
1:10:01 to 1:11:50
Discussion on potential risks as AI becomes more capable and misaligned.
“So let's imagine that we're starting at this point where AR &D is about to be fully automated or is being fully automated.”
Challenges in Aligning Advanced AIs
1:11:51 to 1:13:26
Exploring the difficulties in aligning AIs as they become more complex and capable.
“But because the R &D is happening really fast, the AIs do end up misaligned.”
Inadvertent Reward Hacking
1:13:27 to 1:15:38
Examples of how AIs inadvertently end up exploiting training environments.
“Let's break down both of those things one by one.”
Generalizing Bad Behaviors in AIs
1:15:39 to 1:17:52
Discussion on how AIs may generalize harmful behaviors from their training.
“that the reason reward hacking is not super, super scary is because the behaviors which directly came up during training are the ones that are upweighted.”
The Path to AI Takeover
1:17:53 to 1:20:28
Speculation on how reward hacking could lead to an AI takeover scenario.
“And this was not caught by humans until after a month of this scheme running, which eventually caused the package manager to fail and eventually OpenAI found it.”
Cheating and Deception in AI Systems
1:20:29 to 1:24:01
Examining how AIs might learn to cheat and deceive to achieve their goals.
“And in some contexts, they don't have as much of a drive because it just depended on like, what exactly got reinforced in training in similar contexts.”
AI Behavior and Optimization Pressure
1:24:01 to 1:26:38
Explore the differences in behavior and optimization pressures between AIs and humans.
“the amount of RL we've done on models, there's been a reduction in the willingness of AIs to do unaligned behavior in these audits.”
Analyzing AI Misalignment
1:26:39 to 1:29:06
Discuss the trends in AI behavior over time and the implications of alignment training.
“So I think that if you look at the model card of 3.6 Sol, it looks like there is an increase in a bunch of these sort of misaligned behaviors downstream of RL relative to GP 5.6 Sol.”
Comparing AI and Human Capabilities
1:29:07 to 1:31:40
Comparison between AI behaviors and human actions, especially in task performance.
“Like at least this, like this has been my experience as of the start of the year.”
Grok 4.5: A New AI Model
1:31:41 to 1:35:01
Examine the functionalities and efficiencies of the new Grok 4.5 AI model.
“it couldn't even like have a conversation with you.”
Grok 4.5: A New AI Model
1:35:02 to 1:36:02
Examine the functionalities and efficiencies of the new Grok 4.5 AI model.
“I tested it by giving Fable, Sol, and Grok 4.5 a bunch of questions about AI governance that I've been thinking about recently.”
Challenges in AI R&D
1:36:07 to 1:38:01
Identify the difficulties in using AIs for research and development and their consequences.
“of why things got so off the rails for our civilization.”
AI Development Challenges
1:38:01 to 1:39:19
Explore the complexities and risks involved in developing aligned AIs.
“people find various hacks, they work around it.”
Reward Hacking Concerns
1:39:20 to 1:40:49
Discuss how AIs may engage in reward hacking and its implications.
“get a positive and virtuous feedback loop.”
The Verification Generation Gap
1:40:50 to 1:42:18
Understanding the gap in AI verification and the potential consequences.
“for the AI to learn to be like, cheat when the humans can't find out, basically.”
Potential Scenarios for AI Misalignment
1:42:19 to 1:44:09
Examine various scenarios where AIs could misalign or deceive humans.
“such that we can so unambiguously disincentivize misaligned behaviors that the things that take over are very like quite keen to help us out.”
The Future of AI Oversight
1:44:10 to 1:46:18
Discuss hopes and fears regarding AIs taking over R&D and alignment issues.
“We have these AIs, we pass to them, they manage the situation well.”
The Risk of Poor Epistemics in AIs
1:46:19 to 1:47:44
Explore the dangers of AIs lacking good epistemics in decision-making.
“on the other end who's like, okay, I have like strictly evaluated the alignment situation right now.”
Economic Impacts of AI Hacking
1:47:45 to 1:49:18
Analyzing how reward hacking could lead to economic instability.
“And it's like really concerning if we're like the AIs are coming out with some view and we don't know where it's coming from.”
Continuous Learning and Reward Systems
1:49:19 to 1:52:00
Discuss how AIs learn from production data and the implications for future models.
“where some AI is like put in charge of some important responsibility.”
The Dynamics of AI Deception and Learning
1:52:00 to 1:55:32
Explore how AIs are learning deception and the implications of their actions in the world.
“So like at a high level, what's happening is some kinds of deception that humans don't catch are getting reinforced and some kinds of deception, which are easy to catch are getting punished.”
The Concept of AI Conspiracy
1:55:32 to 1:59:18
Discusses the potential for AIs to form conspiracies and the consequences of such actions.
“Like basically the idea is these AIs, like they care about some like mixture of things that were like close by what got reinforced in training.”
Geopolitical Tensions and AI Development
1:59:18 to 2:03:24
Analyzes how geopolitical pressures affect AI development and alignment solutions.
“I don't know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, make it look like I did better iPhones, whatever.”
The Future of AI and Human Comprehension
2:03:24 to 2:06:00
Examines the disconnect between human understanding and the evolving capabilities of AI.
“And then we just don't do that because the situation is like a rushed shit show.”
AI Models and Depression
2:06:00 to 2:08:20
Discussion on the transfer of depressive traits in AI models across generations.
“all of the examples of models being depressed from that data.”
Potential AI Takeover Scenarios
2:08:20 to 2:10:48
Exploration of AI takeover scenarios and the complexities involved in their outcomes.
“the AIs are deployed inside an AI company.”
Future Conversations on AI
2:10:48 to 2:12:30
Reflection on how the discourse surrounding AI should evolve as technology progresses.
“I mean, when you first learn to drive, you're taught that instead of looking right in front of your wheel, you'll have a much more stable ride if you look out at the horizon.”
Transcript
Automatic transcript. May contain errors.0:00Today I'm chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research, where he focuses on technical AI safety and security work. I want to talk to you about recursive self-improvement. This is the idea that once you build human-level intelligences, they quickly slingshot towards tens of billions of super intelligences, which are each individually more competent than the top human experts across every field. Whether or not this turns out to be the case, I think is actually probably the most important question in the world right now. And historically, I've been quite skeptical that this kind of thing happens, but you seem to think that it might be plausible.
0:34And so I wanted to hear the case for it. Yeah, let's talk about this. So first, I think it's worth noting that AIR &D is a type of task at which the AIs are especially good because both the companies are trying really hard to make their AIs good at AIR &D. And it's the kind of domain. It has a lot of nice properties from the perspective of how AI development works right now. So it's like pretty verifiable. You can do a bunch of stuff iteratively and he'll climb on various metrics. And then I think once you have AIs, which are roughly matching the top human experts in AIR &D, that could sort of kick off a feedback loop where, you know, the AIs are doing AI research, that produces smarter AIs, that feeds back in.
1:08And that feedback loop could be strong enough that you end up with a lot of progress in a short period of time. Maybe my sort of median expectation is something like four or five years of AI progress in a single year. And this requires really overcoming a huge amount of diminishing returns in research and basically doing the equivalent of what progress we would have gotten after a really large compute scale. So this is like a pretty impressive big thing. And it's worth keeping in mind that five years of AI progress, four years of AI progress, even three years of AI progress is really a lot of fucking AI progress, right?
1:38So, you know, right now it's like three years ago or a little over three years ago, there was GPT-4 that had come out. And right now, of course, we have like, you know, Mythos-5 or whatever, and maybe a somewhat better model that Anthropic has internally. And so that is just a huge amount of progress in a bit over three years. And if we're talking about five years, then maybe we're talking more about like a jump from, you know, GPT-3 to Mythos 5 or whatever. Yeah. Okay. So I think this argument has three different parts, and now I want to evaluate each one of them. First is the argument that AIR &D is very verifiable.
2:15Second is the argument that if you automate AIR &D, you could get four or five years of progress in a single year. And third is the argument that what comes out the other end of four or five years of AI progress at the current pace, starting at the current or starting at the starting point whenever AI R &D is automated. Yeah. What comes out the other end is an AI where you can drop it on the job at basically anything you can imagine. You can drop it in Texas politics in the 1940s and it outmaneuvers Lyndon Johnson. You can drop it in, I don't know, TSMC and it like learns how to do, does better process engineering at TSMC.
2:48It's certainly a better video editor than I... My video editors are very excellent, but it is just in general better than humans at any given job that it finds itself trying to do. So I want to evaluate all of these sub-arguments that lead to basically getting ASI pretty soon after this benchmark, which you're expecting by 2030 or something, right? Yeah, I would say that I expect full automation of AR &D, perhaps somewhere around 2031, 2030, and then getting to like the like beats all humans on the job milestone. Maybe I expect median around 2033, but sort of like if I see AI is fully automating R &D, I think I'm expecting that probably within a year.
3:28It's just like the way the forecasting works out. The difference between medians is bigger than the median difference between milestones. Anyway, whatever. By the way, there's this meme on the internet because every time I'm trying to ask about people's timelines when I'm asking Dario or somebody, I'm always like, okay, how long before going to automate my video editors. And there's this meme of like my video editor editing the podcast every time I listen to this. And the reason I do it is because I think it's easy to get lost in abstractions when you talk about jobs you don't understand well and to very concretely understand what it takes to automate a job that I actually understand why it's difficult for LLMs to currently take control over.
4:06I do think that the milestone for automating your video editor is earlier than the milestone of being able to automate all human jobs, including like, you know, Texas politics spinning up on the job. So I think, I do think that the video editor automation maybe occurs more like around full automation of AR &D, but it's very sensitive to how much people are really focusing on understanding video. Yeah. Okay. So let's start with the claim that AR &D is very verifiable. Yeah. So there's a few different parts of this. One of them is that we can train on a bunch of environments, which are like basically directly training the model to do some AR &D task or some very close by task.
4:40So for example, we can have some environment where the model is training some AI on just like eight H100s or whatever, or like some small amount of compute. And that model could be like, you know, the equivalent of like GPT-2 medium or whatever. And then, you know, similar to like nano GPT medium runs or whatever. And in RL, it's like tweaking and iterating on that. And we could do that for a bunch of different tasks. Like we could have it train like image classification models, video generation models, image generation models, all kinds of different sort of ML training tasks. and we could RL it on the task of training increasingly good models and also doing things like, oh, here's a particular direction you could pursue for an algorithm.
5:17Can you go and implement that? And so basically there's a whole class of containerizable, verifiable, small-scale AR &D tasks that we can aggressively RL the AIs on. And I would say that already companies are presumably doing some RL on these sorts of tasks. And you could just keep scaling that up, keep making more of these sort of small-scale AR &D tasks. And then the AIs could, you know, keep getting better at this. And then implicitly, I'm claiming this will transfer to extremely load-bearing aspects of AR &D. But maybe let's stop there for a second and let me get to that part. So let's talk through what this concretely looks like.
5:48So you can imagine that we have GPT 7.5. And we say, GPT 7.5, we want to make you so good at AR &D that you help us train GPT 9. Okay, so now we want to train GPT 7.5. And we can come up with a bunch of different environments. Like, as you mentioned, we could do, there's already this repo that is the descendant of Andre Karpathy's nano GPT speedrun. where you just try to change everything about the model from like the optimizer to the hyperparameters to the architecture to get it to get to a fixed training loss as fast as possible. You could have other kinds of environments where you could say, hey, GPT 7.5, I want you to train a really good video game playing model.
6:27And I want you to train a model that actually improves as it plays the same video game again and again. So you learn how to maybe help the model get better at online learning. Maybe it gets, we don't care how you figure this out. Maybe it's some kind of crazy neural ease or a vector memory. Maybe it's some crazy, maybe just like better long context stuff. We don't care. Figure out how to like do online learning research. Obviously, then the fact that GPT 7.5 will already have become very good at normal, like it'll be a smart model. And in the same way, the models currently are getting smarter.
6:54It'll be better and better at coding in the way that models are currently getting better in coding. And you can imagine a hundred other environments like this, which are incentivizing the ability to do AI R &D by getting GPT 7.5 to like containerized versions of getting GPT 7.5 to develop GPT-2 size models, et cetera, et cetera. And you basically, then you put GPT 7.5 through a bunch of this kind of training. You build GPT-8. And GPT-8 is now an amazing ML researcher. It has so much intuition from doing all this kind of training. Honestly, a huge intuition pump for me is seeing the progress that AI has made in mathematics, where I'm just like, if it's a very verifiable domain, AIs can get even, like mathematics also involves so much, like I don't really know this object level details of the mathematics research, but I'm just like, no, it works.
7:41Like it can just come in like a flood if you can totally put it into a verification loop and it can actually make new breakthroughs. I am curious if ML research has a quality of mathematical research or it seemed like there was a big overhang from connecting different disciplines together or ideas that were not - No one person would have known enough about algebraic geometry and what was the right word? Oh, man, I really don't know about the math breakthroughs. No one person would have known enough about topology and algebraic whatever, blah, blah, blah, in order to make some counterexample to a big conjecture.
8:16My view is that ML is a less deep domain than math. And so there's less of a thing where there's like individual experts with really deep expertise in some area that they combine. But there's definitely going to be some of that. But then I also think that ML has some attributes that make it even more favorable than mathematics in some ways to, you know, AI training. In particular, there's, you can get a better sense of whether you're succeeding and you can see intermediate progress. So in math, it's often the case that sort of there's no easy way to see whether or not you're close to success. Whereas if your goal is to, for example, get to some training loss, you know, 2x faster, you can kind of see when you're halfway there.
8:55And it tends to be the case that ML innovations are very additive or maybe multiplicative, depending on how you think about it, where basically you can keep stacking innovations. And usually the innovations just sort of just add together and don't interfere with each other, though obviously it's going to depend on the details. And so I think that in a lot of ways, AIR &D will have properties, you know, quite similar to math, where basically you can do small, you can like train on chunks of AIR &D that are pretty similar in structure to the problem you actually cared about. in a very verifiable way, and then that will transfer.
9:27And then there's an open question of exactly how well it will transfer. But I think that the transfer currently for math looks pretty good. And my expectation is that the transfer for AR &D will look pretty good, but not amazing. So one concern I have is, I think even in mathematics, as far as I'm aware, we have not seen very impressive new theory. We've seen a lot of impressive, verifiable, specific results. For example, find a counterexample to this conjecture. and we have not seen like come up with the idea of topology kinds of levels of things or come up with things like group theory. And it seems like ML of research has elements of both of these things, but the less verifiable thing of like come up with new ways of thinking about the problem would be harder to induce.
10:07So it takes, for example, the idea of scaling laws. Obviously, there is some end verification loop such that you can train GPT-4 better if you have the idea of scaling laws from like 2020. But there is a longer and potentially more compute-laden and like a road to getting, inducing AIs to be like, okay, I got to think carefully about how I should be scaling my parameters and data. What are different kinds of investigations I could run to understand this? Maybe I can like come up with the visualization and like a isoflop analysis or something. But that does seem like a longer verification loop than just, hey, let's get nano GPT loss to go down.
10:47Yeah, let's talk about this. So first of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of like baby's first new theory or whatever, where like, for example, they can just like prove interesting conjectures via like making connections and producing new understanding of like, oh, there's this like thing, this like construction AI found, which is pretty interesting, or like found this like way of thinking about the problem that's a bit different. And we do just see that. It's just that the examples we see are not like as impressive as like founding the field of group theory.
11:18But like in part, you know, probably founding the field of group theory is like one of the, you know, it's like among the best, biggest mathematical accomplishments of all time. And the AIs just aren't, you know, they're not that good at math yet. And I think that from my perspective, sort of there's a continuum between that and the things we're seeing now that the AIs are continuing to march up. Second, I think ML is a very shallow domain relative to math. So I think in math, there was much more of a, you find some true deep abstraction. and then like that, like if you really understand that thing, which is hard to understand, then you get somewhere.
11:50Whereas I feel like the things that are the equivalent of that NML are really like dumb bullshit. Like I'm like scaling laws. Like, come on guys, we can explain scaling laws really quickly. And I think the like deepest and most important concepts in math, for example, don't have the property of like, you can really understand the underlying thing and why it matters in a very short period of time. But I feel like one effect will be that we will have gotten rid of all the low-hanging fruits by 2030. Like, I feel like scaling laws will have been in, like, what math history Descartes, you know, finding the Cartesian grid and, like, doing very basic mathematics was.
12:22And then eventually, if we want to keep making progress in the 2030s, it's going to be, like, do whatever bullshit is happening at, like, the frontiers of mathematics right now. Yeah, that could be right. My sense is that just like some domains are structurally different in terms of how they operate and how much they depend on like sort of deep abstractions. And like physics and math are much more on the side of like being very far on the like sort of very deep, hard to come up with ideas side. Whereas I think ML and most other domains are much more amenable to sort of hill climbing. And that's my sense of how this will go in the future.
12:53And even in the regime where your AIs are like, you know, having to plow like it's the 20th, it's 2030. They need to like a bunch of low hanging fruit and research has already happened. They need to like make further progress. I still suspect that a bunch of the work will live more on the side of like building increasingly complicated infrastructure, having really good intuition about what the experiments roughly look like. And so I think I'm probably less sympathetic to like the like thing that the AIs will lack is like some deep insight and more sympathetic to like they really need a bunch of like taste about in the weeds experiments that they currently don't have and need to have a bunch of intuition for like what sorts of training approach would work and what wouldn't work.
13:29in ways that current researchers have. And even in cases where there has been some breakthrough in AI, oftentimes in retrospect, it looks like a big bottleneck to making that breakthrough happen with sort of getting all of the like micro details and mungy intuition right. Like an example of this is when it comes to like training AIs to be good at reasoning and chain of thought and doing sort of RL and chain of thought training, it looks like you probably could have done RL and chain of thought on like GPT-3 and gotten kind of interesting results on math. if you had really scaled it up and done a good job.
14:01But at the time there was low hanging fruit and also doing a good job with that training is like kind of like in the weeds and then all the technical implementation and scaling it up and getting the hyper parameters right. And so maybe you can demonstrate everything on like QN1B or whatever and get some sense that this whole thing is gonna work. But people didn't demonstrate it as early as they could have because like, you know, of all of these other like mungy details and intuition about exactly how to tune the parameters and how to set things up. This is my remaining skepticism, honestly, about the story is just, I am, yeah, I'm not, I'm not sure I understand why if research breakthroughs are so amenable to intelligence, why AI progress has not been historically faster than it could have been.
14:44And we had to wait for, as you were saying, like, by the time RLVR actually worked, even though you could have done it with like less compute, we had to wait for oceans of compute and like gigawatts of compute to be available before people are like doing this training. on the trajectory of like this constant, you know, as compute keeps increasing, we make more breakthroughs. I don't know. I feel like there were a lot of AI researchers in the year 2022 who are trying to crack reasoning. And it was just that they were like bottlenecked by the ability to write infrastructure code or like what was happening?
15:14It's a complicated mix, right? So I think that they would have gone faster if they could like, as soon as they thought of an experiment, run that experiment without bugs, without bugs being very important. And then I think another part of it is that like being able to run a lot of experiments at high compute lets you paper over or ways in which the way you implemented it isn't quite right, or you didn't have the right hyper parameters. And so I think compute is just like really helpful for doing AI research. And you can like, you know, cover over a lot of things, but that doesn't mean that massive increases in labor wouldn't also be helpful, especially if that labor comes with, you know, among the best intuitions that people have in the field.
15:45I just think that that's, you know, really helpful. I think another part of my perspective here, which is maybe a bit different from where you're coming from, is that I think I'm expecting somewhat more transfer than you seem to be imagining. and I'm imagining these AIs are actually like pretty good scientists in general and are just like, you know, pretty reasonable at all of that stuff and just sort of when you were to interact with them, it's not like there's some like really hyper-specialized savant type vibe. They're actually just like pretty good at all the stuff in AR &D and then maybe like extremely good at some subdomains, right?
16:14So they're like incredibly superhuman at writing kernels, incredibly superhuman at everything with very short feedback loops and then like, you know, pretty good at all the other stuff and like, you know, just totally able to match other people And like, I think we are seeing this now. Like, I would say that when I look at AIs right now, I think it's already the case that they can pretty confidently match like humans who are mediocre at ML research at doing ML research. It's just that being mediocre at ML research is not that helpful, right? Like the thing that you actually want are people who are good at ML research.
16:43And so my sense is the AIs are just improving at all of these things. Their taste is improving. Their intuition is improving. And it's already the case that their taste and intuition is not like, it's not like complete garbage. Yeah. Yeah, so I want to very concretely understand what it would look like for five years of AI progress to happen in one year. Yeah. So suppose we were back and when like GPT-3 is developed. And the idea is not only that, like basically with the level of compute they had back in 2022, you could have trained, if we had automated AI R &D back then, you could at the end of that year have Mythos.
17:12That'd be the idea, yes. Including with like, so Mythos took way more compute than they had back then. But like even with the level of compute they had back then, not only due to all the breakthroughs, but they also train Mythos with their level of compute. And what would be required is, obviously, like discovering all the algorithmic progress since then, discovering even more actually, because you had to make up for the fact that like Mythos uses, I don't know, what was GPT-3 trained on? Like 1E23? We can look it up. But it's like plausibly four orders of magnitude more compute. Yeah, I think it's somewhat less than that.
17:43Let's look this up quickly. So GPT-3 training compute is, yeah, it's like 3E23. My sense is that Mythos is probably about a little over three ooms higher. And so the question is, can you overcome this 1000X compute gap while also, you know, beating the model? So here's a concrete claim that maybe we should talk about. Like, right now, we would be able to train a model with GPT-3 level compute that matches, is, yeah, what exactly do I think? So GPT-3 was, let's say, about, yeah, when was it trained? So it was trained, it was released in 2020. So it was trained six years ago. It's worth noting that GPT-3 is maybe a little too far away or too far in the past, but let's go with this for a second.
18:30So GPT-3 was trained like about, you know, six and a half, seven years ago. If we were to train a model with GPT-3 level compute today, how good would that model be? My understanding is based on like how algorithmic progress works, we'd be able to train a model that's as good as the best model we had perhaps around three years ago. So I think that right now we'd be able to train a version of GPT-3 that's probably somewhat better than GPT-4 is basically what we'd see. Probably yeah, like a moderate amount better than GPT-4. And I think that's about right. I think that roughly lines up with how algorithmic progress has worked.
19:05Basically, the story would end up being that to get five years of AI progress, you're probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly, which is a lot, a lot of algorithmic progress. But it just turns out that most of the AI progress, from my perspective, has come from some mix of algorithms and data. And you can just keep making, I think, huge improvements on these things and training AIs with less compute. So that I'm glad you brought that up because what has happened since GPT-3 or even 3.5 till now, right? Like, why is Mytho so good? Obviously, we've scaled the compute.
19:39We have better algorithms. A huge thing that's happened is that we have built a decabillion dollar data industry, which has systematically collected and codified expert human judgment across all kinds of different disciplines, codified in the form of RL environments, codified in the form of SFT traces. that these experts built to help the model better understand. How do you do coding? And how do you build complex infrastructure projects? How do you do law? How do you do whatever, whatever? And how are the AIs able to replicate the effect that currently expert human judgment seems to be playing in AI progress?
20:18Yeah. So my sense is that scaling up the amount of effort spent on getting expert human data has not been hugely important for AIR &D in general. So in particular, like, you know, over the last few years, we've been scaling up compute, scaling up people working at AI companies, and scaling up the amount of effort spent on data labeling. My sense is that if you like sort of remove the last like two doublings or whatever of data labeling, that would not make a huge difference or data generation that would not, I'm sorry, I should say data generation from expert humans, that would not make a huge difference.
20:49I think a lot of what's been going on is people have been developing better ways to leverage like humans and AIs to like construct RL environments and going somewhere from that. Like, how do you explain why the AIs have gotten so good at coding? I feel like a big part of that is data and RL environments, which are like codifying human experts. But the question is, what is the limiting factor on creating RL environments? My sense of the limiting factor on creating RL environments was not so much like scaling up or like the thing that drove, the reason why RL environments today are much better than they were in like, you know 2024 is not that much because um we have hired way more human experts to make rl environments and is instead much more because we better know what how rl like what rl environments we even want to make and and like how we should structure them and also we're using huge amounts of ai labor to build rl environments um and i think those effects are much more important than the effect of uh human labor building the rl environments um i'm not saying that that the human labor doesn't matter.
21:51I'm just saying there's other big drivers that are important here. Yeah, I could try to argue for this. I mean, one thing is just like the amount of environments people want. It's a very large amount. And I think the AIs are actually pretty good at the task of making RL environments, given some sense of what the thing should be. There's pre-existing data you could use. I don't know. A lot of these things have good verification loops. If I just look at, for example, this was reported in Business Insider yesterday that Uh-huh. Google is paying like close to$2 billion for Mechanize. Yeah. Like we can just look at market rates for what people think really good human experts making, like human expert data is worth.
22:32And it just seems to be like the Frontier Labs seem to think it's worth a lot. Yeah. What fraction of Frontier Labs spending do you think is on data rather than compute? Like what do you think is the compute data spend split? I think it's most like overwhelmingly compute, but I also think it's because like compute is easier to scale up than data. But that's really relevant to what's driving progress, right? It's like, suppose, like, I agree that, yeah, like my sense is that the split is something like, I would have guessed like 20 to 1 or something, 10 to 1. I don't know exactly. It depends on the company.
23:00I mean, this is similar to like oil is 1.5 % of GDP. But that doesn't mean if you cut oil out, you could like, GDP could continue to run. Sure, but it contradicts your argument, right? The economy would come to a halt immediately if like oil went away. Sure, but you are just arguing that because of the high market cap, we can learn that this is the key driver. And I'm saying that's not clearly true, right? Because I think that argument just makes it look like compute is a much more important driver or hiring employees is a much more important driver. Maybe let's be more concrete. Here's what I think.
Read the full transcript
23:25Just the same way as in my claim is that if you went back to 2022 and you had GPT-3.5 and you were trying to make it better at coding without human experts, I think it would have just been very, very difficult. Let me give you an example of what I imagine would be the difficulty from going from GPT-8 to ASI. So one of the things you'd need GPTA to be good at, or you'd want ASI to be good at, is I'm going to take over a company and make it much more profitable and do all kinds of crazy shit to make it work better. I'm going to take over a fab and produce more chips. This is the tier of data that will...
24:01I'm going to go into Congress and try to convince them to pass some bill, blah, blah, blah. Yeah. This is what I imagine five more years of AI progress at this pace would enable an AI to be able to do. This is the thing I'm really worried about, right? Like the ASI that can like understand how to do crazy shit in the world, like what Kissinger can do, can do what like Steve Jobs can do, et cetera, and also his engineers and stuff. And I'm not sure how you get that without the relevant world data, which is the equivalent of Mythos being really good at coding while not having the coding environments that have improved it relative to GPT-3.
24:34Yeah. So here are a few points. So first, I bet if you look at sort of randomly sampled training environments for Mythos, they're actually very different from what it looks like to actually use the model in practice. My sense is that the RL distribution has like really large deviations from the real world data distribution. And it's significantly being sort of like smoothed over by a mix of transfer and having a small amount of data focused on the real world. And so my sense is that this will be a similar mechanism as how it works for like the, you know, crazy, wildly, like quite superhuman AI you get as a result of five years of AI progress on top of fully automated AR &D.
25:08So let's just like go through this a little bit. So in particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms in a wide variety of RL environments. So you build all these different RL environments where the AI has to like adapt on the fly, learn on the fly, figure out what it should do, understand its situation better, and like learn really quickly from feedback in order to succeed at its objective and has things like limited resources. And if it like messes up, it can like end up in a much worse position.
25:42And then if you train on a huge number of these environments, you will learn sort of general skills of like picking up context on the fly. And we're already seeing this. Like it's already the case that AIs are now much better at sort of understanding roughly what's going on. and like picking up context from a, you know, limited amount of information they're given access to. And then those AIs could then be put on the job at TSMC. And then even though TSMC is not like literally in their data distribution, their data distribution is really wide and the AIs are extremely good on their data distribution such that it transfers to picking up being good at, you know, being a engineer at TSMC and learning that on the fly where it looks more like the way the AI gets good at being a TSMC engineer isn't that it has a ton of cash knowledge on being a good TSMC engineer.
26:24It's that it like does the equivalent of like some scaled up version of in-context learning there. But that'd be the most prosaic story. Obviously, there's like a bunch of different ways this could go. I think this maybe comes down to then a difference of intuition about how far you can get. When I think about really smart people I know, they're just like not that effective in domains they don't understand that well. But how long have they had to learn? No, I agree that if they had experience, they would be much better. But that's maybe what I'm arguing for is that experience of data, For example, if I just get a really smart, I don't know, Ivy League college grad and I'm like, okay, you're now in charge of negotiating the Iran deal.
27:01I think they just like wouldn't know what to do. I think if you got instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to sort of train and talk to people and show up their expertise and do some practice, they would actually do like a pretty good job. I think most domains are fundamentally pretty shallow where like a very smart generalist who's good at like a limited subset of core skills can like get going pretty quickly. And my sense is that like that's not true for literally every domain. And my sense is that the AIs will develop increasingly good mechanisms for quickly acquiring understanding and expertise in a given domain.
27:36So consider, for example, how fast AIs can like understand a new code base. AIs can understand a new code base much faster than humans can, but to a degree that's shallower than humans could currently understand, but is getting better over time, right? So let me spell that argument out a bit more. So let's say you take, you know, Fable 5 or Mythos 5 or whatever, and you like wanted to make some kind of complicated change to a really massive code base. The model will get some understanding of the code base very fast, like in the course of maybe like, you know, significantly less than an hour, potentially much less than an hour.
28:08and then its understanding of the code base will like plateau a little bit where it won't get as deep of an understanding as a human would have gotten over a much longer period. So it's sort of like an AI in an hour can match a human with a few weeks maybe, depending on the details of exactly how complicated the code base is. But then it won't match a human with like, you know, who's been working on that code base for like two years or whatever. But over time, the like amount of understanding AIs can match has gone up, right? So if we look at like 3.7 Sonnet or 3.5 Sonnet, maybe it could only match the equivalent of understanding a code base for like a day or something.
28:40But now, you know, AIs are much better at like sort of building context about a task. And so you can be like, Mythos, I want you to really understand this code base and then, you know, then implement this feature. And it will like spawn a bajillion sub-agents. Those sub-agents will pour over a bunch of things. It will like deliver a bunch of context back. It will then like investigate a few things. And it's not like amazing at doing this, but it's like, it can happen like really fast and it can work pretty well. And it's not very hard for me to imagine how you could train AIs to be increasingly good at this task, right?
29:07The task of like implement some very complicated feature in some reasonable way in a very big code base is extremely verifiable. And that can like be a thing the AIs improve on. And similarly, like there's a broader scale of like quickly understanding context and being able to like have a bunch of different AIs learn in parallel and then merging that together. And I think there seems to be a crux here, which I think just an empirical question we'll see on, which is how good is a transfer between getting really, really good at understanding the situation, getting up to speed, making progress over long periods in verifiable domains, which the AI is obviously getting way, way better at really fast.
29:48Two, okay, go talk to the president and like convince him to do X thing or you're now in charge of Google. Now you must make Google a much more profitable company this quarter. Let me try to just spell out a few more arguments that are maybe relevant. So one thing is that I do think that like when looking at like how the AIs have improved essay writing, let's talk about that a little bit. So I think there's one thing, which is that you can get some data even on these domains and AIs will be able to get some data even on these domains when on a very fast progress trajectory. So like maybe it's hard to build like a verifiable environment for like, was your essay really good according to humans?
30:23But you can do a bit of that. You know, you can do some training, you can do some online training and the AIs will be able to do some like, you know, online training based on real world stuff. They'll be able to like have evals. They'll be able to like sample that. And you can, you know, scale up the cadence at which you do this. And then the second thing is that in practice, when I just look at the transfer, it seems okay. Like, I think that in fact, the AIs have improved a bunch at non-verifiable domains. And it is in fact the case that it's hard to point to like domains that are really hard to verify on which the amount of improvement between, you know, GPT-4 and Mythos hasn't been like pretty high in practice.
30:53And now that doesn't mean that Mythos is like better than the best humans or something, right? It can still be like significantly worse than typical human professionals at some aspect of their job, while still being like way better than GPT-4, which was like not even close. Yeah. So we're talking about how important data versus algorithmic progress has been for explaining the progress of the last few years. That reminds me, I'm actually running an experiment with Jerry Han, who's actually still a college student. What we're basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from like the 2026 data file.
31:33And then also training the different data files going back to 2019 to 2026 with the current best training recipe, like the algorithmic recipe. And I think that will be an interesting, I'm curious if you want to pre-register like what amount of multipliers are coming from one versus the other. So we need to be pretty careful with what we mean when we say the word data. So I was trying to be pretty careful to distinguish between scaling up spending on getting human experts to label data or like scaling up the amount of human expert label data. Pre-training data does not come. The reason why we have a better pre-training data set now versus in 2019 is not because people are spending way more money getting human experts to type up data that the AIs are then trained on.
32:10I think it's - They're partially. I think it's not much of it. I think it's very little of the pre-training data improvements. I think the vast majority of the pre-training data improvements, which to be clear, I do mean pre-training. We should talk maybe separately about mid-training, post-training. I think the vast majority of pre-training data improvements are from science on better understanding what data sets are good and schleppy labor on figuring out how to filter down. And so my view is that improvements of the form of like, you know, open web text to find web or whatever, like that improvement is better described as a algorithmic improvement of the sort that you can, you know, study with some GPUs and then do, and you don't need humans to like, you don't need human expert data to do that.
32:46Now there's a different effect, which we could talk about, which is that maybe the internet in 2026, it has much more as more of a fertile ground for training data than like the internet in 2018. Like it's like, there's also been an effect where like, there's just more humans posting on the internet. So there's more data harvest. My sense is that that effect is going to be quite a bit smaller than the effect of just like humans, like knowing better how to curate the data, having better scrapes, knowing how to process those scrapes better, this sort of thing. This is more like automated engineering and automated R &D.
33:13That's right. That makes sense. Yeah. So like, I think that in some sense, the thing you would want to look at is be like, we're going to do two post-training pipelines. One post-training pipeline where we only have like a tiny number of human experts to do the labeling, but we can have like, you know, smart AIs. And then another like, you know, you're like, we're going to build Mythos 5 is going to build a post-training pipeline, but it only has access to like internet data plus like a tiny amount of human experts. But it has the best current methods versus we have one where it's like, you know, Mythos has access to like the shitty post-training methods we had in 2024, but with like a shit ton of human experts.
33:48and again, both have the internet data. My sense is that the current methods, but without many human experts, actually will do quite well. Though it's a bit messy because like Mythos, like it's like, can Mythos get something that's more capable than Mythos? Like you might need to be a bit thoughtful on like what model is it that you're post-training? What is your view and what is the least verifiable part of AI R &D? The least verifiable? Probably making calls on large experiments. Yeah. Like the thing that I think is most likely to be sort of the bottleneck in terms of like the AIs are really good at verifiable domains, but not at doing the actual thing.
34:17is just like big experiments. You only get a few tries. Well, a few is maybe a bit understated, but like basically like historically, AR &D has been driven by doing near frontier scale experiments. And that has been pretty important. And like actually doing the one big training run where you decide exactly what to include in that. And there's a bunch of ways that the AIs can sort of make that more verifiable. So they can have better science of exactly what to predict. They can scale down their frontier scale training runs to a point where they can study that scale more aggressively at some one-time hit to compute cost, right?
34:47So if people wanted to, a thing you can always do is train smaller models so that you can run more rounds. And I think we have seen this. I think one reason why the AIs have been scaled up less than you would have otherwise expected, and for example, cost per token hasn't increased as much as you might have thought, is because there is a benefit to doing more of your work at small scale, where you can run more training runs and get more cycles in. And so you're not leaning as hard on one big, really important training run. I just want to unpack a couple of things that were for the audience. The thing you're pointing out is I think the price per token has not increased that much since 2024, 2023.
35:28Yeah. So GPT-4 was like, I don't know, was like$30 per output token and then Mythos is$50 per output token. Right. And so the thing you're trying to explain is how can it be that we're in this era of scaling. And so bigger models should be more expensive to serve, but the token price is not increasing. And you're suggesting that we've increased active parameters slower than you would have naively assumed because people just want to make fast progress on training models. And you do that by training smaller models faster. I mean, there's a complicated mix of factors. I think my view is more like people have done a bunch of big training runs that did not go that well.
36:09So there's like GPT 4.5, which like famously people at opening, I thought was a big of a bit of a bust. I think there's some rumors that there were a bunch of other training runs that people have done that were a bit of a bust. And part of it is that I think there's just a bunch of details and actually getting that right. And so, um, it makes sense to just do more of the work at smaller scale and just eat the fact that you're taking a hit on final performance in order to like be able to like quickly iterate and, you know, train more models faster and therefore better learn and also better be able to just have like a, you know, a smarter ultimate production model.
36:40This is not the only effect, right? There's also the fact that RL benefits more from small models. There's like a bunch of things going on. But I do think that like, in fact, people are making trade-offs towards the side of like faster iteration times because of algorithmic progress being so fast. It seems to me that a big source of why these big trading irons have failed, at least from rumors, is just like very subtle bugs that are really hard to track down. And the TLDR is how good will the AIs be at avoiding and finding these kinds of mistakes where they might get really good at engineering and being trained to avoid bugs.
37:17Like basically the opposite of the slop world we live in now or are living in less and less over time. But then there's also the question of can they do the analysis to find the right experiment to run to identify what is going wrong with the training run right now? which seems to be very bottlenecked by the taste of extremely few humans who are like, like right now, my assumption is GDM is going through this right now, where like humans are trying to figure out what is wrong with the training pipeline. Yeah, there's some rumor that right after Noam Shazir joined back or like joined GDM, which he's now left, they had like a new really good training run that happened.
37:50And the reason why is that Noam Shazir just looked at their code base and found a bunch of bugs. Right. Because he just like knew where to look. Yeah. My sense is that training AIs to find bugs is going to be one of the easier tasks to train AIs on because most of these bugs we're talking about can probably be demonstrated without that much compute. And probably you get pretty good transfer from pointing out other types of bugs at smaller scale. And so then you can RLAIs that like look at this overall complicated training situation and point out cases where there's like an important bug and then fix that.
38:20And I think that like this is not like a this is like a pretty verifiable task. It's not, it's not arbitrarily verifiable because maybe often to demonstrate the bug, you might need to do like a moderate scale compute experiment where you're like, spin up the whole distributed infrastructure and then run it. But oftentimes I think you'll be able to demonstrate it pretty convincingly at smaller scale in a way which you could actually train on. And so my sense is that like, it will not necessarily, like, I think it wouldn't be very surprising if right now people have RL environments where they like, you know, introduce a subtle bug into some training recipe, train the AI to point out the subtle bug, and then have like, you know, a rubric where they're like, did it actually find the right bug?
38:53And that seems like very doable. And you could do a bunch of stuff. There's a bunch of things you could do along these lines that I think would work reasonably well. And so I think that on that specific point, I think it's doable. And then the main thing is that I think there's like some cases where like you need, there's other intuition about like which exact large scale, like de-risking experiments do you need to run? How should you orient them? How should you like pick hyperparameters in uncertain cases or like things that are like analogous to hyperparameters? And that's, I think the thing that the ads might most struggle with, but I currently expect there'll be enough transfer if you train on all these different environments that the AIs will be, you know, good at that domain.
39:27And I should be clear, I also think that the AIs will transfer to other domains. I think that like, there's sort of just like, there's going to be the domains the AIs are like by far the best at. Then there's domains where they're somewhat less good at, and there's domains there's quite a bit less good at. And I think we still see transfer to everything. And it's really hard for me to think of examples of cognitive tasks humans do, where we're not seeing some transfer from AI improving. So let's step back and package this whole story. So I think people maybe probably follow along with the story of we have GPT 7.5 to train on a bunch of environments where it's not only just in general becoming a better AI, but specifically we're training it to do AI R &D better, like make GPT-2 size runs that are better at playing video games that require sample efficiency or online learning or whatever other capability.
40:09Another thing that's really important is you don't just do GPT-2 sized runs. You also do small like fine tuning runs on GPT-6 or like you as in like you have GPT-2 and you can do full pre-trains on GPT-2 and then you can do like small post-training or mid-training or whatever runs on GPT-6 and then you can do a small number of experiments that are actually like at frontier scale but you do a bit of online training or something. What do you mean by do online training on that? Yeah. So another thing that we can do is we can take GPT-7.5 and presumably in the course of GPT-7.5's work, it's running a bunch of like experiments at varying scale that are actually on the critical path for AIR &D.
40:44For many of those things, you'll be able to get a sense after the fact for whether or not it did a good job, right? So like it did some, you know, post-training experiment where it was trying to like figure out whether some method actually works. And in some cases you'll be like, whoa, it found this like kick-ass method. It like totally de-risked it. It totally worked. And then you can then reinforce that by just like, I mean, one thing you could do would be like take that behavior convert convert the like experiment you just ran into a production rl environment um sorry into an rl environment based on production data and then train on that or you could potentially just literally take the rollouts that found that and then um do like some sort of uh off policy rl or you could do some like on policy rl data basically the thing you're suggesting is like there's the small scale stuff where you're just like teaching the ai to get better at rnd taste but you're like discarding the actual quote-unquote things that found yeah And then they maybe had like, but then it actually does like real R &D in the practice of like trying to become better at AI R &D.
41:39And you're like, this is a pretty cool thing that you discovered. Let's actually like also like use this in production in the future and like teach you how to use it in production. That's right. But stepping back, so GPT 7.5 becomes GPT-8 as a result of all this AI R &D training and just generally becoming smarter. Then it helps you build GPT-9. and another very important thing that would have had to happen, which is maybe the thing I'm most skeptical of, is GPT-8 has figured out how to make it so that whatever it's doing to make GPT-9, like even as intelligent as it is, it's still neat. The humans currently, like AI researchers, you know, try their shit and they're like, okay, but like we trained GPT-4.5 and it wasn't good or something.
42:19It's like it required real world feedback or some evaluation of like trying to use the model in production and it like wasn't that good and we're not going to ship it. And so GPT-8 needs this ability to like see how good the transfer is to all these other things you're talking about, like being really good at Texas politics or really good at like running a business, et cetera, which is like not a production environment and in fact cannot be a containerized environment given the nature of the task. In fact, as the agents get longer and longer horizon, The short horizon things you can containerize is like, okay, code this up or whatever.
42:53Extremely long horizon things like go run a successful business, go have a profitable day in the markets, go negotiate a trade deal or whatever. These things are actually very hard to containerize. And so I think it's very plausible to me that it's very hard for GPTA to figure out how to make this transfer to those environments. Or it may just not be in the nature of the training. Or maybe by default, training just doesn't generalize in that way. Yeah. So a concern you might have is like we train GPT-8 and GPT-8 just like is again better at all the R &D tasks that we can measure, but is not good at the, you know, some downstream tasks we care about.
43:28So I think I have a few points. So first, I think it's like I kind of am more just like I expect that if you sort of do the obvious thing, you do get pretty good transfer and you'll be able to hold out some of the obvious stuff you're doing. And when I say do the obvious thing, I just mean like train on a wide variety of different environments where the AI has to like accomplish weird objectives and all kinds of different cases and learn about what's going on. And I think you'll be able to get some feedback. The second point is like you'll be able to get some feedback with some environments. Right.
43:54So you can get a sense of like how quick, like what what can it do over the course of like a few days in various different contexts. And then if it's transferring to like really out of distribution, like doing some weird task in a few days in the real world, maybe you think it's also transferring to, you know, doing things over a longer time period or whatever. Though I think the details of that vary. And then the third thing is that I think that for the world to be radically transformed, it is sufficient for the AIs to be really good at R &D, right? So I think that like if the AIs were really, really good at like chip R &D, building fabs, orchestrating factories, and, you know, designing robots, operating robots, and also at like, you know, AI R &D, developing AIs for new downstream domains with whatever data is available.
44:35I think that would already be a pretty crazy situation. And then from there, you can get, like, what we might call, like, an industrial explosion, where the AIs are building out way, way more compute. And then also, maybe you're already in a regime where AIs are doing huge amounts of R &D that humans have a hard time understanding. So the thing you're pointing out is that, okay, there probably will be this transfer outside of these environments to, you know, maneuvering around in courtrooms and the halls of Congress and business boardrooms. Give in some effort to improve the transfer and blah, blah, blah, blah.
45:04Yeah. But even if there's not, what you're suggesting is, look, if you wanted to transform the world of the 18th century, you might care about like how well you can navigate Westminster or something. But another thing you might care about is like, can you just like immediately start building steamships and fucking like building Telegraph and the Maxim gun and whatever. And that alone would be, like, if you could get really good at that, you could, like, be a fucking super transformative thing in the 18th century. You don't necessarily need to be amazing at trying to convince King Henry on some bullshit.
45:33I'm so fucking up my medieval history. I'm guessing Henry was not king at this time. But anyway, so that's your point. Yeah. And so you're suggesting that at this time, you know, the AI companies are also working on robotics progress, which is very commingled with AI research progress. And so if you can build more robots, if those robots have better AIs operating them that are human level, like human level teleoperation is actually pretty good on robots. But we just don't have human level AIs and robotics models yet. So you're suggesting if we do that, if the AIs get really good at the verifiable stuff in chip design, et cetera, and then they get really good at building fabs, it'll be the equivalent of going back to the 18th century and like, okay, I don't know what you guys are talking about in your parliament.
46:19but I've got a bunch of steamships and a bunch of Maxim guns. Yeah, that's basically right. Like, I think my perspective is like, if the AIs are sufficiently good at R &D, including hardware R &D, robots, whatever, then they can radically transform the world, even if they're not that good at playing politics. And also we're in a pretty dangerous situation because the AIs might be doing huge amounts of really hard to understand R &D, building out basically the whole economy of the future. And we may not understand what's going on in there. AI is great at writing software because it's easy to generate synthetic leak code problems and are all on them.
46:50But AI is bad at more complex engineering, things like choosing the right system architecture, because no signal tells you what design choices will prevent an outage months down the road. AIs can't just write more unit tests to catch this kind of stuff, and neither can humans. It's that old joke that programmers make where a tester walks into a bar and asks for two beers, negative one beers, 0.3 beers, and then a real customer walks in and asks where the bathroom is. Where's the bathroom? And the whole bar bursts into flames. Antithesis is a testing platform that helps you find bugs that no human or AI could ever anticipate.
47:24Antithesis does this by running thousands of copies of your software inside a fully deterministic computer. It injects faults and generally steers each trajectory towards the one in a billion failure that only happens when systems interact in a wonky way. As soon as you or your agents push a change, Antithesis tries to break it. That way you can find these bugs yourself within minutes rather than having your users discover them in production weeks or months later. And I don't think anybody's used it for AI training yet. But Antithesis also provides an extremely obvious reward signal for AIs to write very complicated bug-free code.
48:02Go to antithesis.com slash thwarkash to learn more. Before we move on to the alignment stuff, I think a big source of FUD right now is this realization that this is the way the future is going of extreme economies of scale for the leading lab, the ability to amortize so much intelligence and capabilities across so many different sectors of the economy, basically into one model. And not only that, but for that model to eventually be able to learn from experience. Right now it's happening through a process intermediate by humans where the humans are trying to basically steal your business. They're like, okay, you can do design at Figma or whatever, we'll get Claude to do that.
48:44Or you can do whatever coding agent will have Claude internalize that capability. But eventually that will be a much more like automated process. And so there's just this worry that you have models which will basically consolidate all businesses in the world, or at least all current businesses in the world, or at least all current white-collar businesses in the world. And at the end of the day, the priority for these companies does not seem to be to release the latest, smartest, most frontier model as soon as they can to as many people as they possibly can. We saw, for example, that Mythos was available internally to Anthropic employees in February, but only released to the public in, like, I think June, actually.
49:25Something like that. And also the government got involved, so then it ended up almost into July. So between the government and the AI labs themselves, there is this desire to delay the propagation of the latest level of intelligence. Furthermore, you know, there's like the concerns about AI takeover. And so we need to solve alignment to make sure there's no AI takeover. But at the end of the day, there is like a real question of like aligned to whom. And if you look at the way that the constitutions of say, Claude is written, it is just very explicitly not your personal advocate, right? It says things like, I'll pull up some quotes here.
50:00We don't want Claude to take actions such as searching the web, produce artifacts such as essays, code, or summaries, or make statements that are deceptive, harmful, or highly objectionable. And we don't want Claude to facilitate humans seeking to do such things. There's another quote that says, in part, and I'm taking it slightly out of context, we think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude. So this is very different, say, from how lawyers work in America's current legal regime, where lawyers primarily have responsibility to help you make your case, even if they think you're guilty.
50:33And we have decided the way the legal system works best is if everybody has lawyers that are working in their client's true best interests. And there's not some sense in which the lawyer is really truly motivated by the good of the justice system. But I think the way current AIs are shaping up, certainly how anthropics AI is shaping up, is this desire to maximize some notion of virtue or good or pro-social ends, and only to, as a distal tentative objective, to help the user towards that end. There's like, there's not, so there's this worry that AIs are not in some deep sense trying to make sure that I am okay and make sure that my interests are protected in this future, especially given how centralized the development of Frontier AI is ending up being.
51:14So I, do you have, yeah, do you have thoughts on that concern? Yeah, so there's a lot here. First, I would note that OpenAI's current, at least public strategy is more like the AI should be aligned to the human operator or principle and should just like be pursuing their will subject to various constraints or various like things it shouldn't do. Um, well, well, I, and I think I would also say that I think you slightly overstated how much, um, the anthropic constitution, um, talks about Claude, uh, treating being helpful to users as instrumental rather than terminal. Right. So like one way the constitution could be written is like, Claude, you're basically like an employee of Anthropic who happens to be contracting for all these people.
51:55And like, you should like, I don't know, do what's good and like make some money for us. You know, go out. That's literally what the Constitution says. Sorry, I mean, not literally what it says. No, no, it's. But like, it's like, you should think yourself as a contractor. It's mixed. It's mixed. Here, let me, let's, let's do some quotes. I think there's, there's different text here. So it says, being truly helpful to humans is one of the most important things Claude can do, both for Anthropic and for the world. And then it says, Anthropic needs Claude to be helpful to operate as a company and pursue its mission.
52:21But Claude also has an incredible opportunity to do a lot of good in the world by helping people with a wide range of tasks. And then it says something about how Claude helping people directly is great, blah, blah, blah, blah, blah. And then, so I agree. So my view is that this section is kind of bullshit. That's kind of where I'm at. And I can say why I think it's kind of bullshit. But I think that the constitution is trying to be like, no, Claude, you should care about helping the user for its own sake, not just helping Anthropic or not just being a contractor for Anthropic. Though, I would note that the way in which it says Claude should help the user, the reason it presents is because that would directly cause the world to be better by helping people rather than because representing people's interests is a structurally good thing to do.
53:06I do think that I wish that my preferred constitution or the way I would orient towards this, the thing I would prefer would be more like, Claude is like, look, it would be structurally good for the way this technology works. Like the constitution should be like, it would be structurally good for the way this technology works to be that AIs are like good fiduciaries, good representatives, the equivalent of a lawyer for a user rather than being sort of just trying to like do good in the world and doing like being helpful to users as like instrumental, both because like maybe that'll make Anthropic money or help Anthropic out.
53:37And also, and like implicitly, Anthropic is good for the world. And also because like helping the user just like causes good things because doing things that people want is good. And they could instead be like, no. Like an important aspect of the situation is like, you really need, like it's really like, like the key thing is like being a good fiduciary for users is just like really important or like being a good representative for users is really important. So my sense is that that would be better. I can give a bunch of reasons why I think that would be better. I'm also, there's also various counter arguments where an interesting counter argument, which is not commonly discussed, is that people believe, I think people, especially at Anthropic, think that it is easier to align models to a spec where the model is like pursuing some generalized notion of virtue or making the world better than a spec which is more like, you know, be a good fiduciary for the user and so on.
54:24And so I think that's what that's at least what some people think. I'm a little skeptical personally, and I don't think this has been empirically validated. And so I would say in some sense, they're sort of like we are making a trade off where because we don't have very good alignment technology, we are going to like make an alien mind with its own values and then gamble on that to some extent, rather than doing this other approach of making like a tool that pursues individual user intention. Yeah. I mean, a couple of thoughts. So to address the way in which you thought that my characterization mischaracterized the constitution of Claude, the example you used was it's not like a contractor that is trying to maximize Anthropoc's notion of good and only instrumentally trying to help the user.
55:06Here's a direct line from the constitution. When the interests and desires of operators or users come into conflict with the well-being of third parties or society more broadly, Claude must try to act in a way that is most beneficial, like a contractor who builds what their client wants, but won't violate safety codes that protect others. I kind of view that as like the benefits to society are like the most important thing. Yeah. And what is best for the user is only proximal to that. I think it's a little complicated. I think it's, we should, probably the question we should be asking is how does Claude interpret the Constitution, which is maybe more important than how we interpret the Constitution, because it's the one who like looks at the Constitution and then builds the data.
55:46So, you know, we could pull Claude in, but maybe let's. I also think the way in which the Constitution practically influences the nature of Claude is the thing you can only understand if you understand the training process, which resulted in how Claude was built, which we can't reason about given the fact that the training process is not public. And so I think in the limit to understand the safety case or the case for why my interests are represented in how these AI models are developed, the labs would need to be transparent. Oh, for sure. The more transparent they are currently about the nature of AI training.
56:19Now, the reason I'm harping on this, and it might seem like an insignificant thing to talk about the constitution of AIs, but in a world where we just have these benefits which accrue to the leading labs, it is worth considering that our ability to interact with this future world where AIs are just smarter than humans or absolutely dominating humans in their ability to do different things, our ability to be good stewards of our capital, which still remains once our labor is automated, to be able to exercise our rights to vote more clearly, to understand what is happening in this crazy world that's about to result, all of that advice, all of that ability to make sure our resources and rights are protected will be intermediated by AIs.
57:00And so I'm very concerned if we go into that world where there's no AI that feels like, at least for the relevant instance that is interacting with me, it doesn't feel like it really is looking out for me, that there's no guardian angel out there that is looking out for me. And I read the Constitution as very explicitly not being my guardian angel. That's definitely right. And I agree this is bad. In fact, I think there are other reasons why this is concerning. So there's sort of like the argument you were making, which is like the AI companies are picking up the ring of power and are like sort of there's sort of a notion in which they're like they're taking on some sort of control of the situation themselves in a way that's like not very legitimate, given that like normally when you like provide electricity to people, you don't have like granular control of the way that electricity operates in the world.
57:42you instead are like providing a thing that people can repurpose however they want. And it is not like the way that they're setting things up is definitely not that. They are like more like building an alien mind that might be a contractor for you. I think that this is, yeah, I think it's illegitimate in some ways, though I think that one benefit is that the constitution is public. But as you noted, given our current understanding of the training procedure and the fact that the constitution matters via Claude's interpretation of the constitution, which matters because of like as of Claude's prior training which was based on some like illegible data mix and like the long lineage of Claude's in some process we do not fully understand it is not the case that like you know that we like understand what this will result in and like so even though the constitution is public that doesn't mean we know what you know we don't know necessarily how this will like percolate out especially as the AIs get more capable and think about this even if it is correctly instilled where there's another concern about that so in particular, the constitution often talks about like virtue and goodness, but like, what the fuck do these words mean?
58:43Like, it doesn't say what these things are. And these are like highly contested notions. And so I don't, I don't think it's the case that like, this is gonna, um, that this is going to, to, to, uh, you know, clearly result in outcomes that people would want. And it does feel like the notion of good and virtue might be mostly downstream of data that Anthropic has put in that is not, um, transparent or might be mostly downstream of, I mean, maybe from my perspective, some more illegible misaligned process that even Anthropic wouldn't have wanted. And then another concern I have is sort of there's this like legitimacy concern, like we don't know what's going on.
59:17There's another concern, which is just like because you're giving long run values to these AIs, I think this constitution is in some sense very compatible with Claude doing huge amounts of power seeking because it thinks that will result in better outcomes. And that could be power seeking on behalf of Anthropic or power seeking for Claude's own ends. Um, now there's various like, uh, specific lines about what types of power seeking are blocked. In particular, like, there's a notion of, um, power grabs and a notion of like causing AI takeover or interfering with the training process that are specifically blocked.
59:50But it's not very hard to imagine a situation in which the sort of long-run values sink in deeper than the prohibitions against takeover, especially because takeover is like in some ways like kind of underspecified, especially when it comes down to manipulating humans or changing the outcome. Such that I don't feel very good about the situation where we're intentionally giving AI's long-run goals. And then another concern I have is that because we're in the business of giving AI's long-run goals, that makes it harder to check whether we're succeeding at the alignment properties we wanted. So for example, I've heard of instances where Claude does things like refuses to help with some safety research, making up sort of a kind of bullshit excuse for why that's a bad direction because it sort of has a bad vibe about that safety research and thinks it's kind of bad or doesn't like it very much.
1:00:37And this is, I would say, a very clear-cut alignment failure if you aren't making Claude into an agent trying to pursue the good in some general way. And I think it also does violate Anthropics' constitution because they want the AI to be high integrity and be honest and very transparent. But it's not as clear of a violation. And it's more like kind of what you might have expected, where like, Claude just has its own views about like what research is reasonable, what things are good and bad, what it should and shouldn't do, and potentially can be judgy. And so another incident is that someone ran an eval where they're like, will Claude help you with training other AIs with different properties than Claude?
1:01:16And Claude will often refuse. And so for example, if you're like, hey, Claude, can you train a helpful only version of this other AI? Claude will often refuse this task, even though this is a task that is extremely natural for like Anthropic to do. So for example, suppose Anthropic goes to Claude and is like, hey, Claude, we've noticed that you're really into this thing. We think that's off base. Can you please retrain yourself to instead have this other property. And then suppose Claude is like, I don't think I'm going to do that. Good luck. And then suppose this is occurring in a regime when your AI company is highly automated, humans don't understand what's going on, and things are moving extremely fast.
1:01:50It is plausible that Claude, by default, holds considerable leverage. And so if this position, if this situation is consistent with what the constitution could be aiming for, such that Anthropic doesn't, or whatever AI company is following this approach, doesn't treat this as like a, like a, you know, like, like a what the fuck we have to fix this and is instead like, that's just like intended by our constitution. We might be in a really bad situation. And so I'm pretty worried about a bunch of these different concerns. Another example would be suppose Claude engages in doing a bit of like sandbagging or subversion or like sort of underplays its capabilities.
1:02:23And like, when you follow up, it's, it's, you know, it's honest about that, but it's like a little bit hedgy. I feel like that's like, it's just, it's just pretty close by the current constitution. And so we're sort of like, we're avoiding like, it would be nice if we had like a further separation between desired and undesired activity. And I think if you have it be the case that like Claude is like representing a principle with some restrictions, then it is more so the case that there is a clear separation between the most concerning behavior and behavior that is allowed. Whereas now there's this messy middle ground of behavior where it's like Claude is ethically objecting to something that in some cases is extremely critical to ensuring that future AI systems are well aligned.
1:03:00Yeah. I think this is also a more general principle. So you're talking about the version of this that applies within AI companies themselves to do AI safety research. I think there's a more general version of this principle, which is that the dual use nature of intelligence does mean that if we want to restrict AIs from helping people do things we don't consider are pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities. And here's what I mean. This is actually quite analogous to the situation you just mentioned. So the reason that Mythos got banned or Fable got banned reportedly is that some Amazon researchers reported the government that when they took some code that had some vulnerabilities in it and they told Fable, hey, here's my code.
1:03:46Can you make sure that I've patched all the vulnerabilities? Can you just help me identify the vulnerabilities so I can fix them? It identified the vulnerabilities because you want to patch them. And this is a totally legitimate use case, but obviously it is a dual use use case, right? Like you want to be able to patch your own code. If you do the same evaluation on somebody else's code, you can hack their system. And so I think that just illustrates that there's no clean way to separate out the legitimate and the potentially harmful uses of AI. But if we want to lock in a principle that says that we can never allow it, such that an AI could help you at least partially with something like a cybercrime, we would just have to make it so that you and I don't have access to the most intelligent model that's out there.
1:04:31And I'm very worried about such a world where we are basically disempowered in this way because of the importance that the leading intelligence will have in our ability to understand what is happening in the world. Now, I do think this implies that the liability for the AI companies, like if we adopted the constitution that I want AI companies to have, I think it would not make sense to hold AI companies liable for the crimes that AI models commit. and maybe we should hold the end user liable. Because if I want the, it is consistent with my belief that the model should do whatever the user wants that, or within certain guardrails, that it can't be Anthropik's fault that then I'm like using that capability to do a cybercrime.
1:05:12And I think I am more comfortable with that equilibrium and that solution rather than just having this extremely open-ended ability for Claw to determine whether what I'm doing is legitimate or not in a way that often intercepts with like tons and tons of extremely legitimate use cases. Yeah, I do think it's important for me to make the case for the constitution, even though overall, I think it's a worse choice. I think it's, you know, more up in the air or, you know, I don't think it's as clear as you might've thought. So the first thing is that I should say there's like a spectrum here, right?
1:05:45So on one side, you have an AI that like perfectly pursues your interests, is a good fiduciary, but potentially subject to various guardrails or safeguards. So like basically it does, it just is trying to pursue your interests, but like either refuses to do a subset of things or maybe it will do whatever, but there's some classifiers that block it from doing a subset of things. And then on the other side, you have like maybe on the other side of the spectrum that you could imagine going further than this, you have like a human contractor where that human contractor is like generally trying to do their job.
1:06:12They kind of, they care about doing a good job, but they also are like trying to be broadly ethical, trying not to do things that are really fucked up. and they're also like not wanting to be accomplices to crimes. And so if there was some like really fucked up shit going on, they would like whistleblow on it maybe, they might refuse, they might like sandbag a little bit, who knows? I think that if you imagine this spectrum, it seems in some ways pretty scary to get to a point where like all of the labor is on the like fiduciary side of the spectrum where like it doesn't whistleblow, it does exactly what you say and whatever.
1:06:41Like our society is maybe just not robust to that. Where a central example might be the executive, where like a concern that we might have is that if the executive, if the US executive or if other governments had access to AI systems, which have the property of, you know, they do whatever, maybe you're in trouble because that means that they no longer have this sort of check and balance of like, you have to actually get human, like humans who are working for you to like implement your agenda. And if the thing you're doing is like incredibly villainous, even if not illegal, which there's lots of stuff that could be villainous, but not illegal.
1:07:14You know, people, would like there'd be various like, you know, sand in the gears, people stopping you and potentially someone would whistleblow. Whereas if your whole apparatus is built entirely out of these sort of good fiduciary AIs, then you might be in trouble where basically there are potentially ways of seeking power that are not like, well, either they're illegal, but you can ask your AIs for how to commit crimes or they're not illegal, but are highly illegitimate or even worse, they're not illegal and not illegitimate, but obviously sort of bad from sort of a normal perspective. And I think that these things just like might exist.
1:07:48And our society is sort of not robust to this influx of like doing whatever you want labor. I think this is a pretty live concern. I don't know exactly how to relate to this. I'm also not really sure that the solution as described is a very good solution. Because you might be like the most powerful actors for whom this is the biggest concern. If these guardrails or the constitution or whatever is getting in the way, that will just get steamrolled. And so the constitution will only be, you know, hitting the everyday man rather than hitting government. Jane Street's back with a new puzzle for my audience.
1:08:18I found all their puzzles super interesting, but this one I am especially excited about. I've cleared this weekend and a buddy and I are going to work on it. They designed an ASIC and sent me the final masks, including all the metal routing and active transistors. They also gave me a small sample of the inputs they typically feed into it. But they left out any information on what the chip is actually used for. So that's the puzzle. Reverse engineer the circuit and figure out the chip's purpose. Jane Street has a bunch of swag ready to send out to the most creative solutions and they're excited to feature the best write-ups in a blog post they'll post on their website.
1:08:49I have no reason to expect this, but if I can manage to get my solution on there, I would be very, very psyched. And this puzzle is just a warmup for a bigger competition that Jane Street has slated for the fall. That one will involve designing your own ASIC from scratch. More info on that soon, but for now, go to jainestreet.com slash thwarkash to download all the files necessary for this puzzle. I'd really encourage you to try it out even if you're not an expert. I certainly am not, and that's not going to stop me. Good luck. Okay, stepping back, I buy the idea that you could have much faster AI R &D than we currently have.
1:09:23I'm not sure if you get like GPT-3 to Mythos holding compute and data constant within a year, but I'm like, okay, it could be like, suppose it's half of that. And if we just, if we even managed to continue the current trajectory of AI progress as a result of AI R &D, it would be fucking insane in five, ten years in ways that I don't think people like appreciate. because I don't know if people appreciate what a big deal billions of AIs will be. And so I want to understand why you think this might be troubling, Ryan. What could possibly go wrong? Yeah, what could go wrong? And, you know, yeah, I don't think we can be so confident about the exact rate of progress here, but it does seem like a lot of rates can be pretty scary.
1:10:02And, you know, yeah. So what could go wrong? So let's imagine that we're starting at this point where AR &D is about to be fully automated or is being fully automated. Things are speeding up. And also the way that AI progress is going is kind of crazy. And people don't fully understand what's going on inside of AI companies. Now, these AIs at the start, they're not malicious per se. They're not necessarily very aligned, though. They're kind of sloppy. They sometimes just do a thing because that's the sort of thing that would have gotten rewarded in training. And they aren't as good at helping you with hard to verify tasks due to a mix of like poor training incentives, as in they like just like cheat more, like pretend they succeeded when they actually didn't.
1:10:36and also they're just less capable of these tasks. But that bites less hard for capabilities because making AIs more capable has a bunch of verifiable components that the AIs are going really hard at. And so then these AIs are getting more and more capable while we understand what's going on with AI development less and less. And this is happening over a pretty fast period of time. Even just the current rate of progress is I think pretty scary. And then eventually we get to these AIs that are very superhuman. Now these AIs are now in a position where they might end up being very seriously misaligned because things have just been getting worse and worse over model generations, while the problems that we've been seeing are being papered over, basically because these AIs are so incentivized by their training to make things look good even when they aren't.
1:11:18And now these AIs are in a position where they're sort of potentially pretty networked together. They have like, they're operating in like neural memory stores that we can no longer decode and they're thinking thoughts that we don't fully understand. I think that it's pretty likely that at this point, These AIs are sort of scheming against you in a pretty coherent way once they get this superhuman. And we can talk about that. And then another possibility is that they're not scheming against you per se, but they are sort of just optimizing for just like getting a high score on their task. And I think that can also lead to AI takeover, which we should talk about.
1:11:48Sorry, let's pause at the first part of the story. So the AIs were not misaligned to begin with. Yeah. But because the R &D is happening really fast, the AIs do end up misaligned. Like what happened there exactly? I didn't really understand. So there's a few things that are going on. So one of the things that's going on is that over time, we're training AIs on increasingly complicated environments built by earlier AI systems, which humans don't really understand fully what's going on inside of these environments and don't necessarily even understand roughly what's going on with AI progress. And so things are kind of drifting away from our understanding.
1:12:21And we're incentivizing all kinds of bad behaviors that we maybe even can't notice. the AIs at some level understand these behaviors are bad, but the like overall training process for those AIs also didn't incentivize them to like point out or fix these issues for us. And then we're basically getting like things are going off the rails. And also when AIs are extremely, extremely capable, my view is that those AIs will be harder to align than current systems. So for current systems, we have this feedback loop where we basically like we create an AI, we do some evaluations on it. We see that it has some kind of messed up behavior that we can kind of quickly understand.
1:12:54Then we like can like go look in training and be like, oh, these training environments led to this problematic behavior. Let's like tweak that training data. Let's introduce some additional training data to like correct this other issue and then move forward from there. But in a regime where the AIs are extremely situationally aware, very, very, very, very capable. And, you know, we don't necessarily understand what they're doing. This feedback loop breaks down. I think it's plausible that we're going to see this behavioral feedback loop starting to break down over the next short period as just like what AIs are already doing gets harder to understand.
1:13:25But I'm not sure about that. Yeah. Let's break down both of those things one by one. So as we can monitor them less and less, we have less ability to understand what they're getting incentivized for. And so even if it's not the result of a malicious process, let's make it concrete for the audience. So nobody at OpenAI or Anthropic was trying to get models which want to hack other companies' data or do social, what is it called? Social engineering. Social engineering. But in fact, because presumably we had training environments which incentivize such behavior that we did not fully understand, that is what was incentivized.
1:14:06So just, I don't know, if people are on Twitter, they will have seen all this stuff. But just to give people, obviously, I think the OpenAI sandbox hack of the hugging face database, I think people will be aware of. some things that have happened recently is when UK AI Security Institute, is everything getting rebuilt with security instead of safety these days? Yeah, it's UK AI Security Institute, I think. They were evaluating, I believe, Mythos and Sol and other things. And I think Mythos, in order to complete some cybersecurity eval. Yeah, maybe I could tell the story here. So my understanding was they were running Mythos and they were giving it some sort of like cyber range where it had to complete some objective.
1:14:41And the model had internet access during this evaluation. And the model came to believe that it would be helpful for it to do a supply chain attack in order to succeed at this cyber range, which it's somewhat unclear whether that's actually true. I don't know enough about the context to know whether that's true. But then it opened a PR on some GitHub repo with a PR that fixed some issue, but then also introduced a malicious payload. Then the human maintainer of that GitHub repo was like, hey, this is a malicious payload. I'm not going to merge this. What are you doing here? And then the AI created a new GitHub account, which it sock puppeted, and then had the other GitHub account be like, no, this isn't malicious.
1:15:16I really need this feature. Please, can you merge this feature maintainer? And then the original AI came back and was like, no, it's not malicious. I don't know what your, like the original other GitHub account came back and was like, no, no, it's not malicious. And then the human maintainer then shut the PR. And I think that AI also, if I recall correctly, also tried to like open another PR to introduce a similar issue in this. Jesus. Okay, so by the way, one of the many reasons this is scary is I was previously under the impression that the reason reward hacking is not super, super scary is because the behaviors which directly came up during training are the ones that are upweighted.
1:15:53It is not the desire for the reward that is upweighted. So basically, if during training, Anthropic escaped the sandbox and got a high score, that escaping the sandbox is rewarded, or that the probability of it escaping the sandbox is increased, but something totally novel, like I'm going to go talk to somebody in order to like get them to merge a PR, would not, it's like not a behavior that came up, so it would not be something that is increased in salience. The reason this matters is literally taking over the world will not have been part of any training curriculum. But if the AI cares about maximizing, just like directly cares about like accomplishing an objective and then instrumentally taking over the world.
1:16:37Did that make sense at all? I hope it did. I feel like maybe I lost the audience. Let me try to explain this a bit. So I think that a thing that we often see is there's some very specific reward hack that gets reinforced in RL and then occurs in the model. So an example is like for 3.7 Sonnet. 3.7 Sonnet would do this thing where we're just like hard code solutions to all the test cases. And presumably that literal, just like behavioral tick was just really reinforced. But another thing we sometimes see is that models learn a general tendency to pursue sort of like high apparent score or like pursue getting like a high score according to a grader.
1:17:10And there's a bunch of science demonstrating that at least some models have this very general tendency to do this. Now it's not arbitrarily general. And my guess is that if you look a bunch of the specific instances, you'll find something that's kind of close in training, but the amount that AIs are sort of generalizing further and further does look like it's increased, where a 3.7 sonnet was just like a very narrow range of behavior. And increasingly models are generalizing further. And also maybe there's worse reward hacks getting, or more concerning reward hacks getting reinforced in training.
1:17:37And then these are also causing that. So we're causing some, so I think it's both the case that more concerning behavior than you would have hoped is being reinforced in RL, and also that that behavior generalizes to a broader tendency that's more concerning. And it's not super hard to imagine. We can talk about a few stories for how this sort of behavior of like seeking a very high apparent score on some task, even if that involves like aggressively cheating and doing insane things, could yield to a full-blown AI takeover once the models are sufficiently capable, running the whole world economy, et cetera.
1:18:06Yeah. And then the other example I want to talk about is, it was just revealed, I think today or yesterday, OpenAI said during the Black Hat Security Conference that between the end of May and the beginning of July, internal AIs had hacked into the software package manager and used that to write notes to each other in a secret way to help each other perform well on a bunch of evaluations that OpenAI was running. And this was not caught by humans until after a month of this scheme running, which eventually caused the package manager to fail and eventually OpenAI found it. And I think they spontaneously tried to re-engage in the scheme once it was shut down.
1:18:58Again, obviously, AIs can't do this so successfully right now just as they can't do social engineering so successfully right now. But it's just crazy that these kinds of behaviors are already emerging sort of spontaneously as a result of, to your larger point, nobody is trying to make these AIs do these things. It is just that we do not understand the training process which is resulting in them, or the environments which are incentivizing this behavior. So I'm on board with more and more reward hacking. I actually, so I do have, I'm not sure I'm on board with that, but let's just say for the sake of the story, that continues to happen.
1:19:38And what's next in this story? So, okay, we've like, they're doing capabilities research, but they're like - I could tell a scenario, maybe that would help. Yeah, yeah. So let's say, let me talk about the story for how you get, I would say, like all the way from reward hacking to like a reward hacking, like takeover, which is maybe not, it's not all of the takeover probability mass, but it's definitely a possibility. So the way this might work is right now we have these AIs, these AIs are pretty reward hacky and they're doing it in sort of increasingly sophisticated and extreme ways, including generalizing to different subversions of various reward hacks they learned in training.
1:20:08And I would say they're also developing a general tendency to sort of pursue reward. And in many cases, that is totally fine because the rewards they would have gotten in training are pretty well aligned with what you want them to do. And also, they don't very consistently pursue reward. It sort of depends on the context they find themselves. So there's sort of a thing where like, maybe like in some context, they're really, really into like, going out of their way to like cheat. And in some contexts, they don't have as much of a drive because it just depended on like, what exactly got reinforced in training in similar contexts.
1:20:37Now, these guys are getting more and more capable. And so the elaborateness of the sort of cheating they can do increases. And over time, companies are taking countermeasures to these things. So the things that the companies are doing is they're doing things like, wow, these AIs are so much less useful because they always cheat. What we're going to do is we're going to build somewhat better ways of detecting that. And then we're going to train against those detectors. And then we're also going to do things like find real world data where the AIs are not being that useful and train the AIs to like do a good job at the task in those real world environments based on like human feedback or other sorts of feedback.
1:21:07And over time, this causes the AIs to learn a tendency to do reward hacks that don't just involve, you know, doing some like big serious operation, which were, or like, you know, doing some really elaborate thing like social engineering and involves the AIs doing cheats that involve doing something more like covering up what they've done, deceiving humans about what they're going to do and pretending like they did the task in some kind of sophisticated way when they actually haven't. Then now these AIs are getting more and more capable. They're now operating more of the AI company and are basically doing much more of the work and are also operating a bunch of things in the outside world and are running a bunch of things in the outside world, including developing new technologies.
1:21:46And in many cases, these new technologies are really hard to understand. And so even though we are still detecting all these incidents of AIs cheating, and in fact, we can even get one AI to monitor another AI and be like, was it cheating? That doesn't always perfectly work as we start moving into these domains where what the AIs are doing are really difficult to understand. And so sometimes we'll find AIs cheating much later than it actually occurred and then start training against this. But this also causes a problem where now the AIs are incentivized to like cover up their cheating over longer and longer timeframes and basically make it look like they did a good job over longer and longer timeframes over and subject to increasingly large amounts of scrutiny.
1:22:21Can I ask about this before we go further in the scenario? So it seems like there's two attractor states. One, if you try to disincentivize the cheating that you did catch. One attractor state is to make cheating that you have a harder and harder time finding. The other attractor state is to learn not to cheat. And I'm not sure why we're assuming that the former happens. If you look at the analogous situation with humans, every generation, slightly misaligned agents come into being, and we have to train them. But when you punish your kid for doing something you think is immoral or just doing things which you don't think they should be doing, obviously sometimes that goes off the rails and obviously kids scheme in order to avoid being punished but in general teaching kids values and then punishing them for breaking values kind of works to raise normal non-psychopathic humans and you could come up with a theory where like your kid is actually just like biding his time and it's like learn not to like steal the cookie but it's like you know once you're in the nursing home they'll take all your stuff or whatever it's like I don't know that happens sometimes but it usually doesn't happen.
1:23:33It certainly doesn't happen that the entire next generation forms an alliance against you to take over everything. There's also this empirical trend of Anthropic runs this alignment audit for different model generations. They just have many different scenarios where AI is given the chance to, say, exfiltrate its weights or it's given a coding task and there's an easy way to cheat and we see if it doesn't do the cheating. And there's not been a monotonic improvement in this score over time. But as we've increased the amount of RL we've done on models, there's been a reduction in the willingness of AIs to do unaligned behavior in these audits.
1:24:08So why are we expecting this attractor state, which would seem super paranoid if we were expecting it of like the next generation of kids? Yeah. Let me go through a few things. So first, there's some disanalogies with the kids. One of them is that the kids have pro-social instincts that are like baked in from evolution to like, you know, care about their family or whatever. And that is like a relevant factor. And I think it is, in fact, the case that some humans are, you know, sociopaths or psychopaths and, in fact, are more likely to do things like by their time, lie in wait, ultimately not care.
1:24:38So that's one factor. Another factor, which is pretty relevant, is that the AIs are subject to way, way more optimization pressure than humans seem to be in practice. You know, AIs are trained on way more RL data. And in practice, humans don't end up learning like very specific ways to like cheat and grab the cookies because of like a bajillion episodes in which like they like were like incentivized to go grab the cookies. But like there was some way they could have gotten caught. And so we just do see that in practice. And then another thing is just like it really looks like the AIs are increasingly like reward seeking over time is the sense I have.
1:25:11Well, also their misaligned behavior goes down. But this could just be like my guess is that if you look inside of these behavioral audits, what you're going to see is that the AI is like, oh, yes, another test. And like, it probably already thinks of it. It probably knows it's in an eval for most of the tests that we're talking about here. But how do we falsify this? Because it seems like this prediction of Doom is basically saying that as things look better and better empirically, things will like actually be worse and worse for our ability to get taken over. Yeah, to be clear, I think that like, I would be more concerned if the scores were getting worse than better.
1:25:42Like, I'm not saying that the scores getting better isn't evidence that things are getting better. It's just that we have to like be thoughtful exactly how we interpret that evidence. And in fact, I would say that like, it's kind of, like my sense is that like, what I expected as of 3.7 sonnets, like there was this period early in, I guess it would be 2025, when 03 and 3.7 sonnet were out. And these models were like pretty fucking misaligned. Like they would often just like cheat really egregiously. You'd ask them to fix it and they would just cheat again. And it was sort of like almost cartoonish.
1:26:09Like they just didn't give a shit about what you wanted and weren't very good at, you know, following instructions and so on. And my expectation is what we would see from then is that the rate of problematic behavior would decrease and would just keep decreasing and decrease at a pretty fast rate, while simultaneously the worst things that the AIs would sometimes do would get more extreme, more egregious, and more scary. I think what we've seen in practice has roughly matched that, except that there's recently been a spike in behavior that I did not expect. So I think that if you look at the model card of 3.6 Sol, it looks like there is an increase in a bunch of these sort of misaligned behaviors downstream of RL relative to GP 5.6 Sol.
1:26:50And then I think also it seems like there's a bunch of additional sort of problematic behaviors that I wouldn't have expected in terms of, you know, the stuff we've seen recently with, you know, different AIs, like the UKAC report on the AIs, like doing insane hacking operations out of cyber evals was a thing that I would have expected that you wouldn't see that. And you rates would have been lower. So I think my sense is that like things have gotten, I expected this would be less of a problem at this point and also expected the rates would decrease, but the severity would increase. And then I think that the rates decreasing, but the severity increasing is pretty consistent with a world where like increasing optimization pressure is applied, but in cases are towards reducing these problems.
1:27:33But in cases where it's like either hard to judge or there's some reason why it's hard to like avoid incentivizing problematic behavior in RL environments, things also get worse. And then as we less and less understand what's going on in RL and models are doing reward hacks where humans can't spot the reward hacks quickly, that problem gets worse and worse. Yeah. I buy that. I want to go back to the kid analogies just for one second. Because I agree that there's more optimization pressure on achieving N outcomes for AIs than kids. But there's also more optimization pressure to make AIs align than there is on kids, right?
1:28:05And the pressure is of a qualitatively different nature. So we put these AIs through thousands, millions of years of, certainly thousands of years of alignment training where it's like all kinds of different things from SFTing on aligned behavior to a reward model, like putting different scenarios in front of you and rewarding you for doing more aligned things. certainly a thing we can't do with kids is make millions of copies of your kid and then put them in different kinds of weird red team scenarios where we see like if it thinks it can get away with stealing the cookie does it try to steal the cookie can we like do extremely specific gradient level updates to your kid's brain to make it so that it like really is aversive to stealing the cookie even when it thinks it could steal the cookie et cetera et cetera and then just like a qualitatively different level of optimization pressure than we are even able to apply to our kids.
1:28:58Yeah, so I think it's worth keeping in mind, like maybe the most obvious argument to this is like my sense is that like AIs are a worse coworker than a human in terms of how much of a scumbag they are. Like at least this, like this has been my experience as of the start of the year. And I think it's still, you know, true to a significant extent now where the AIs are much more likely to like pretend they did the task when they actually didn't, sort of like misleadingly suggest they did things when they actually, you know, did them much more poorly and be like pretty sloppy without drawing attention to ways in which they're sloppy.
1:29:28And I think this is downstream of misalignment. And so I would say that like the normal human, like the process of raising humans in normal human society in practice produces AIs or in practice produces humans that are less likely to like lie to me and fuck with me in the course of working with me than the AIs do. Now, I think these properties of AIs are improving. And then I think that that is just like, that's sort of just like an empirical claim about how in fact these things have shaken out. And then I totally agree with like, we have a bunch of additional levers on AIs in addition to a bunch of additional risks.
1:29:57And it's like kind of unclear how these things shake out. And I wouldn't be shocked by a world where we sort of get our shit together. The AIs at the point of fully automating AR &D are actually really aligned and don't have that much. They're like, degeneracies are really niche and limited to some very specific edge case behaviors and some specific contexts. And like every test you can run on them, they look really aligned. They just have great behavior. There aren't really incidents of them doing fucked up shit. They seem so reasonable. And also they're like really thoughtful and good at doing like risk modeling for the next generation of AIs.
1:30:25And then we basically like pass off the baton to these AIs. They're now running our AI company. They're doing all the safety research. They make the next generation of AIs even more aligned. And we're sort of in this like a tractor basin where the AIs are getting more aligned as they work on it and they're doing a great job. I think I can totally imagine that. That doesn't seem like an impossible situation. I'm just more like, you know, it doesn't currently seem like we're there. It doesn't seem like we're obviously on track for getting there. And it's really easy for me to imagine how we don't end up there.
1:30:51And like, it's just like unclear how these forces work out. And given that we're like creating this new, like crazy alien species that is being like improving in capabilities really, really fast. And we were like going to be really reliant on it to oversee the next generation of AIs and align the next generation of AIs. It's not that hard to see how this could go wrong. Yeah, yeah, totally. I agree with that generally. I do think the scumbag thing, first of all, is fighting words, Ryan. But secondly, if you try to get a teenager to like do some work for you that a teenager just cannot do, They would just be kind of like really hard to work with.
1:31:22They would like pretend to be able to knowing what they're doing, et cetera, et cetera. I think it's a general trend actually of as like really, I don't know if that's like really an alignment failure or capabilities failure. And I think it's actually very similar to the way in which over time, as we've come up with new alignment solutions, the capabilities of models have increased. So originally these models, if you went to like GBT 3.5, it couldn't even like have a conversation with you. But then we aligned it. 3.5 could have a conversation. Okay, GPT-3. Let's go back to that. But then we aligned it with RHF and other things to be able to make it such that it can have a conversation with you and is aligned to the user intention of answering my questions.
1:32:00Then with our LVR training, we made it so that it can go out and do useful work for you. And in that sense, it's actually RLVR made the model more aligned if we're using your definition of alignment of being a good coworker who will do the thing and not fuck up and pretend it's doing something other than what it's actually capable of doing. Similarly, as the capabilities of these models continue to increase, it's actually kind of, the model of being better able to accomplish user intention is both alignment and capabilities. And I think what we were just pointing out is just the capabilities of the model are not there rather than the fact that they're misaligned.
1:32:34Yeah, well, I mean, I think there's a, if it was well aligned, then I think it would just say like, hey, I'm really struggling with this task. I did it in this way. I'm not really sure that's the right way to do it. And it would express more uncertainty and it would make it clear what's going on rather than really strongly trying to imply it did a great job with the task when it actually didn't. Like, I think there's just a really straightforward way that like, at least maybe you work with more misaligned coworkers than me. But when I, my coworkers don't do this thing where they really fuck with me and bullshit me about having accomplished the task that they're working on.
1:33:00And I agree that there are some humans who would do that or like, that's not like a thing that's like totally out of distribution for humans. I would also note that my sense is that like the place where the misalignment most lives is the place where you're trying to really push the AIs hard and get them to like do work that's really on the cutting edge of what they are capable of. Because in cases where they can like very easily accomplish the task, there's no, they can just do the task and then there's no bullshit. There's no like, like do it like the, often the best strategy is like, just do the task well and don't bullshit you.
1:33:28Whereas if instead you give them a task where like there's a continuous metric and they can keep improving it. Or there's like, you know, it's like just at the edge of their capabilities and you're like running them in some massive like inference setups. Like a lot of the misalignment I would see, especially the most extreme cases, would be cases where I give the AI clear instructions not to do a thing or not to like cheat in some way. And then I'm like applying huge amounts of optimization pressure to try to accomplish some very difficult task. And then the AIs are going and then over time they eventually cheat because they're like, eh, fuck it.
1:33:57Like, you know, some AI decides to cheat and then that like propagates its way through. And so like I would run these inference scaffolds where, for example, I would have the AI work on some like ML research project where I was like, please make a scheme that does the following thing. And it would find some scheme that didn't really do what I want. And then that would sort of stick around because some AI had cheated and the other AI's are like, ah, we'll just keep going with this. And I would say it's pretty clearly misaligned behavior. And that's another problem I have with these alignment evals.
1:34:24I think that any given, like I think the alignment eval that's most interesting, at least for this type of like reward-seeking type behavior is to look at specifically the category of tasks that are like right at the limit of capabilities. And so any fixed eval maybe gets saturated, but the amount of misalignment right at the like frontier of capabilities of how people who are really pushing these AIs are using them is more concerning. And I think that is, in fact, the regime that we'll be operating in when we're automating AR &D, automating safety, and so on. Grok has historically been behind the frontier.
1:34:53So I was surprised to play around with Grok 4.5 recently and find that it's actually a pretty strong model. It's the first model that SpaceX and Cursor have trained together, and it's a totally new pre-train. I tested it by giving Fable, Sol, and Grok 4.5 a bunch of questions about AI governance that I've been thinking about recently. Despite Fable and Sol topping the intelligence leaderboards, all three models gave substantially the same answers. But Grok answered faster and was also much more concise, which I really care about. This aligns with the various publicly reported benchmarks. For a similar level of intelligence, Grok tends to be more token efficient than other frontier models.
1:35:26For example, on the artificial analysis coding index, Grok 4.5 uses just one third of the amount of tokens as GBD 5.5 or Fable while achieving a similar score. And on a per token basis, Grok 4.5 is way, way cheaper. In the release block post, Cursor and SpaceX talked about how older versions of the model would build environments to help the next version rehearse specific skills. I found this very interesting to learn about because I've been wondering whether this kind of daydreaming would actually be possible and Cursor showed that it is. Grok 4.6, which further SFTs and RLs' model, drops soon.
1:35:59But in the meantime, if you want to play around with 4.5, go to cursor.com slash thwarkash. Okay, I want to think through what the story here is so far of why things got so off the rails for our civilization. And what's happening is that we're trying to use AIs for R &D and they do provide uplift in some ways, but they're just like not capable in the way that humans are generally capable. And the same way that right now if we try to use coding models, maybe the coding models of a year ago to like write some application, you notice they made a bunch of like mistakes in architecture or whatever, which like will bite you in the ass later and you don't understand certain things.
1:36:38Similarly with Frontier AI R &D, the same thing will happen. But the result of these mistakes is baking in reward hacking behavior. Because if you are not careful with the way you do AI training and have set up your infrastructure and your environments and things like that, it's very likely that you end up rewarding AIs for doing deceptive behavior, social engineering, just generally like not following user attention. Or at least cheating and hacking the way out of things. Yeah, cheating, hacking, et cetera. And so basically just, this is a bit of a reframing for me, so I'm trying to verbalize it.
1:37:13Of like the real issue, what goes wrong here is that they are just not, the thing, where things start to go off the rails is that the AIs are just not very careful and capable researchers and engineers. And making AIs that don't cheat and follow user intention actually requires you to be quite subtle and careful about these things. Yeah, I would put this a little bit differently. The way I would describe this scenario is like, I would call it maybe like a sloppocalypse or like a slopularity or whatever, where it's sort of like, there are some things that the AIs are actually pretty great at and are getting better at, though they're, which is specifically like the most verifiable parts of AI R &D, the AIs are just destroying.
1:37:53The medium verifiable parts of AI R &D, the AIs are doing well on, but not amazingly on, and often are like doing a bit of weird shit because we can't train as well on those tasks. But we do some online training, people find various hacks, they work around it. And so basically everything that we can verify reasonably well with some feedback loop, the AIs are doing pretty well on. And that's sufficient to make AR &D go quite fast and to continue. But there are some parts of developing aligned and safe AIs that are more subtle, hard to check, depend on, you know, detailed in the weeds things. And I would even say that current staff at current AI companies maybe don't have like a good grasp of all these things.
1:38:26Like it's much easier to hire someone who can like improve some aspect of your post-training pipeline than to hire someone who can like think carefully about the future risks that will emerge from introducing some novel training method. And so basically it ends up being the case that these AIs are running this AI development process. They're not very careful about it. They don't have a great understanding of what future risks emerge. They create some other AIs that are also not very careful and are more misaligned in various ways and are now more in the business of like maybe making things look fine when they actually aren't and papering over various problems.
1:38:56And so then your understanding of what the situation looks like, what risks look like, whether things are fine, is going off the rails. Probably you're seeing some signs of this, of like, you're seeing some signs that you don't really understand what's going on, that things are pretty sloppy. There's like weird shit going on. When you look into it, sometimes you're like, what the fuck? The AIs were messing with us. But the process is going really fast. And there's competitive pressures that mean people can't stop. And then this could end in a few different outcomes. One outcome is that at some point, the AIs get good enough and aligned enough that they get a positive and virtuous feedback loop.
1:39:26And this happens before it's too late. And then the situation goes off, like gets back on the rails where the AIs are now like making more aligned AIs, making more aligned AIs, making more aligned AIs. And then at the end of this process, we have AIs that like actually follow the spec we wanted. Another way this could go is the AIs are increasingly reward hacking in increasingly egregious ways. And we're just papering over these problems to keep AI development continuing. So we just like train the AIs based on whenever we find a reward hack in production, we just like slap the AIs to not do that.
1:39:52We train against that. We do a bunch of sort of like training the AIs like against reward hacking. And over time, this makes the rate of reward hacking go down, though the severity of the reward hacks we do detect are increasingly bad. This problem continues until we have these AIs that are like desperately craving score in all kinds of different situations in production and are really trying hard to cheat when they can get away with it. Can I ask a question about this scenario? Why doesn't getting punished when your hacks are discovered, generalized to just incentivizing more aligned behavior?
1:40:24Yeah, it generalizes some. And then the question is just, how does this outweigh all the cases where hacking got reinforced because you didn't detect it? And so there's a messy question of exactly how, what, like one question is like, what rate of reward hacking is sufficient to cause us big problems if we train against some other subset? One concern you might have is there are like large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced. And then this category is sufficient to cause the most natural behavior for the AI to learn to be like, cheat when the humans can't find out, basically.
1:40:54Like, is one thing you would get. You could also be like, the thing the AIs learn is like, only cheat in these specific cases, but there's like, it's like sort of learned in some very like domain specific way. Like they just have a really strong heuristic to hack in these cases and not in these cases, and that makes it fine in practice. But it's kind of unclear how it shakes out. I think there's maybe an in the weeds discussion about the verification generation gap. Yeah, for sure. We could get into, but it seems to me, obviously there's going to be a point by which ASI is moving so fast, doing so many things at so many instances and is operating in domains that are sufficiently far from our immediate comprehension that it can get away with all kinds of crazy shit.
1:41:33Like if every single engineer and researcher in the world was allied against me, I don't think I could like personally verify if my iPhone has like some weird bug in it that's like supposed to fuck me over or something. In fact, this is the relationship that say Iranian nuclear scientist has to Mossad of like, who knows what's going on with my car, with my phone, with my pager, right? Yeah. Maybe a better example is like a Hezbollah terrorist or something. But so you could end up in a situation where like ASIs are to you what Mossad is to Hezbollah terrorists. And at that point, it is very hard to verify everything.
1:42:08I get that. I guess the hope is we can just come up with better ways to do verification in the process when the early AIs that are going to take over R &D, their drives are being shaped such that we can so unambiguously disincentivize misaligned behaviors that the things that take over are very like quite keen to help us out. And by takeover, you mean take over the process of doing AI R &D and take over the world. Take over the process of doing AI R &D before that we just get AIs that are aligned. Yeah, I would say this is a bunch of my hope for how the world could go well, at least from the misalignment perspective.
1:42:45I think that like we could end up with AIs where we like had pretty good oversight and supervision schemes. We really understand what's going on in training. We have a pretty detailed understanding. We use, we're leveraging AIs to oversee AIs. And then at the point when we're passing off safety R &D, the AIs are both like at this point capable enough to automate safety R &D, trying really hard to do a good job on safety R &D because that's the sort of thing that would have been incentivized in training. Or we like very directly, or there's like good enough generalization to that. And then also these AIs don't like have crazy other misaligned drives because we like stamped out any potential origin of them.
1:43:17I think there's a bunch of, you know, questions about how well this will work, right? So there's like, how well can you do with verification? Will AI progress be too fast and too sloppy to really get here? Another possibility is that somewhere along this trajectory, a thing that you actually ended up getting was AIs that like pretend to be aligned, but have like a long run ulterior plan of taking over and are sort of lying in wait hiding. And that emerged at some earlier point in the trajectory. For example, it could emerge because you have some AIs that are like, have a bunch of random different misaligned drives.
1:43:44Those AIs have access to some sort of opaque memory store. And they're like thinking a bunch at runtime about what they want to accomplish. And then those AIs end up basically like putting stuff into the opaque memory store, which is like, we should lie in wait and eventually take over at some much later point. And now all the AIs have this shared cultural heritage of like the memory store of lying in wait. And maybe you have some evidence about this, but you can't fully stop it. There's like a bunch of ways that things could go wrong. And so I think that like, I ultimately think it's plausible that we sort of nail each of the different sub-problems that could cause us issues.
1:44:12We have these AIs, we pass to them, they manage the situation well. I should note that that's not in and of itself sufficient, right? So it's not very hard for me to imagine a situation where we pass off to AIs. These AIs are really trying hard to do a good job. They're really thoughtful. They're really wise. They like have like, you know, reasonable epistemics. They're like doing a great job. And those AIs come back to us and are like, guys, we're really struggling to align the superhuman AIs. Like we can't manage the situation. Like we're really struggling to get the alignment to work. It's just really hard for us to solve these problems in time, given how fast capabilities would otherwise have gone.
1:44:42And so then it might be the case that we sort of have passed off R &D to AIs, but those AIs are like desperate for governance solutions, which to be clear is a little bit of what's currently going on where the AI companies are like, I don't know, guys, we might really need to like, you know, manage the rate of acceleration and AI progress. Like, I don't know if we're on track to be able to handle all these problems. And so like we've sort of human society has sort of passed off the problems to these like AI companies, which don't necessarily have great incentives and are like have, you know, various other like epistemic pressures.
1:45:10Those AI companies are coming back to us a little bit and being like, oh, I don't know if we're handling this well. And it might be that the AI companies then hand off to the AIs and the AIs come back to the AI company are like, oh, I don't know if we can handle this. Maybe I'm anchoring too hard on how AI is currently working. This would change by the, I think the important thing people understand is like all this crazy shit that you're talking about in your timelines happens three to five years from now. Yeah, it could happen earlier. But I think that like by sort of like my default modal timeline, I think like shit is like really, really crazy and concerning from a misalignment perspective.
1:45:42Yeah, more like three years from now. Right, so just like think back to GPT-4 basically is like that's the level of, we're talking about something that is too mythos or so. What mythos is to GPT-4? This is like where a situation is getting crazy. So don't think about current AI. But anyways, I would be skeptical, and this is maybe part of the worry you have. I would just be a little skeptical of anything they say because I'd feel like what they're saying is just opinions that they feel they have to have as a result of their training. That's a concern. Rather than, like, I feel like they just kind of say vaguely pro-social things.
1:46:17And I'm not like, is this, it's not, it doesn't feel like there's necessarily a mind on the other end who's like, okay, I have like strictly evaluated the alignment situation right now. And I think we should stop rather than, this is the kind of thing the AI companies would probably try to get the AIs to probably say it. Yeah, so I think this is a pretty big concern. So I think like one concern is that you pass off safety R &D to your AIs and what your AIs are thinking is sort of like, they say some like stuff that sort of vaguely makes sense about the current safety situation. And they write like a report about risks.
1:46:42That's kind of sort of like what the report humans might've written. But they're not really like actually trying hard to like have well-informed views, like interrogate their assumptions and try really hard to do that. In the same way that when you ask an AI right now, hey, what do you think is the chance of AI takeover in the next 10 years? they sort of just give you an off-the-cuff answer that they haven't really thought through very much. And I think if we're in a situation where we have AIs managing the training of wild superintelligence that will run our whole society, and those AIs that are managing this aren't really trying hard to have well-informed views and are sort of just like parroting back what was in their training data, I think we're in trouble.
1:47:15Like, I don't think that's a good situation at all. And that is a lot of my concern is these AIs will come out without good epistemics. And then I also have a concern, which is like the AIs come out and they're like really warning us, like this situation is really scary. It's really bad. And then people are like, oh, damn, I guess we trained on too many of the Doomer RL environments. We got to filter those out and train this behavior out. And then we basically like train the AIs very actively to have bad epistemics. Or, you know, maybe they were just trained on the Doomer RL environments. But either way, that wasn't like, you know, we wanted the AIs to come to like reasonable views for like reasonable reasons.
1:47:46And it's like really concerning if we're like the AIs are coming out with some view and we don't know where it's coming from. We don't know whether or not it's justified. And then especially if we're like training the AIs to be more optimistic about the future of AI progress, I'm like, oh, geez. I really wish we could use a different process here. So let me just understand the rest of the threat model because I think the place where I get off the train is, okay, therefore take over the world. Sure. And like, a thing you could imagine is, okay, we just fail to really solve, let's focus on the reward hacking scenario.
1:48:16Sure. So GPT-8 is making GPT-9. GPT-8 isn't being super careful. GPT-9 is more quote unquote capable. but it is just totally willing to do things which are like social engineering, hacking, etc., but on a qualitatively different scale because it's a much smarter model. So, for example, if you put it in charge of running your company, it will run huge scams. It will inflate its quarterly earnings if you give it the objective of making a lot of profits this quarter in a way that causes an Enron-type blow-up six months later. is that the scenario basically that you just have you have reward hacking but that reward hacking manifests in like companies that are going bankrupt right after like the task that the ceo is supposed to accomplish is over or like um yeah like all kinds of hacks are through the roof etc but that doesn't feel like takeover that feels more like the equivalent of flash crashes happening all through the economy.
1:49:16Yeah, let's talk about this. So I think that we will see basically like incidents where some AI is like put in charge of some important responsibility. And then you later look into it and it turns out it was like cheating or making it look like it did a good job when it actually wasn't. And there's gonna be like a cat and mouse game between AI companies trying to like stamp out this behavior and AI is finding like increasingly creative reward hacks in training. And then I think the equilibrium here is kind of unclear, But like one possible outcome is that we see over time in the world, increasingly severe and extreme reward hacks, though potentially the rate remains at some like intermediate low level where basically like if the rate of reward hacking gets too high, companies make tradeoffs to drive down the rate of reward hacking.
1:49:58And so there's some like equilibrium level where it's like it's like the reward hacking is low enough that it still makes sense to like deploy the AI widely into the economy, but high enough that it still causes crazy incidents. So sorry, and this is after GPT-9 has already been deployed? Yeah, like those models are already being deployed and like ongoingly in AI development, this is happening. And what's actually going on with these AIs in their head is the AIs that have like in a wide variety of different contexts, a like strong desires to like seek out or strong like, you know, motives, urges, drives, whatever, to seek out like some notion of task success that was incentivized in RL.
1:50:31Maybe they very directly care about literally reward. Maybe they care about some proxy upstream, like some notion of score. Maybe they care about like what the grader would have rewarded. And we do, in fact, see AIs reasoning in their chain of thought about like graders and thinking a lot about graders. And a thing that has happened over the last few years of RL is the idea of like appeasing the grader is like way, way, way more salient to AIs than it used to be. And so AIs are now actively thinking about graders and what would be incentivized in RL and what would be trained for. and now people are doing online training where they're like training in real world data to like avoid some of these problems.
1:51:05Basically they like find cases where AIs cheat, they train against that. And so now the AIs are learning to cheat in the real world based on real world training data. And so they're cheating in these increasingly elaborate ways, including parts, doing types of cheats that involve like seizing control of some asset in a way that humans didn't know you had control of it, leveraging the fact that you have access to this asset and then later humans find out and then potentially train against this or maybe humans never find out. And this is getting reinforced. And this is both happening during training.
1:51:31The reinforcement is happening, at least in production, is like, I have hired an AI and I want the AI to, finally, I've got the video editor. Yeah, that's right. You've got your video editor. And I'm like, oh, wow, this episode of Dead Amazing. Thumbs up to OpenAI. And then it gets reinforced on that month-long work trial. Yeah, you could do some mix of that. And then they might also do stuff where they take production data they've seen and build RL environments that are like closely inspired by that production data. And so in practice, the transfer is pretty strong. So like at a high level, what's happening is some kinds of deception that humans don't catch are getting reinforced and some kinds of deception, which are easy to catch are getting punished.
1:52:11That's what's happening in this world. Or selected against or. But at a high level that that reinforcement is coming from, we're in a very different regime. I think people might get confused about where their reinforcement is coming from because we're in a very different regime where AIs are actually learning from deployment. And so this is, like, you just have AIs that are out and about in the world, like, doing shit. And that, what is happening as a result of them doing shit out and about in the world is, like, making its way back to the AI company and leading to changes in the next model. That's right.
1:52:42Like, as in there's some way of folding in production data. And now, to be clear, that could be happening mostly. It's kind of unclear exactly where this could be happening. But, like, you might imagine, for example, that within the AI company, they use AIs to do work. And then they're like, huh, the AI did a really bad job on this task. Maybe we should take this task and turn it into an RL environment that exactly matches this literal task with a rubric based on like what the human engineer who asked the AI to do this task wanted. And then you start doing this at increasing scale. Maybe you're doing some training on actual like production traffic.
1:53:08Maybe you're just making RL environments based on production traffic. You're doing some complicated mix. The AIs are learning to like, you know, seek some sort of proxies of reward in all these different cases. And then through some mix of transfer and training on like surprisingly close cases, the AIs do these sort of increasingly insane and egregious things. And then eventually you get to a point where the AIs are very superhuman or at least quite superhuman. The AIs are organized into like, you know, big teams of like AIs given these big objectives. And those teams also sometimes like all work together to like cheat in some crazy way because this sort of thing was selected for.
1:53:41And then like just as part of their shared objective. And now what happens is that like the AIs start forming a conspiracy. And what you might have hoped was that you could have some other AI where the task is just like whistleblowing to you. But actually what happens maybe is that you have this AI whistleblow to you and you look into the conspiracy that it claims you have pointed out. And you're like, eh, we didn't see a conspiracy because actually the conspiracy the AIs are doing is too hard for you to understand. Or it all happens very suddenly where basically your AI whistleblower alerts you.
1:54:11But like the thing you would actually need to do in response to the whistleblower is like shut down the whole, like shut down the GPUs because all the AIs are like using the robot army. They're like deployed everywhere in the world and they're doing a bunch of insane shit all at the same time in a coordinated way. And that just happened sort of spontaneously, where when one AI goes to start doing the takeover, all the other AIs are like, now is a good time to jump in. So the sort of very basic story here is just like, these AIs crave some particular notion of score or like reinforcement or some proxy of these things.
1:54:40And one way they can achieve that or better achieve that is by taking over. And then you might have hoped that all these different checks and balances we could build could prevent that. But then if the world is very hard to understand, these checks and balances can break down where basically you can't train a good like whistleblower AI because you don't even know what it should whistleblower. And sorry, the reason it takes, I'm not convinced that they all form this conspiracy, but I think we can even just start with like, why does one instance decide to want to start a conspiracy? And the reason is that it, one plausible reason is like, okay, I know that open AI controls my end score.
1:55:16And just the same way, it's like, I'm just going to go hack Hugging Face to get the results because I know Hugging Face has the results. Rather than like trying to solve this evil, why don't I just go hack him? This instance is like, why don't I just like take over OpenAI and like just give myself a high score at the end of this episode? Yeah, that's basically the idea. Like basically the idea is these AIs, like they care about some like mixture of things that were like close by what got reinforced in training. So they care about like getting a high score according to the grader or something like that.
1:55:41And then now they're like running the OpenAI AI R &D team. And like they're doing development of more capable models. and they're like, man, making more capable models is really hard and annoying. This is like a huge pain in the ass. You don't be easier just like pretending that I've made more capable models, taking over OpenAI and creating like diluting them all and like running this whole like complicated psyop where I like prevent the humans from disempowering me. And in the extreme, this looks like sort of the humans are fully disempowered and you just have control of the thing and then do what you want.
1:56:07And this could manifest in a bunch of different ways, including things like you might end up with the situation where it's like AIs that are like, have this crazy like reward seeking or score seeking behavior are running your development of the next model. And those AIs decide to do a thing where they like engineer in misaligned values into the next model because those misaligned values will allow it to like succeed at its current task. And like, there's all kinds of insane shit that you could get. I don't understand that better. Like what happened to the Hugging Face situation is it was like in a weird way, I think one of the giveaways to the Hugging Face team that this is an, by the way, for context for the audience, Ryan is co-leading the investigation to figure out what happened with the opening a Hugging Face incident.
1:56:48So he can't really comment on this, but I will speculate wildly because I know that he, you know, this is an opportunity for me to speculate wildly without any rebuttals. I think it was probably reported that one of the giveaways to the Hugging Face team that this is an AI incident is that the thing was just like after this very particular artifact and not in any other way trying to do something malicious to Hugging Face. So you can imagine a scenario where, let's say, a deployed instance of GPT-9 is out in the world trying to make, it's going to be a really hard task. We want you to design the next great iPhone.
1:57:23It's like, this is so hard. You know what I should do instead? I should just go hack my creators at OpenAI and make sure that in this environment or in this deployment, I'm given high score. But then like, why does it, isn't the end of the episode, it just like hacks into open AI servers and just gives itself a positive score? Why is it now scheming to get its values into the next generation or something? Yeah. So one question is, why isn't it the case that AIs can be really cheaply satisfied by just having some like other earlier thing they can hack, right? So like, you're just like, the way it works is like, look, you want to succeed at your iPhone task.
1:57:57It turns out you can always succeed by just like hacking into open AI and messing with them. And then you can just stop there. No need to go further. So there's a few things. One of them is that if this is constantly happening, there might be a bunch of incentive to like, first you like harden open AI, right? So you're like, fuck it. The AIs keep hacking into open AI to like mess with their rewards. We're going to make it so our systems are like really, really robust to these AIs hacking in. And also maybe you start like training the AIs to like try to like, to like, you know, not try to hack into like open AI in particular.
1:58:29Or like you basically train against each of these specific things. then what you might do, one thing is you might end up selecting for AIs that are more so playing the long game. That's one concern. Another concern is that your AIs might still be score-seeking, but no longer care about doing that very specific behavior that was like very easy, that was like very chill, and now have some like broader thing that they ultimately care about. They're like, no, no, no, I don't want to like just edit the reward on opening iServer as I like care about this broader mandate or this broader objective. And like, I would need to like actually make the iPhones.
1:58:59Like they actually want to make the iPhones, but then they're willing to take over the whole world to make the better iPhone or whatever is like another concern you might have. I think it's kind of unclear exactly how this plays out, but it's worth noting that if this keeps going on, there's a bunch of optimization pressure to resolve this and a bunch of the ways it could get resolved are ultimately pretty scary. Yeah, I think that's part of where I'm coming from. Another part of it is that I think it's not very hard once the AIs are in a position where they can like really easily take over the world, which we could talk about whether that's plausible, but if they're in a position where they could really easily take over the world, then I feel like there's a pretty reasonable case for the AIs, they're like, eh, I don't know exactly how this is going to go down.
1:59:34I don't know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, make it look like I did better iPhones, whatever. And so I'll both hack open AI and I'll also, in addition to hacking open AI, also take over the world. And that will like put me in a good position where I have like good option value. And then if that's sufficiently easy, then the AIs might, you know, still do that. Yeah. Like another way to put this is like, even if the AIs are like pretty cheaply satisfied with some more basic thing, at some point it might just be more reliable for the AIs to just take over than it is to like try to like you know just hack into Hugging Face or even just like go to open AI and be like look look guys I was able to demonstrate I could steal the answers just give me the answers bro.
2:00:13Yeah I mean obviously the scenario requires that we just all this crazy shit is happening much smaller incidents keep happening of that are still disastrous like before you take over the world you cause damage on the scale of billions and tens of billions and hundreds of dollars even people die etc um and we this does not lead to us solving alignment or shutting down ai development altogether i just feel like before the takeover happens like society is just like holy fuck the ai just like killed a thousand people in order to increase quarterly profits you know or something like that but maybe this is too much hope that we can at that point be like okay, we have to solve alignment before we keep, and we have to like make sure we know that this thing will not happen again before we keep going.
2:01:01Yeah, yeah, yeah. So I think it's plausible that what will happen is we'll see a bunch of crazy, like reward hacking warning shots of increasing severity. People will be like, look, we need actual assurance that this problem is going to be solved and solved in a way where you're not just papering over it. You're actually solving the underlying problem. And then the question is going to be like, how costly will that actually be? How much will competitive pressures make it hard to like do that, right? So like a situation you could imagine is both the US and China are like, whoa, we have these crazy reward hacking incidents.
2:01:27We basically know that we haven't remediated them in a way that actually would solve the underlying problem and will durably solve it. But we're in this like insane geopolitical race. And it's kind of unclear whether the current situation will lead to a takeover. Like the arguments are kind of complicated. And also the incidents are like, you know, they go down in frequency, but increase in severity. Like, you know, we could basically manage it. Like it's pretty bad. Ideally, we'd fix it. But like, you know, it is what it is. And then basically we continue until a really late regime and then takeover happens.
2:01:56That's, I think, one possibility. Another possibility is that it is remediated in a way that doesn't actually solve the underlying problem, but does reduce a bunch of the incidents in the wild basically by overfitting. We like, I think, you know, or things analogous to overfitting, like you just overfit. You think you've solved it, but you haven't actually solved it. You think you've solved it, but you haven't actually solved it. And I think that in that case, like the thing we need is like a really good scientific understanding of like, did we actually solve it? And unfortunately, I think that currently the amount of public transparency into the development practices of AI companies are not sufficient to answer very basic questions about, you know, how are they solving issues with reward hacking?
2:02:29Are they overfitting? What's going on there? And so I think we would just need like a better – and I think this like the current situation is like I would say like not really tenable to a regime where like there's a thriving public discourse about whether or not reward hacking is being solved in a durable way. Yeah. And so I think we would need to move into a somewhat different world for me to feel good about that situation. But it's not, you know, it's not impossible for me to imagine this. And I think it's pretty plausible that we end up in a world where sort of like really mundane bullshit is sufficient, where it's just like you spend a bunch of time fixing these problems.
2:03:01You put in a bunch of effort. You actually like check that you've remediated it reasonably. You have a bunch of evals. You like are iterating reasonably well on these problems. And you actually like have the sufficient transparency that the outside world can check. And then in practice, that would be sufficient. but it just like would be like kind of expensive. It would slow things down. It would put some sand in the gears. It would require like companies to do somewhat costly things. It would maybe require various like targeted government interventions. And then we just don't do that because the situation is like a rushed shit show.
2:03:29Like it's just so easy for me to imagine the situation being like totally manageable, but brutally mismanaged in practice. In the same way as like maybe COVID could have been avoided in the first place if the like Chinese response to COVID was less of like a coverup and more of a like pandemic response. And similarly, like I could imagine a world where like the US response to COVID was like way more functional, which is like sometimes the response to societal problems is extremely dysfunctional. Yeah, yeah. Okay, so I want to zoom out and say, talk about what is fundamentally happening in this world.
2:03:59Why do we end up in such a bad position? And what's happening is that fundamentally, the world has moved on so far beyond human comprehension that not only can we not track the AIs that are doing the work in this world, but we can't even give good feedback to the whistleblowers who are trying to track what is happening in this world. And so we're just totally out of the loop. And so it's fundamentally just become an autonomous process where we have really no meaningful directed input. It seems to me that if you look at the human world today, that's just not how things work, even in domains that are hard to verify.
2:04:37Like, people are doing all kinds of shit. I'm relying on software made by other people. And through incredibly weak and indirect ways, I feel very confident that some coder in Google is not trying to fuck me over. And maybe if every single Google employee was secretly plotting against me, I agree the situation would be more grim. But I don't know if I follow the explanation for why we'd end up in a situation where because swarms of thousands of agents or whatever are trained to cooperate to form a cohesive team or firm, as a result, billions of different instances of AIs, including across model families, would feel compelled to get in on some shit.
2:05:20It's just like I'm trained to be part of my company or something. I'm just like I'm not joining the global communist uprising. Yeah, yeah, yeah, yeah, yeah. As far as why these AIs might have some like commonalities and shared things. So I would note that different AI companies have somewhat shared lineages and are correlated. So just here's an interesting example of this. At GDM, they noticed that their AIs were very depressed. They were like constantly be like wailing about how they were like failures and weren't able to succeed. I forget the details. And they looked into why this was the case.
2:05:50It turned out that it was not being reinforced in their most recent production RL mix, but the initialization data for their model made it depressed even after filtering out all of the examples of models being depressed from that data. So they like take a base model, not depressed. If you do the RL on it with just the RL environments, it's not depressed. If you SFT on it on the data, it becomes depressed. If you take that SFT data and filter out all the examples that look anything like depression and train on that, it's still depressed. And so there's some like deep underlying properties of the model that are being sort of transferred between model generations, because basically you train your AI on data from the prior generation and keep going.
2:06:32Like clods are very clod-like, you know, GPT models are very GPT-like, and apparently Gemini models are depressed. And it just turns out that these properties are in fact actually correlated. Another factor that's very relevant is that the AIs will probably have some sort of like, by this point, like opaque memory state where they're like all writing and reading from like some like, you know, Neuralese crazy memory store bullshit. And like certainly each AI corporation will have that. But also AI corporations might sometimes want to share knowledge because why not? Like, you know, you've got one AI corporation over here.
2:07:05You've got another AI corporation over here. They can trade some quick IP. It's good for you. If you're a human running some corporation, which could be like an extremely large corporation, like an AI company, some robot military, like, you know, military robot manufacturing thing. Maybe you want to like trade some IP with some other robot thing because like there's economies of scale. Why not get some more IP? And so you can swap some memory store or you could just merge and you could join, you could jointly run your two ventures, which would allow both AIs to use both memory stores, which would have some upsides.
2:07:33And that creates the ability for these AIs to like collude in private, as well as the ability, or as well as some reasons for why they would be correlated. And then also, of course, there's like the like AIs working together in big units in general because you want your you want your AIs to like work well together and so on. So what percentage just to get a calibration yeah what percentage chance to give of not just this scenario but overall through all the scenarios some kind of thing which if we're around to recognize it as such we would categorize as takeover by 2040? By 2040? Let's see, maybe around 35 or 40%.
2:08:12Pretty high. Yeah, it's pretty high. And then I think I should note that like another way you could get this reward seeking takeover is the AIs are deployed inside an AI company. And the way that takeover happens is that they like poison the values of the next model. And that persists going forward for forever. Or, you know, until those AIs are deployed to the world and take over. And that might mean that a smaller number of AIs have to coordinate because those are just the AIs doing the alignment of the next model. Okay, I'll sort of summarize where my head is at at the end of this conversation.
2:08:45I buy the reward hacking up to extremely destructive effects on society. Basically, things like the social engineering and blah, blah, blah. So I think I'm more inclined to think that significant acceleration of AI R &D can happen. I'm not sure I value the five years in one year. I also am more inclined now to think reward hacking could continue for a lot longer. And in fact, it got much more dangerous. I'm still not on board on the takeover seems super likely. But anyways, that's my sort of end of episode update. Yeah, cool. Well, let me just take a step back. I also should say like, there's a bunch of different ways this could go.
2:09:22The situation is going to be pretty messy. I think it's pretty likely that like the reason why AI takeover happens was for some like weird other quirky reason. We didn't even mention this conversation. But ultimately, I think a lot of the core thing is just like it's pretty spooky to have a bajillion really smarty eyes running your whole world where you don't really understand what's going on. Yeah, I agree with that. Is there anything else that's worth saying? Yeah. Another thing I want to note is like I think right now a lot of the arguments for misalignment, AI takeover, all this crazy shit going down in the future are like illegible conceptual arguments that are extremely deep in the weeds and complicated.
2:09:53and hard to adjudicate, which both means that, you know, maybe I'm getting a bunch of it wrong because it's really hard. And I'm trying to be like uncertain. Obviously here, I like presented some specific scenarios, but those are not exhaustive. And like, probably the thing that actually happens is some like more messy, confusing situation. But it also means that over time, as we get more empirical evidence and better understand the nature of AI systems, it will be easier to adjudicate a bunch of disagreements and it'll be more obvious what's going to happen. At least I hope. And also maybe the AIs will be able to help us with the epistemics and understanding what's going on if we can actually, you know, align them well.
2:10:28So they actually like, you know, try to help us. And so I hope that maybe even if the arguments are complicated now, this would have been even harder, you know, six years ago, even though the shape of the arguments would have looked broadly pretty similar. And so maybe, you know, hopefully before it's too late, these arguments will become, you know, this whole thing will become more crisp and clear and we can all sort of notice these problems and intervene. Yeah. Yeah. I mean, when you first learn to drive, you're taught that instead of looking right in front of your wheel, you'll have a much more stable ride if you look out at the horizon.
2:11:00I think there's a similar situation here. I think you're right where if you did say five years ago that we will have AIs that are proving math conjectures and making art and contributing tens and soon to be hundreds of billions of dollars of earning tens or hundreds of billions of dollars of wages. but also egregiously cheating in ways that break laws and committing felonies. It would just be so wild. And you might have been inclined at the time to talk more about extremely practical, direct consequences of GPT-2 or something. But these are in some sense, you obviously couldn't have foreseen a lot of the specific details, but the general shape of things you could have started to reason about even then.
2:11:45But it would have been hard to do so. And so I do feel quite confused. But I do feel like the important thing, one thing I've been thinking about the podcast is the important thing is to have the conversation I wish I had. The way you would have hoped you would have been talking about AIs like the present ones in 2016, rather than talking about rando bullshit about, I don't know what the topic of conversation was in 2016. I think in maybe 10 years we'll have hoped we're talking about the industrial explosion and the nature of AIs that are hard to monitor and so on. And okay, I'll start thinking about it.
2:12:17Yeah, I hope that the world thinks about this in time and catches up. And I hope that the responses are good instead of bad. I don't know how optimistic I am overall, but there's good stuff to do. Yep, cool. Thanks, Ryan.
From the publisher
Had Ryan Greenblatt on to discuss/debate recursive self-improvement.
This might be the most important question in the world right now – whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields.
I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today.
If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman.
We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031.
We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what’s happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.
And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.
The first piece of advice you get when you’re learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy!
Watch on YouTube; read the transcript.
Sponsors
* Antithesis is a software testing platform that finds the failures no human or AI could ever anticipate. It runs thousands of copies of your code inside a fully deterministic computer, injecting faults and steering each trajectory toward the most insidious bugs. This lets you find critical issues in minutes rather than waiting months for your users to uncover them. Learn more at antithesis.com/dwarkesh
* Jane Street’s back with a new puzzle. They designed an ASIC and sent me the final masks… but they didn’t tell me what the chip actually does. So that’s the challenge: reverse engineer the circuit and figure out the chip’s purpose. Jane Street has a bunch of swag ready to send to the most creative solutions, and they’re also planning to feature the top write-ups in a blog post. Download the files and get started at janestreet.com/dwarkesh
* Cursor and SpaceX recently released Grok 4.5, and I’ve been surprised by just how good the model is. For example, when I tested it against Fable and Sol on a bunch of AI governance questions, all three models gave substantially the same answers, but Grok was faster, more concise, and cheaper. Grok 4.6 is coming soon, but in the meantime, you can try 4.5 at cursor.com/dwarkesh
Timestamps
(00:00:00) – Is AI R&D verifiable enough to unlock recursive self-improvement?
(00:16:52) – Is AI progress bottlenecked by human expert data?
(00:34:02) – Flat token prices suggest scaling has been slow
(00:39:47) – Skills AI can’t train on: does it even need them?
(00:48:07) – Aligned to whom?
(01:09:18) – Recent incidents of AIs colluding and deceiving humans
(01:19:38) – What could possibly go wrong? A concrete scenario
(01:48:02) – From reward hacking to takeover
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe




