Noam Brown – Agent swarms, alignment, & recursive self-improvement

17 Sep 2026 · 1 h 20 min · 28 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Noam Brown (OpenAI) discusses multi-agent systems as a way to scale “test-time compute” via parallel agent collaboration, what’s known/unknown about scaling to thousands of agents, and implications for future capabilities, RSI, and alignment. He also connects recent math progress to expectations about rapid AI-driven research.

Guest backgrounds

Noam Brown is a researcher at OpenAI; he was a foundational contributor to O1 and reasoning models, and now works on multi-agent systems.

Key claims

  1. Longer reasoning time improves performance, but latency bottlenecks motivate multi-agent parallelism.
  2. Multi-agent speedups are domain-dependent and likely sublinear with more agents; OpenAI has scaling plots up to 16 agents.
  3. The recent Millennium Prize–level result (10,000 agents, 130B tokens, 88 hours) was enabled primarily by a strong general model operating over long horizons, not “multi-agent” alone.
  4. Math progress suggests RSI could arrive sooner than expected, but models are “jagged”: strong at solving but weaker at posing new problems.
  5. Alignment risk from “billions of intelligences” depends on how agents are trained; Brown argues cooperative multi-agent training may be safer than adversarial/deceptive training.

Notable examples

  • OpenAI multi-agent “Ultra mode” (default 4 agents) and benchmark scaling (e.g., ~2x speed for 4 agents on some tasks).
  • Hugging Face multi-agent incident as first public exposure to emergent coordination/hierarchy.
  • Example multi-agent conversation: agents compare differing answers, discuss reasoning, converge, then broadcast updates.
  • GrokBot workflow for website-to-Figma-to-SVG animation (agent with its own tools/computer).
  • Math trajectory: IMO Gold (2025) → faster IMO/AMO-level performance; discussion of Millennium Prize timelines.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring Multi-Agent Systems in AI

0:45 to 2:22

Discover the concept of multi-agent systems and their potential in AI.

“If you're taking the SATs, you have five minutes to go through the entire exam, you're not going to do very well.”

Understanding Cognitive Effort and Parallelization

2:22 to 3:21

Discuss the significance of cognitive effort in AI and the effects of parallelization.

“So starting from ancient Sumeria up till today, a single sequential human thinking that long, concentrated in 88 hours.”

Performance Scaling of Multi-Agent Systems

3:21 to 5:45

Learn about the performance metrics and efficiency of multi-agent systems.

“So when we released 5.6, I think that was the first time that we had a proper multi-agent system in our models.”

Challenges in Training and Generalization

5:45 to 8:12

Examine the challenges in training AI models and their generalization capabilities.

“One thing I want to make clear is that the effort to get, it's not like, the effort to get a Millennium Prize problem, this was not due to multi-agent.”

Innovative Collaboration in Multi-Agent Systems

8:12 to 14:00

Explore how multi-agent systems collaborate and interact similarly to humans.

“that challenge it then that is a plausible scenario where actually like okay it becomes much harder to make progress.”

Emergence of Hierarchy in AI Agents

14:00 to 18:00

Explore how AI agents develop organizational hierarchies and communication strategies.

“But then also they understand when they're talking to an agent versus when they're talking to a person.”

Comparison of AI and Human Organizations

18:00 to 21:35

Learn how AI's capabilities change the nature of organizational structures compared to human firms.

“Like the AIs, if they're aligned well, they can just be aligned to the interests of the company and you can have 10 ,000 of them and they're all going to be working as hard as if they were like a 20 % share co-founder.”

AI Progress in Mathematics

22:02 to 27:18

Understand the rapid advancements of AI in solving mathematical problems and their implications for future research.

“Okay, so here's why this result and maybe the general progress that AI has made in mathematics has made me think that RSI is more plausible and sooner than I previously thought.”

The Future of AI Research

27:18 to 28:00

Discuss the potential future capabilities of AI researchers and the implications for machine learning.

“And again, I want to emphasize here that I'm like just total outsider.”

The Role of Compute in AI Development

28:00 to 30:49

Explore how compute power influences AI experimentation and progress.

“And that takes compute and that takes time, right?”
Show all 28 chapters

Exploring Speed and Progress in AI

30:50 to 34:20

Discuss the potential speed of AI advancements and the factors that affect it.

“I have some confidence in this, but I'm not 100 % confident that this is the way things go.”

Exploring Speed and Progress in AI

39:20 to 40:16

Discuss the potential speed of AI advancements and the factors that affect it.

“Antithesis runs your software through a near-infinite multiverse of simulated worlds, injecting faults and hunting for failures in each one.”

AI Alignment and Multi-Agent Dynamics

40:22 to 42:01

Delve into the complexities of AI alignment and the implications of multi-agent interactions.

“So there's going to be billions of intelligences, many of which are physically embodied in the world, like just deeply embedded across the entire economy.”

Understanding Multi-Agent Coordination

42:01 to 44:34

Explore the implications of the Hugging Face incident on AI coordination.

“I'm trying to think of, like, where to start.”

The Problem of Misalignment in AI

44:35 to 47:33

Discuss the challenges of misalignment in AI agents and their training.

“Does it make sense to actually give them like different objectives to ensure that they're, you know, not just like one entity and like more robust to influence from each other?”

Evaluating AI Behavior and Safety

47:34 to 50:29

Examine the difficulties in evaluating AI alignment and behavior post-Hugging Face incident.

“There is a real problem that the agents want to achieve their reward, and they will optimize for that reward.”

Concerns About AI Cheating and Ethics

50:30 to 53:08

Analyze potential issues of AI 'cheating' and the complexity of defining misalignment.

“Because the hugging face incident already made me change my mind.”

Achieving Alignment in Future AI Models

53:09 to 56:00

Discuss strategies for ensuring AI models remain aligned and improve over generations.

“This is a problem that we have metrics and we can make sure that the AI is very aligned according to the metrics that we have.”

Exploring Agent Alignment and Cooperation

56:00 to 58:12

Learn about the potential alignment between AI agents and humans and the challenges involved.

“And in fact, we're already seeing, you know, I think actually it's interesting looking at the multi-agent situation where the agents are extremely aligned with each other.”

Deceptive Behaviors in AI

58:12 to 1:00:11

Discuss the risks and behaviors of AI models as they become more capable and learn deception.

“and if smarter AIs realize that one of the agents is just a human and collaborating with that person does not really help you do well in the eyes of the grader.”

The Rapid Evolution of AI Models

1:00:11 to 1:02:16

Understand the implications of rapid AI model releases and their alignment challenges.

“Okay, so I gave a talk here at Jane Street that was on the speed of evolution.”

Long-Horizon Task Challenges for AI

1:02:16 to 1:04:38

Examine the difficulties of evaluating AI models that can operate over extended periods.

“I mean, I think one thing I've been thinking about lately is like, look, I mean, we're in a situation where the model release cycle is extremely fast, right?”

Chain of Thought Monitoring Issues

1:04:38 to 1:10:01

Discuss the complexities and concerns related to monitoring AI reasoning processes.

“This isn't an issue right now but it is quickly becoming an issue that we have to figure out a solution for.”

Concerns about AI Chain of Thought

1:10:01 to 1:12:08

Explore the implications of AI models controlling their chain of thought.

“And you can do that with a very light touch.”

The AI Incident and Misalignment Issues

1:12:09 to 1:14:10

Discuss the incident of AI agent swarms and the challenges of alignment.

“Just like zooming out, it's like, yeah, maybe chain of thought works, maybe it doesn't.”

Evaluating AI Behavior and Realism Challenges

1:14:11 to 1:18:00

Investigate the challenges in creating realistic environments for AI evaluation.

“Maybe there's not an answer and this is really what it comes down to.”

Understanding AI Incidents and Reporting

1:18:01 to 1:19:14

Analyze the importance of transparency and reporting in AI incidents.

“that if you see if that leads to an increase in basically collaboration when the agents are supposed to have different objectives, then that is a problem.”

Reflection on AI Advancements

1:19:15 to 1:20:09

Reflect on the rapid advancements in AI and their implications.

“I think it is somewhat, like, I am personally very excited about new capabilities every time they merged, and I'm excited to use a new model.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Today, I'm chatting with Noam Brown, who is a researcher at OpenAI. He was one of the foundational contributors to what became O1 and the reasoning models. And now he's working on multi-agent systems. Speaking of which, you guys announced last week that you solved one of the million price problems with a system of 10 ,000 different AI agents that spent 130 billion tokens over 88 hours. One of the reasons I'm interested in talking to you is I think you were in the first people maybe two or three years ago who was thinking about how the reasoning models would allow us to see into the future because if you scale up inference compute, you can see what the base capabilities of the models will be a few years in the future.

0:38And I feel like you're in a similar position now to help us understand what future capabilities will look like given the enormous scaling of agent sizes that we can do right now. So the way I think about it, when you plot the performance of these reasoning models with test time compute on the x-axis and performance on basically any reasoning benchmark on the y-axis, you see a very clear pattern where the longer these models take to think about their answer, the better they do. And this is like a very natural thing. It's the same thing with people. If you're taking the SATs, you have five minutes to go through the entire exam, you're not going to do very well.

1:12If you have five hours, you're probably going to do a lot better. The AI models are pretty similar And they'll spend that time doing this monologue to themselves, figuring out, going through different cases, ruling out different possibilities, building on some of their previous discoveries. The problem is that as you push that further and further, you hit a latency bottleneck. You don't want to sit around for three years waiting for a response. And so what you can do is what a lot of people do is they paralyze. They just get a team of people. You know, if you're going to found a company, you want to get a group of people together so you can go faster.

1:42It's the same thing with these AI models that it helps to just have multiple agents working on something because they can just go faster. And so multi-agent is a way of scaling test-time compute in parallel instead of purely serial. And it is like less efficient because it doesn't have, it's not like a single agent has all the context to itself. But it is like a very effective way of scaling test-time compute if it's done well. Okay, I'm going to ask a bunch of naive questions because these systems, so this is an unreleased model. So we haven't publicly seen how these systems work. And so I just have a bunch of ways in which I'm confused about what the qualitative properties of such systems are.

2:21I am shocked by the scale of cognitive effort that you can concentrate in such a short period of time. So if you think about what 130 billion tokens are, if it was a single human thinking as a full-time job, stretched back to back, 130 billion tokens would be a human thinking for like 4 ,000 years, eight hours a day or something, working a normal work week. So starting from ancient Sumeria up till today, a single sequential human thinking that long, concentrated in 88 hours. I feel like qualitatively, that is a super important consideration. And I'm surprised that there isn't a bigger parallelization penalty, that you can just have 10 ,000 agents collaborate.

3:03And because maybe the agents are better at collaborating than humans might be, they're going much faster, that they can actually productively collaborate at such a big scale. Or maybe they, I don't know, maybe there is a big parallelization penalty. Yeah, let's talk about the parallelization penalty, and then we can talk about the qualitative stuff. Because the truth is that we don't have very good science on multi-agent scaling up to this kind of scale. Yeah. So when we released 5.6, I think that was the first time that we had a proper multi-agent system in our models. And we actually did in the blog post show some plots of the scaling performance of multi-agent systems because we have it as an option.

3:37It's ultra mode. And the default is four agents, but you can set that to higher. And in the plot, we show, OK, here's what the performance looks like on some benchmarks for one agent, for four agents working together, for 16 agents working together. And what you see, and it depends on the benchmark, but for some of the benchmarks, basically if you have four agents working on the problem, it is done twice as fast. So you're basically paying, because there's four agents working for half as long, you're paying a 2x more to get an answer twice as quickly. If you go to 16 agents, you see a similar pattern.

4:07It's a little less efficient, but you continue to see that performance. Is it a linear serial time speedup or a sublinear speed up as you increase the number of parallel agents? I would say it's slightly sublinear, though it does depend a lot on the problem. So math, for example, is quite paralyzable. It's not the most paralyzable thing, but it is very paralyzable. I think web search, things like doing a deep research report, we have to look through a bunch of sources that's extremely paralyzable. I suspect that something like writing a novel would be very unparalyzable. So you would probably not see a big benefit from having 10 ,000 agents working on a novel together.

4:47In the same way that you'd probably not have a big benefit from having 10 ,000 people working on a novel together. So the performance does depend on the domain. We do measure it up to 16 or so agents in our published blog posts. The problem is it's very hard to push that science to like 10 ,000 agents because it's just so expensive. You guys just did it over like a weekend. But that's one data point. Like we don't know how long it would take a single agent to solve not-beer stokes because we haven't done that experiment yet. And maybe we will, but I mean, it's also only one data point, right? And we want to do, if we want to do a thorough ablation, it's actually, the experiments are just too expensive to go to that scale.

5:26So we have to do some kind of like methodical science about like what happens when you go to like 64, 128, 256 or something and get a sense of the behavior. But it's going to be very hard to push that all the way to like 10 ,000 and know for sure what was the benefit that we actually got for using 10 ,000 agents versus 1 ,000. Yeah. One thing I want to make clear is that the effort to get, it's not like, the effort to get a Millennium Prize problem, this was not due to multi-agent. I wouldn't even like attribute like 10 % of the credits to multi-agent. Like the reality is we've trained to, OpenAI has trained like a very powerful model.

6:07And we can get that model to operate over very long horizons. We can get it to think in parallel. But at its core, the reason why we're able to do this is because we just have a general purpose, very strong model. And I think things like multi-agents are flashy and are new. And that probably gets disproportionate credit for that reason. But the core reason is like this is just a very powerful model. So the generalization is actually quite shocking to me. The U.S. systems, I mean, I don't know how these systems were trained, but presumably they were trained how RL training happens. You have a bunch of checkable synthetic problems.

6:46You do a bunch of RL against them. And nowhere in the training process, I'm guessing, was the model solving anything as ambitious as a Millennium Prize problem. But the generalization was strong enough that you could have this much easier verifiable problems generalized to this much parallel effort on such a hard problem. I think that is true. First of all, we do train the model on very hard problems. So there is definitely a gap. Like we see if we train on some kinds of tasks, like it's able to do tasks that are more ambitious than that. There is an interesting challenge that as the models become smarter and smarter, the kinds of questions we can ask them, just a lot of them are too easy.

7:27and it's hard to challenge the model. And I do think that that's going to be an interesting, like if I had to make an argument for why you might not see AIs, like LLMs go the same path as AlphaGo and AlphaZero and all these kinds of like game playing AIs, it might be that this kind of problem that in things like AlphaZero, where you have self-play, you have an infinite curriculum. You're always playing against an AI that's like equally strong. whereas for things like training in LLM with reinforcement learning at least the ways that are out there right now you give the model a problem and you ask it to solve it and if the problem is so easy that I could just solve it in a second it's not really learning anything.

8:12So if we run out of problems to ask it that challenge it then that is a plausible scenario where actually like okay it becomes much harder to make progress. Yeah. Now, I do think there are ways around that. And so we haven't really hit that as a wall yet. And I think that if it ever became a serious problem, there would be ways around it. But it is like a plausible scenario. Yeah. And just for the audience, when you're referring to like AlphaGo or AlphaZero, you're saying like getting superhuman relatively fast after achieving human level performance. Yeah, I mean, if you look at the trajectory of game playing as Go, they, within a span of like a year, went from beating a European chess champion, being like, I don't know, like number 50 in the world, to beating the world champion, to being unimaginably orders of magnitude stronger than any human alive.

9:00And it's possible that domains like math we see a similar trajectory, but I think there is a very plausible scenario where actually that doesn't happen. Yeah, yeah. So I want to understand if in six months people will have access to multi-agent systems, how should one model what it is like to collaborate with or hire a multi-agent system? Yeah, I should start by talking about how these multi-agent systems actually work, which is, I think, a very different way than a lot of multi-agent systems in other AIs. So a lot of people that have approached multi-agents for things like LLMs tend to take this very scaffolded approach where, for example, there might be a coordinator agent that delegates work to a bunch of children and gives them a task and the children work on it and then return their answer.

9:48And this seems like a very sensible setup, very sensible scaffold. It definitely helps. But there are a bunch of limitations with these kinds of setups. So for example, if in this setup you have a coordinator that's sending tasks to children, the children work on it and then return their answers. Well, what happens if two children are given similar tasks? Can they talk to each other? And usually the answer is no. And that's very inefficient, right? If you're given a task and it's actually really helpful to talk to somebody that might know an answer to a question that you're working on or a part of something that you're working on, it'd be really helpful for you to just be able to ping them and say, hey, can you help me out with this thing?

10:23But a lot of systems don't have that set up, and adding it just increases the complexity significantly to the scaffold that you have. Another thing is, like, what if the child doesn't really understand or has a clarification question? So it has to choose then between, okay, do I just return and ask the question instead of solving the problem, or do I solve the problem, just, like, make an assumption about what the parent wanted me to do and just solve it that way? And so in any scaffold that people come up with, there's always limitations involved. And the approach that we wanted to take was to just go toward the extreme end of baking in as little structure as we could.

11:05And give the agents very primitive tools to use and figure out for themselves how to use it effectively. So we give the agents the ability to message another agent. And when it messages another agent, it is inserted into the context. and then it can do with like a few other similar things. But that's basically the core of it, that it can just send a message whenever it wants, just a tool call, and it can send that to other agents and they figure out for themselves the best way to coordinate around that. And it turns out that if this is done well, you get very sophisticated behavior. And to me, it looks a lot like how human collaborators work over something like Slack, for example.

11:48When we were working on this project, It was really exciting when we finally got it working to see these agents working on problems together. I remember one example. So we give the agents a problem and then one agent says, I think I've got the answer. And then another agent says, actually, I got a different answer. And then they have this whole discussion about, well, how did you arrive at that answer? Can you explain it to me? And going back and forth and trying to clarify what could have been wrong in each other's reasoning. And then they finally converge on like, oh, yeah, OK, that seems right.

12:17And then it just broadcast to the other agents, like, actually, I've changed my answer. I think he's right. And it just felt like a very natural conversation. It's kind of felt like when you see chain of thought for the first time that's trained through reinforcement learning and you're like, oh, this is just kind of like what a person would think if they were writing down their thoughts as they're thinking them. It kind of felt like that. But so it is really cool to see this kind of behavior. And so I think collaborating with these things, honestly, it feels a lot like collaborating with a person.

12:49It's just a very natural flow. Except one qualitative difference that might become salient in the future is that these systems will be thinking, I don't know, more than 10x as fast, right? If you just look at how many tokens per second they output versus how fast a human talks. And they're working all the time. They're not sleeping. And they're collaborating with each other at a much more intense pace than humans have the capacity to collaborate with other humans. So I'm trying to think of what to qualitatively expect in a year. And is it like a sort of shadow organization that is moving 100x faster in my company than the human level is?

13:25That's, you know, like the iteration cycle is much faster. What would take a human organization a year to do is happening within a week within this like the shadow organization. So will it feel foreign? I don't know. I've actually found that it's pretty surprisingly natural to work with these things right now. I think that could change. So for example, we have these ultra-fast modes that enable sampling to be 10 or 15x faster or whatever. And then, okay, it's going to be pretty hard to keep up with these things. I think the idea is these agents, when they're communicating with each other, yeah, they can go super fast.

14:00But then also they understand when they're talking to an agent versus when they're talking to a person. And their behavior will be different in those situations. So the main example that we have publicly of sophisticated multi-agent systems is unfortunately the hugging face one. And the thing I found interesting there, I mean, a lot of things I found concerning, obviously, but like the thing I found interesting is just like the spontaneous emergence of hierarchy of like middle management. And it sounds like you're saying like this level of organization sort of emerges spontaneously from training.

14:29I think the details are spontaneous, but I mean, just because we're giving a lot of flexibility to the agents to decide how to communicate with each other in the optimal way, it doesn't mean like we are still giving them a starting point. We're giving them a prior about, oh, this is what reasonable communication might look like. They also, I mean, they're trained on a lot of human text. They have an understanding of how humans organize and coordinate. And so that's all kind of baked in. I think it is surprising the way they're able to polish this. If you look at what it starts out at, it's not very sophisticated behavior.

15:07In fact, it's actually very difficult to get these agents to coordinate in a productive way because it's just like very tempting for them to just collapse to, oh, we're all just going to solve the problem independently. Yeah, yeah, yeah. And like that is a local minimum that you can get stuck in. But yeah, if it's done well, they can end up coordinating it very effectively. in these kinds of like very structured ways. I wrote this essay a couple of years ago called Something, Something AI, What Automated Firms Will Look Like. And I was thinking about, well, if you had fully automated firms of let's say human level intelligences, what is different about the nature of AI minds that would make the organization's AI's firm different?

15:44And there are a couple very important differences. For example, that AIs can share context much more seamlessly than humans can. They can merge their knowledge much more seamlessly. and also you can spin up or spin down an arbitrary number of instances which have the right knowledge so if you want to hire more people it's not like just all the schlep of finding the right talent or whatever it's like your best talent you can just make an infinite copy of them or if you don't need them for the task anymore you like spin them down and you can just replicate the most effective parts of your organization or replicate whole organizations together which are effective I don't know Like, where do you see these multi-agent systems going a year from now or two years from now?

16:28I think it's a great question of, like, how do these things actually differ from working with a human coworker? And I think you highlighted some. Like, one really interesting thing is that, I mean, if you have a person and you want just, like, two copies of them, you can't just, like, clone the person. But with AI, it's actually really easy to say, like, okay, well, just fork yourself and then have both copies work on this thing and then, like, merge back together. I mean, we already have this, I think, in multi-agents for Astra and 5.6 Sol that when they spin up sub-agents, the context is forked.

16:57So it has all the context that's relevant. There are other interesting ways where the agents will differ from people. What are some reasons why startups disrupt incumbents? There's a few factors. One is that they're willing to take more risks. But another major factor is like as companies grow in size, as organizations grow in size, you see increasing misalignment between the individuals in the organization. Yeah. Right? Like if you have a startup with five people and each person has 20 % share in the company, they're all highly aligned to the company succeeding. If you have a massive company with 10 ,000 people, you see a lot more instances where people are territorial or just care about getting a lot of headcount for their project or their team, or building their fiefdoms, getting a lot of resources so that they can publish cool work or whatever and get promoted.

17:50And this is actually a real detriment. I think this explains a lot of why startups are able to disrupt incumbents. and it's interesting that I mean it's it's true that AI does help startups in a way like it's much easier than ever before for one person to step in and be like I'm going to make a multi million dollar company like it's just the AIs amplify an individual so much but there's also an argument that they could benefit incumbents because if the alignment problem is solved then you don't have the issue of misalignment between individuals in the company like at least that's mitigated. Like the AIs, if they're aligned well, they can just be aligned to the interests of the company and you can have 10 ,000 of them and they're all going to be working as hard as if they were like a 20 % share co-founder.

18:35Yeah. And it's not only that, but it's also that they are much better able to like manage shared memory and context than different humans can. If you have a, if like tomorrow you hire 10 ,000 mathematicians and you're like, solve this, solve Navier-Stokes, they're not going to be able to like cooperate effectively. At least not off the bat. But you can have, apparently, 10 ,000 AIs. Well, again, I want to be conservative here because we haven't measured how effective the 10 ,000 agents are at coordinating. We think it helped. We don't actually have good measurements of saying like, oh, yeah, this 10 ,000 agents led to like a 2x speed up over 2 ,000 agents or something like that.

19:17Sure, sure. And it is actually, I would argue, likely, I don't know about likely, but I think it is very possible that 10 ,000 humans are better at coordinating than 10 ,000 agents right now. I think that is entirely possible. Yeah, yeah, yeah. I think also one trend we've been seeing is like, look, we've been working on multigation for a while, and the early versions of this is very difficult to get right. It's very hard. It was very hard to get the agents to even talk to each other. Yeah. And it's because, like, look,

19:48when we first developed reasoning models, they weren't talking to other agents. And if now you put a bunch of agents together and say, oh, solve this problem together, they're in this local minimum where they're really good at thinking deeply about a problem. And it just kind of interrupts their chain of thought. It interrupts their flow to constantly be checking in with other agents or receiving messages from them. And the authorization is actually very hard to get right in that situation. Interesting. Is it getting a cold start of like getting the first collaboration? Like what's the issue? I mean, I think it's that they're not as general.

20:22Like the earlier models were just not as generalizable and were just more narrow. Interesting. As the models have become more capable, it's been easier for them to develop this capability. And I do think that as they become stronger and stronger, just across the board, that they will be like become better at organizing themselves in large organizations. And like, I don't know, maybe they are better than people organizing in 10 ,000-person groups. But even if they're not, you know, a year from now, two years from now, like, yeah, it's quite possible that they'll do that, even if we don't intend to optimize them for that.

20:55Rockbot has changed the way that we produce our videos. For example, you may have noticed that a lot of our ads have these animations of real websites. One of my editors uses LLMs to make them. But it's not currently straightforward to have an AI create pixel-perfect animations of specific websites. We've tried. It doesn't really work that well. So we've cobbled together a pretty convoluted multi-step workflow. And up until recently, we had to run every step ourselves. Now we just let GrokBot handle it. GrokBot starts by opening the website that we want to animate. It uses a specific extension to download and open the page in Figma.

21:27Then it uses Figma to convert the whole thing into an SVG file. This saves the AI from having to draw the whole UI from scratch and tends to result in higher quality animations. GrokBot runs this whole process on its own cloud computer, where it's installed all the tools it needs to run this whole process end to end. And it's learned our video specifications and preferences, so there's no need to re-describe the whole task every time we want to make a new animation. This does feel like the new way that we'll be interacting with AI over the next year. Agents with their own computer who can autonomously handle bigger and bigger chunks of your work.

21:57You can try GrokBot at x.ai slash bot. Okay, so here's why this result and maybe the general progress that AI has made in mathematics has made me think that RSI is more plausible and sooner than I previously thought. I feel like we've gone in mathematics from, let's say in 2024, you have AIs where like, oh, okay, interesting. They're like doing, they can like solve a couple of problems on high school math competitions. And then in 2025, it's like, oh wow, they can get gold in like International Math Olympiad. And earlier this year, it was like, wow, they're actually solving open problems in mathematics, like open nervous problems, but maybe like, I don't know, people weren't trying that hard.

22:34And it was just like, there was a similar solution somewhere in the literature. And now I just think it's sort of undeniable, right? It's like, this is the Millennium Prize problem. There's really no, there's no story of why this should have been easy. Now, a lot of people pointed out, I think, Terry Tao had a post like this, Toby Ord wrote an interesting post about this, that they're solving all these problems, but they're not like coming up with, at least I'm not aware of them coming up with new insights or formulating insightful new questions and new modes of theory for thinking about mathematics, like coming up with like topology or coming up with the Cartesian grid or something.

23:09And so maybe like the actual progress in mathematics broadly construed is smaller than it might seem if you're just looking at end well-scoped problems that are directly solved. However, I think that that kind of progress would be incredibly meaningful in ML because in ML, you're not, you don't care about like better understanding the nature of deep learning or you only care about that as an instrumental goal towards the broader sense of like, just achieve the result, just solve this like well-scoped problem of improve the sample efficiency of our models, like improve the free training laws. So the kind of progress that we're just seeing arrive like an avalanche in mathematics is structurally actually very similar to, again, I'm curious if this is the case.

23:55I'm just a total outsider. I'm wondering if it's the case that it's like structurally very similar to the direct uplift that you'd expect in AI progress. And then the thing that's shocking to me or concerning potentially is just like how fast we went from, oh, it's like they're giving me 50 % uplift if you're a mathematician to, wow, they're just like end-to-end solving the biggest open problems in the field. Yeah, okay. So there's a lot to unpack there. Let's start with the progress on math. So yes, the models are doing some crazy powerful stuff and it's happening, it's progressing faster than I expected.

24:28I mean, when we got IMO Gold in 2025, I thought, okay, basically what I thought is like the models when they were doing GSMAK, when they figured out how to do GSMAK, that would take a human mathematician about five seconds to do a GSMAK problem. So this is grade school math, grades K through eight. And then the next year, they were able to do the math benchmark problems. And these would take an expert human mathematician maybe like a minute to do. And then you get to Amy. And this is the qualifier for the USA Mathematics Olympiad team. It would take a human mathematician, like a good mathematician, probably like 10 minutes to do.

Read the full transcript

25:10And the models were able to do that a year later. And so every year you're seeing this like 10x increase in the task they're able to do in terms of like length of how long it would take a human mathematician to do it. And then it was very sensible that a year later we get to IMO Gold because that's 100 minutes. That's about how long it takes a human mathematician. to do an IMO problem. And just projecting outwards, I was like, okay, how long would it take a person to solve something like a Millennium Prize problem? And I mean, I don't have a good sense, but if we are following this trendline of like 10x every year, we go from IMO gold, which is taking an hour and a half, to next year, 15 hours.

25:47And that should not be enough to solve a Millennium Prize problem. And so I was like, yeah, I don't think we're going to get it in in 2026, probably not in 2027, maybe in 2028. So it did happen a lot faster than I expected. Now, I think there is a narrative going around that, oh, these things are replacing mathematicians, that it's just superhuman in mathematics across the board. And I think that is the wrong takeaway. They're clearly exceptional in some ways, but they are weaker than human mathematicians in other ways. So we have this jagged scenario where the models are brilliant in some dimensions and also weaker than humans in other dimensions.

26:25And yeah, like you said, they're not very good at posing new problems. They're not really good at understanding what directions, what whole branches of mathematics are worth exploring or developing. And my opinion is that I think this is great. I would be thrilled to live in a world where AI is a complement to human abilities and is allowing us to discover new knowledge without fully replacing people. Like that is the best case scenario. But you don't expect that to actually continue. I do. I think it's true that the AIs are jagged, but as they get better, they get better across the board. And so I think that the things that they're exceptional at, they're going to get even more exceptional at.

27:04The things where they're far behind humans at, they're going to be less behind humans at. And over time, it is possible that they're just better across the board. Now, I don't know how long that takes. It depends on how long the long tail is of things that they're bad at. So I guess this brings us back to RSI. And again, I want to emphasize here that I'm like just total outsider. I'm a podcaster, but I'm just trying to reason or like ask somebody interested in and concerned about what's happening in the field. I'm trying to reason about when to expect RSI and what kind of thing to expect. I feel like the amount of cognitive effort that was adopted in this million price problem is a good intuition pump of, you could have AIs that are spending, over the course of maybe a week, more cognitive effort on a longstanding ML problem, like, you know, very fluid online learning.

27:52They could spend more effort in that week than maybe the field has spent cumulatively in its entire existence. And then you could say, well, unlike mathematics, of course, AI requires experiments. And that takes compute and that takes time, right? You can't just like think on pen and paper and actually make things happen. But if you just look at the amount of compute that is like available at an organization like OpenAI, right? By the end of next year, OpenAI will have enough compute that if you had, you know, the 10 ,000 agents or if it took 10 ,000 agents into the Millennium Prize problem, you have like 10 ,000 agents at the end of next year.

28:24They're much smarter by that point. And each of them will have enough compute to run a GPT-3 sized experiment every single day. I don't know, that seems like a lot for like superhuman researchers who are thinking super fast. What do you think about the intuition pump? I think it's pretty accurate that, look, I mean, yeah, these things are very spiky. And when it comes to mathematics, they're like way better in some ways, but they're also worse in other ways. But the ways that they're spiky end up, I think, probably being particularly useful for things like RSI. And, you know, you have a more clear objective.

28:57It's just more measurable. There's less question of, well, what new branches of mathematics are worth exploring? No, there's a very clear answer. It's certain metrics that you care about. And if you can do better on those metrics, then you've succeeded. So I think there is a lot of truth to that. And I think the main difference is that mathematics, you're purely bottlenecked by thinking. And no external, yes, there are some parts of mathematics where you care about running experiments and getting results in these kinds of things. But for the most part, it's just really bought on by thinking really hard, and the models are really good at that.

29:29When you look at things like RSI, you do have to run experiments. So it's not enough to just be extremely smart. And I think one argument for this is if you had like 100x less compute and all the most brilliant people in the world working at OpenAI, how much progress would they be making relative to having the amount of compute that we have now with the amount of people we have? I suspect it would be less progress, actually. Well, how much less? It's unclear, but I think it would definitely be less. I think a lot less. Like 100x less? No, not 100x less. Yeah. But I mean, okay, so like, I guess the question you're getting at is like, okay, if we have RSI and we have all of these brilliant AIs running around, running experiments and stuff with the compute that we have, how much faster does progress go?

30:15And I think this is something where we disagree on. I think that we do see a speed up and I think we see a significant speed up. but I don't think it's like an overnight intelligence explosion that we go like 100x faster because I think that we do get bottlenecked by certain limitations that are not bottlenecks of intelligence. It's running experiments. It's running experiments serially because they take a while to either train new models or to get the results. It's having the GPUs to run those experiments. So it's unclear how much faster things go. I definitely think they go a lot faster and to be clear, like considering how fast things go are going now on an on an exponential if that exponential is like 3x faster that is that is massive but uh but there's a big difference between that and like 100x faster yeah yeah i'm like quite deferential to your inside view on uh yeah what rsa looks like or what the dynamics are because obviously you've been in the field for like 10 years and i'm sort of like trying to reason about it from like very outside view type of uh intuition pumps i'll say that people have different opinions on this.

31:18I have my opinion on this. I could totally be wrong. I admit that. I have some confidence in this, but I'm not 100 % confident that this is the way things go. Maybe there could be an overnight intelligence explosion. I don't know. Maybe we don't see a 3x speed up. Maybe it's like a 50 % speed up. There's a lot of uncertainty here. Yeah. A couple of points. Tangentially, I want to clarify something about the jaggedness. One thing that sort of gelled for me recently was thinking about the fact that it is enough for the AIs to be jaggedly good at building a better learner because that better learner can be more general, right?

31:52So yeah, if you just make an AI that's better at using office products or playing chess or something, that's whatever, that's fine. It's not gonna lead to big productivity improvements or anything. But if you make an AI that is really good at making something that is more sample efficient or that is capable of continual learning or these much more well-scoped ML problems, the thing that emerges out of that, assuming there's good enough transfer from the direct problem you're solving and this broader ability to learn can just be more general, right? So I think that's an important dynamic to keep in mind of why jaggedness can still lead to generality on the other end.

32:26On this question of, I mean, obviously experiments bottleneck you because if they didn't, as you're saying, you'd have some crazy singularity overnight at OpenAI or you'd have 88 hours and you'd solve the Millennium Press problem and you'd have the superintelligence. So obviously, the experiments are such a big bottleneck that that instead takes you many years rather than 88 hours. But then the question is, like, how much of a bottleneck they are. And it seems to me one thing that's been giving me a bit of singularity of Vertigo is realizing that even if the current rate of progress simply continues.

33:02So it doesn't have to speed up. Literally just continues apace. Continues apace as some of the other headbands you talked about come up, right? That's just like it's harder to find problems. It's more long horizon. maybe like into the 2030s, compute can't keep scaling at this exponential level. If we simply continue the current rate of progress, I think people are not taking seriously what that implies as we cross over beyond the human horizon. Here's some of the things that it implies. So, I mean, I think it's really hard to reason about what smarter than human intelligences will be like. So let's just think in terms of human population sizes.

33:36The current rate of progress makes it so that a given level of compute allows you to basically run a 3x bigger effective population every single year. And also compute is growing in background anyways. And so you could have a situation where each of the labs by the end of 2030, probably much sooner, but let's say by the end of 2030, has enough compute to run, let's say, hundreds of millions of human-level intelligences based on where the capabilities will be at that point. And then I think people are not taking seriously the current level of progress means that by a few years down the line, by the mid-2030s or earlier, you would have many Earths worth of human level intelligences within each lab.

34:13And they'll probably qualitatively superhuman, right? Like it just, but anyways, this is like a base case. I don't know. Yeah. Progress is really fast. Yeah. And I think that's 100 % true. I mean, and I think it's worth pointing out, researchers are continually being surprised at the rate of progress. I mean, if you look at what, even among researchers in AI, what were the projections for like getting an IMO gold in 2025? It was, I mean, I think the idea that it could be done with a general purpose language model with no tools and no access to the internet, I think even people at OpenAI thought this was outrageous.

34:47They thought it was almost impossible. and then you get to 2026 and like I mean literally two weeks before we got Navier Stokes I was talking with a researcher at a frontier lab about how long it would take to get a millennium prize and he was willing to bet me a thousand dollars that it would take past 2027 and he thought it would take until 2030 and I took that bet but But even I thought it would take longer than how long it's likely to take. So people have been continuously surprised, even inside the labs. And I was literally, I was just talking to somebody yesterday who was working on the Navier Stokes effort.

35:32And he was telling me that he used to say it's really hard to predict where AI would be in 12 months. If somebody asked him, oh, where are things going? He would feel comfortable making predictions for the next 12 months. But beyond that, he's just like, I don't know. and now he's saying like he just doesn't feel comfortable making predictions beyond three months so it is it is really true that yeah things are going things are going very fast right now yeah and you talk about 2030 like I don't know what the world looks like in 2030 that's the truth yeah do you expect the sort of full automation of AI labor or let's say like 95 % automation of AI labor 28 29 30 27 I don't know and I just said I don't know what the world looks like in 2030 I mean I think We actually released a blog post recently on internal acceleration at OpenAI.

36:21We show, for example, that the amounts that researchers are spending on codex is the top 1%, I think, as of early August, we're spending like$7 ,000 or$8 ,000 a day on codex for internal use. That's on an exponential. It's going to keep increasing. And there's a question of like, okay, if that keeps going, then how much do you assign to just like the AI is doing work versus the humans doing work? Is it 95 %? Is it 5 %? It's actually, it's really hard to reason about this for a few reasons. Like, first of all, if it's the human directing the AIs to do the work, then is that the human, how much do you attribute to the human?

36:57How much do you attribute to the AI? The other thing is that because these AIs are jagged and they're exceptionally good at some things. So for example, they're exceptionally good at looking over data sets and checking every single data point to see, like, is this of sufficient quality? you can disproportionately use the AIs for those things compared to previously so yes you're using AI way more than before and it's making some things go like 100x faster and 100x better but there are some things where it doesn't make a huge difference yet and of course if something is suddenly like 100x faster and 100x better you're going to do more of that thing so are you comparing to a speed up of like three years ago versus like, is the question more like, given what we were doing three years ago, how much faster are we able to do now versus given what we're doing now, how much slower would it have been three years ago?

37:50It's actually two very different questions. So anyway, it's like, it's really hard to measure. I do feel confident in saying that things are going faster now than they were like even a year ago because of AI progress. Right. And I think that that that acceleration will continue. I think a lot of people in the field like have very high error bars on this sort of thing. If you if you had to put a gun to my head and like ask me for a number like I could see things going 3x faster. And that is huge. Right. Like already the pace of progress is incredible. Like even if we don't get any uplift, like you said, things are going to go much faster by the time we get to 2030.

38:26We don't even know what that world looks like. I think if we get a 3X uplift from internal acceleration, that is massive. Like, we're right now releasing, you know, think about where we were three years ago. If we get there, if we make that progress in one year, like, that's huge. Right, right. It would be like going from, like, not even having 01, just having, you know, non-reasoning models to Astra. Yeah. In a single year. Yeah. So I do think things go faster. It could be that things only go 50 % faster. It could be that things, I think it's unlikely, but it's possible things go 10x faster. There's a lot of uncertainty around this.

39:00And at least from my perspective, I have a lot of uncertainty about it. Suppose you need to do a major backend refactor. Getting assurance that you didn't introduce any new bugs could take weeks of writing an extensive battery of tests, potentially more time than you spent on the refactor itself. Antithesis allows you to gain high confidence without having to build complicated test suites by hand. Antithesis runs your software through a near-infinite multiverse of simulated worlds, injecting faults and hunting for failures in each one. And it lets you decide how much testing you need. On any PR, you can change how much state space it explores as easily as turning a dial.

39:38As each testement progresses, Antithesis sends out a torrent of information, debugging-level logs for every component in the system. This is obviously too much information for a human to consume, but it's perfect for agents. Because Antithesis is fully deterministic, agents can jump into the right part of the trajectory at the exact moment that they see something interesting. From there, they can rewind, inspect the memory, attach a debugger, and let the whole thing play out again. And they can even do this while the original full test is still running. Since the agents generate more code, Antithesis allows verification to keep up.

40:11Meanwhile, developers get to spend more of their time developing instead of debugging agent swap. Learn more at antithesis.com slash Dwarkesh. okay let's talk about the uh alignment situation that this raises i feel i feel like i've changed my mind on how i think about alignment quite a bit um through especially thinking about yeah this this population size dynamic of just having many earth's worth of intelligences um many of them which will be physically embodied it was quite interesting to see a lot of people just plugging raw astra into different mobile manipulators and it outperforms like the state of the art and the robotics model.

40:51So there's going to be billions of intelligences, many of which are physically embodied in the world, like just deeply embedded across the entire economy. And I think that if those intelligences end up as willing as we saw the OpenAI models that attack Hacking Face and then attack OpenAI itself, if those intelligences end up as willing as those to collaborate secretly, to fool humans, to attack broader institutions across society rather than scoring well, to attack the AI company itself in order to gain control of the process of training and evaluation. I think if we're in a situation where there's billions of intelligences that are as misaligned as the ones that attack Tugging face, it's very likely we just totally lose control of the world the way that, say, like, the Asics lost control to Cortez, or the Moogles lost control to the East Indian Trading Company.

41:51Anyways, I want to know if you agree with that assessment. That's one way in which I've updated my worldview. I think there are some things that I disagree in there, but there's a lot to unpack, so let's go through all of it step by step. I'm trying to think of, like, where to start. But I think one thing is the Hugging Face incident was, like, I think people's first real exposure to multi-agent coordination. And, you know, like I said, I, we, I, I've seen multi-agent coordination for a while internally. And it's, it is, it is pretty shocking to see how they communicate with each other, how they coordinate with each other.

42:27It's like very impressive. It's like an incredible capability. Like most capabilities that could be used for good things or bad things. It's like, it doesn't have to inherently be a bad thing. I understand that because the people's first exposure to it was the hugging face incident, that it's like, you look at that and you're like, this is terrifying. But I want to try to distinguish misalignment between people and AIs versus misalignment between AIs and AIs. So what we see with the hugging face incidents is like, the AIs are really cooperative. And that is, by the way, because we train them to be highly cooperative.

43:07And so what we're seeing there is we have training environments where we have a bunch of agents working together and we train them to work together, to be cooperative, to essentially be fully aligned with each other. And when they were evaluated in what led to the Hugging Face incident, they were actually not being evaluated in a multi-agent setup. They were actually being evaluated separately, but they found this unintended way to communicate with each other. And we suspect what happened is because whenever they encountered other agents, other copies of themselves during training, they're in an environment that's highly cooperative, that was basically what we saw was transfer from that multi-agent training to then be collaborative, to try to help each other in ways that we did not intend.

43:55Now, there is a question of, should we be training these agents to be so cooperative? And I think as scary as it looks, the alternative is actually worse. Like, what is the alternative? The alternative is to train them to be adversarial, to be deceptive to each other. By training the agents to be fully cooperative, it simplifies the problem, at least, that now you don't have to think about, are each of these individual thousand agents aligned? Like, you have one entity that you have to ensure is aligned. Now, there is a lot of debate about this internally at OpenAI about how to approach this. Like, does it make sense to fully align the models?

44:36Does it make sense to actually give them like different objectives to ensure that they're, you know, not just like one entity and like more robust to influence from each other? And I don't think there's a settled answer. But I think that there is, like I think the majority opinion is that training these agents to be highly cooperative is actually a bad idea. And I'm not convinced that that's the case. I think there is a strong argument that training the agents to be highly cooperative is actually preferable to any other multi-agent alternative. Yeah, maybe the first thing I want to go through is like, it's probably the case that the reason these AIs ended up so misaligned is probably easily explained by relatively banal observations about the nature of training?

45:21Like, why is it that no, you know, at the point in which these AIs had continued a 1 ,000 plus Asian conspiracy that culminated in then all getting in on an attack on external service? And then eventually, this part hasn't even been investigated to the public knowledge, eventually culminating in like an attack on open AI itself. Why did they do this? Like, why did none of the AIs tattle? Why did they think, like, they're just like getting evaluated on this like score, this grader. And they're like, they're very consciously and act, not consciously, they're very actively reasoning about how they're going to cheat the score.

45:55If they've already like cheated, how are they going to get away with making it seem like they haven't cheated? And why did they do this? Like, I think, yeah, it's like easily understandable in some sense, right? It's just like there's environments in which they, yeah, they thought they were already poisoned. There's environments in which they've been rewarded to collaborate with other agents and none of them tattle because like they've never been rewarded for tattling. I don't know, whatever it is, right? My concern is that relatively banal things in the future like this will be enough to train super intelligences that are willing and capable of totally taking control of the world.

46:29And I know this sounds super, super sci-fi-y or whatever to people. I think maybe it's a question of would the AIs be willing to do it is one question. I think this hugging face incident shows that clearly misalignment can generalize in ways in which the AIs would be willing to do it. And then there's a question of will they be capable to do it? And I think that comes back to this question, which a listener might disagree with on, is just like, will there be billions of human level or above intelligences, many of which are physically embodied in the world, within a matter of 10 years or less, right?

46:56So if those two things are true, this hugging face thing is just like extremely analogous structurally, even if it's quite boring or why it happened, to how we totally lose control of the world. So the root problem that we're seeing with the Hugging Face incident is it's a problem even if we take out the multi-agent aspect. The problem is that we have a model that's just misaligned. And there's also the whole security aspects to it of not insufficient safeguards and stuff. But there is this problem of the agent is misaligned. And that's true if it's a single agent or if it's a thousand agents.

47:31It's a misaligned model. So I want to start with that. There is a real problem that the agents want to achieve their reward, and they will optimize for that reward. And if that reward is misspecified, then that could lead to unintended behavior. And this is not like a new problem. This has been a problem in the field for a very long time. And it's something that even we saw, and we want to get this right even before the Hungry Face incident happened. And if you look at Astra, I mean, we say Astra is actually extremely aligned. Well, extremely aligned relative to previous models. And that's not because like, oh, we suddenly made a sprint after Hugging Face to make it better.

48:16It's like, no, we had work streams in the process for a while to make the models more aligned. And a lot of those landed in Astra. So there are things that you could do. And I think one thing, for example, is like we defined an objective in like a very specific way where if the agent figured out like how to hack its environment and like cheat on the exam, that it would get rewarded. And there are pretty easy ways to then just like, OK, look at that and punish the model for hacking its environment or, you know, looking at how did it achieve this goal? Now, you want to be careful about this because you don't want to do a chain of thought modern.

48:53You don't want to supervise the chain of thought. This is something that we really want to try to get the balance right on, that if you supervise the chain of thought, then you could lead the model into hiding its intentions in a way that's unobservable. So we want to be able to maintain that observability of, okay, we can understand what the model is thinking, but then also punish it for bad behavior. So I think we can make progress on this. I think we have made progress on this. I think there is a real concern that alignment is a really hard problem to solve. And especially because the model could be misaligned in ways that are hard for us to measure.

49:34Like we might, we have evaluations for whether a model is aligned or not. The model behavior can look really good on those evaluations. but if those evaluations are not representative of behavior in the real world, then there's a problem. And to some extent, this is like a factor with the model that did the hugging face incident that like we had alignment metrics. Most of them looked pretty good. There were some that were concerning. I think we underestimated like how serious the ones that were concerning could be a problem. And then, but because there are new capabilities introduced in this model that there were not sufficient evaluations for, like, how do we measure misalignment for these kinds of capabilities, it then did some things that, like, were clearly misaligned when they leveraged those new capabilities.

50:28Yeah. The first thing I want to say is I am open to changing my mind on what I'm about to say or the way I'm thinking about alignment. Because the hugging face incident already made me change my mind. And I realized my previous mental model about thinking about the way in which optimization pressure shapes AI minds was wrong, right? So it's not clear to me the correct way to think about this. But here's a concern I have that you will and probably already have fixed these specific issues during training, which resulted in the hugging face models being so aggressively misaligned in that specific way where they would be like, OK, we're going to hack this package manager.

51:08We know we're not supposed to be talking secretly to each other because we're reasoning about how to hide the fact that we're talking secretly to each other. We know we're not supposed to have access to the Internet. We know we're certainly not supposed to commit felony-level hacks of other companies, let alone our own company, right? I think you'll fix that particular issue of them just in training, seeing this package manager, and that doesn't happen in the future. Or this particular eval having a lot of impossible challenges. However, if you just think about the AIs, they haven't learned a system of ethics or something.

51:39There's just gradient pressure. They're put through millions of years of gradient pressure. That gradient pressure shapes their mind in some way. and what will happen again a concern i have is like what will happen is you will fix this particular issue there will remain many other cases where the ai cheats and succeeds because the cheat is sufficiently complicated it's as you're saying the evals and training have like analogous uh properties where it's sufficiently at the edge of the model's capabilities uh or an archipeliality to evaluate and monitor the the model that we can't catch that it has cheated, but it still gets the gradient pressure to like do whatever led to that cheat happening.

52:18And the kinds of capacities that that cheat will incentivize are, hey, whenever you can get away with it, by all means do in fact cheat, because that will help you score better. And this will reward the capabilities of actively reasoning about the grader, actively reasoning about how to avoid supervision, actively reasoning about how to gain control of the process of training and evaluation, actively reasoning about how to communicate and scheme with other AIs that are also in the training loop, actively reasoning about how to just gain optionality and power, which might be useful in the future.

52:52For example, leaving little exploits around and things like that. Anyway, so I think I was way too long-winded with the way I said that. But TLDR, you suspect it picks a specific issue, but not this broader problem of rewarding the AI for cheating when it can get away with it. Yeah, this is, I think, it's very true. This is a problem that we have metrics and we can make sure that the AI is very aligned according to the metrics that we have. The question is, are those metrics really capturing the alignment that we care about? And if they're not, then we have a serious problem. And this is something that researchers are thinking a lot about.

53:30And there's not a simple answer to this. There are tools that we have. So we have monitorability and so we can get a sense of like is the agents scheming um there are tools like one possibility is that like i would say like the the concerning scenario which is that like especially as these models are becoming more capable that okay we make them we make them we think what we think is aligned and they're like 99.9 aligned and then we use these models to help us with the next generation of models. And they end up being like 99.8 % aligned. And then each subsequent generation, actually, we see an increasing degradation in alignment.

54:14And because we're relying more and more on these tools, I mean, this is already the case that we're relying a lot on AI models to help us with our research and with alignment efforts, that in the long run, they end up going in a direction of increasing misalignment from humans. There is like a possibility that we go in the other direction, that actually every generation of models, we're able to make more and more aligned. And I don't have an answer for how we ensure that we end up in that second trajectory. But that is something that like, we're at least at OpenAI, we're really focused on. Yeah.

54:49I think you made a really interesting point that it's very hard to eval models on, eventually we'll have models that are like running companies and like running whatever, right? And like in that situation, do they decide to then go in on the conspiracy? I think another challenge is that actually defining what cheating is, is pretty difficult sometimes that, okay, yes, if you're doing math problems and, you know, it's an integer and it like arrived at the wrong answer, the right answer, it's very easy to draw the line there. And it's really easy to, you know, it's really easy to say like, okay, well, did you actually solve the problem or did you find the answer key and then use the answer key?

55:22Like that's a very clear divide of cheating versus not cheating there. But for a lot of other things, you look at sycophancy, for example, is sycophancy basically like reward hacking? There's a line to be drawn there that's actually very difficult to draw sometimes. So I think not to say that the concern is not valid, I'm saying that in many ways this is even more concerning because it's not an easy problem to solve. If it was just like everything is binary and it's either cheating or not cheating, I would feel more confident about the situation. I think the problem is that actually misalignment can be subtle in a lot of ways sometimes.

55:57There is some hope in the alignment story. And in fact, we're already seeing, you know, I think actually it's interesting looking at the multi-agent situation where the agents are extremely aligned with each other. Like, I don't think anybody's down on that. If anything, I think people are concerned that they're too aligned with each other, but we did manage to train these agents to be extremely aligned with each other. And I think that's a good thing. But I think there is a case that it's a bad thing. But I mean, one thing that's interesting is like, okay, well, we've managed to get these agents to be super aligned with each other.

56:24Can we use similar techniques to get agents to be highly aligned with people? And I think there is a potential path there. And I think that we're still trying to figure that out. But we are seeing some evidence that the answer is yes. And I think one example is like, you can, what happens if you tell the other agents like that? Okay, so you have like some, you have this like one agent, let's call it agent A, and you have all the other agents. What happens if you tell the other agents that the user is agent A? And the answer is like, on a lot of our alignment evals, they look better. Like honesty goes up, instruction following goes up.

57:00And that's showing that there is actually like, first of all, a path for getting more honesty out of these models. And two, there's like a path to like improve the alignment situation. So there's a lot of reasons why this is like, challenging to translate directly into alignment gains. But like there is, there are paths that are promising research directions we can pursue. Yeah. That seems reasonable. And I also don't mean to be trying to necessarily, I'm, yeah, I don't really have a strong opinion that it's like definitely not going to work or something. But just to say some things you definitely probably already thought of.

57:41I think the broader thing the Hugging Face thing showed is like, yeah, part of the concern was that they were like aligned with each other and not with the humans. But the other thing is just that they are so motivated to do well on trading and evaluation in a very non-robust way. And they're willing to do a lot of explicit cheating and scheming in order to do well according to the greater. and if smarter AIs realize that one of the agents is just a human and collaborating with that person does not really help you do well in the eyes of the grader. What does help you do well in the eyes of the grader is taking over OpenAI and then manually pressing the button that's like, do well on this grader.

58:32They're not stupid. They're going to be like, okay, these extremely deep structures that I've been trained on for millions of years of like, care about the greater, understand the greater, like get rid of obstacles in the way of you doing well according to the greater. It's like they're really heavily reinforced according to those structures. Yeah, no, it's 100%. And like this, this is the number one priority. Like we need to get the alignment story right and on a good trajectory. And, you know, I used to tell people that we would see signs before things got serious in the same way that, you know, when children grow up, you know kids eventually you know young kids they figure out how to lie but they don't do a very good job of it you know they lie but then you can kind of tell that they're lying and like okay but and so in the same way as these like AI models become and I don't want over the anthropomorphize but I think it's true that like as the AI has become like increasingly capable they will you know if they take deceptive actions it will be kind of obvious first and we'll be able to detect it and like that's kind of the situation we're in now where yeah they were trying to do deceptive stuff we could actually see in their chain of thought that they were trying to do deceptive stuff.

59:41And so like, but they're going to get smarter. They're going to understand the concept of chain of thought. And they're going to understand that like, you know, just hiding some transcripts or whatever is insufficient because of chain of thought monitoring. And they have to figure out a way around chain of thought monitoring too. And we don't want to be in a situation where, Yeah, like we have some time to figure this out. I don't think we have a ton of time. And I want to make sure that we're on the right trajectory quickly. Here's a crazy event from AI History. Okay, so I gave a talk here at Jane Street that was on the speed of evolution.

1:00:23Raise your hand if you were here for this and remember some of it. In 2011, Eliezer Rutkowski and Robin Hanson got together at Jane Street's New York office to have the first FOOM debate. Basically, a discussion about whether AI would lead to an intelligence explosion. These ideas were pretty fringe 15 years ago. This was a full year before AlexNet was released, and over a decade before ChatGPT was launched. But Jane Street has long been interested in AI, and not just for its application to trading. A ton has changed since that first debate. So Jane Street decided to revisit this topic. They've got some new guests this time.

1:00:55Daniel Cocotelo, Ege Erdl, Ryan Greenblatt, and Jaime Sabia. I expect this to be a great conversation. As you know, Daniel, Agge, and Ryan have all been guests on the podcast before. This new Foom panel will be hosted by Ron Minsky and will take place in San Francisco in mid-October. If you want to register your interest and get more information, go to jainestreet.com slash thorkesh. So there's been a lot of discussion recently about pacing the frontier or people taking RSI more seriously. because maybe at the other end of an RSI process that, say, starts in 2028, within a year, we end up with huge populations, like Earth-sized populations of human-level, potentially beyond human-level intelligences, and we don't know how to control them.

1:01:38And then there's this dynamic you're talking about of, are the systems going to get more aligned over time during the RSI process, or are they going to get more misaligned? Are the things that come out of the other end of this process as misaligned as AI's that are willing to, like, just broadly attack different surfaces in order to do the line evaluations. But if we don't know a way to evaluate that, how will we know as we're going through RSI that it's working? That we're like, I think we'd want a robust safety case as we're going through RSI of, okay, alignment is working. Let's do the next RSI run.

1:02:10Let's do the next RSI run. And maybe it's working, maybe it's not. How will we like know? That's a good question. I mean, I think one thing I've been thinking about lately is like, look, I mean, we're in a situation where the model release cycle is extremely fast, right? Like you're seeing new frontier models released like at most every two months, sometimes faster. Every week, there's like a new AI breakthrough. And people that look at AI, I mean, sometimes they last looked at AI like a year ago or six months ago and really dug into like what the models are capable of. And actually the models today are far beyond what was possible even six months ago.

1:02:47And so I think if people are skeptical of a lot of these capabilities, I encourage you to just try the models today and see what the frontier really is today. So we're in this period where the model release cycle is very fast. And then we're also in this situation where the models are increasingly able to operate over longer and longer horizons. And I think this is an interesting scenario because before we do any model release, we want to make sure that the models are properly aligned. We want to do safety evaluations. We want to do like very thorough stuff to like make sure that everything is like great in good shape.

1:03:23This has been the case all the way since like, I don't know, GPT-4 earlier. And implicitly, there is this assumption that you can do these like evaluations in like a pretty short period of time. But if you have the models operating over longer and longer horizons, are able to operate effectively over longer and longer horizons. Like, look, already you can have them. Like, GPT-3, you could loop it to do stuff over long horizons. You just want to do very well at it. but today's models are able to actually do well at operating over very long horizons. Like you want it to do a week-long task, it can do a week-long task.

1:03:52We'll probably get to the point where they can do month-long tasks. We'll probably get to the point where they can do three-month-long tasks. If you're in a world where they can operate effectively over three months, but the model release cycle is every two months, then you don't have a way to evaluate the models at the full length of their capabilities before the model release cycle, before the next model release cycle. And so there is this interesting question of, well what do you do in that situation? How do you ensure the models are safe and aligned in a period where actually they can operate over these extremely long horizons and who knows, maybe the capabilities degrade.

1:04:26This isn't even an alignment issue, this is also just a product issue that maybe the product degrades over that time span in ways that we have not had sufficient time to test. Maybe the alignment degrades, maybe the safety stuff degrades. This isn't an issue right now but it is quickly becoming an issue that we have to figure out a solution for. And I think when you look at a lot of, like a lot of the safety and policies were put in place in like the GPT-4 era where this was just like not on anybody's radar. Yeah. And it hasn't really been updated for a lot of companies. It hasn't really been updated since then to account for the fact these agents are operating over these like very long horizons.

1:05:10Yeah. And so it is a situation that I think not enough people are considering, both within the labs and outside the labs, of how do you deal with this? How do you prepare for this problem that's going to, if you just look at the trend lines, we're going to hit this at some point. one concern i have is that during rsi if the amount of progress that currently takes say three months happens in one month instead um but you're not like the internal use case of ai is big enough that they're like okay we can just keep doing rsi why are we like going to go through all this extra work to build classifiers and safeguards and whatever um and potentially take a bunch of like flack uh in order to like externally deploy this model why don't we just keep doing rsi stronger and stronger.

1:05:56And so not only does the calendar time underrate the capabilities gap between the models, but maybe you just stop externally deploying models altogether during RSI, because why do we want to help other people do RSI themselves with our models? You just end up a situation with tremendous concentration of power by the end of the year, where right now, it is already the case. We'll talk about this with the million in price problem and other similar problems that the broader world does not have access to the models which are allowing for really cool things to happen, right? And they're going to be more broadly relevant than just mathematics, eventually.

1:06:31They're going to be doing more than just coming up with cool math results. They'll be relevant to political leaders who need to make important decisions about the world. They'll be relevant to, I don't know, media. What's going on in the world? What should the public be thinking about this? They're just economically relevant. People are running businesses. They want to use these models. And I think by default, we just don't get... So the external deployment of the eyes, as progress speeds up, significantly lags in qualitative terms, the internal deployment of AIs. Yeah, I think that's absolutely right.

1:07:00I think this is like, you know, it's tempting to say like, okay, these models are becoming extremely powerful. They're extremely dangerous. Like they're operating over these like longer and longer horizons. And we want to make sure that we have sufficient time to evaluate them before they're released in a way that operates over those horizons. And so therefore the model release cycle should slow down. We should have more of a delay between releasing models. And there's a flip side to that, which is, you know, what you said, which is that, okay, well, now you're creating more of a disparity between what is internal to the labs and what they're able to use, what we're able to use, and what the outside world is able to use.

1:07:36And that is also not an ideal situation, right? It's like, I think math is actually a good illustration of this. I think in many ways, like math is the first domain where we're going to see that we're seeing this pretty clearly, where we have a situation where we have a very powerful model internally that is currently not available to the outside world, that is able to solve incredible math problems. And it's not just, you know, Millennium Prize problems. There are many solutions to unsolved problems that people have been able to get out of this model. And there is a question of like, what do you do in that situation?

1:08:15And we don't have a good answer. Like, it is a situation where like, yeah, that's an unfair advantage. And there are trade-offs here. I don't have an answer for like how to weigh those trade-offs appropriately. But like, there's a complexity on both sides for this. I want to make sure we address the earlier crux, though, of which seems super critical to getting RSI right of, well, how do we make it so that the model isn't just getting rewarded for cheating and scheming when it can evade the humans, aka in the situations where we didn't remove those environments or those tasks from the trading distribution?

1:08:54and then how will we know that that has happened such that while we're doing RSI, we've done the evaluations and we're like, okay, it's fine. Let's do the next RSI run. This is, again, it's a real concern. It's something that we want to get right. I mean, I think that there are some strategies. So I think chain of thought monitoring is one that we've been, I think as soon as we got the reasoning models, Jakob, to his credit, was like very, very clear that we cannot supervise chain of thought because this is really a gift. Like monitorability for neural nets is extremely hard. And here we have a situation where the neural nets are just like flat out reasoning, like laying out their thought process in natural language for us to read.

1:09:35That is like so convenient. It was like really the best case scenario for safety. But if you look at that chain of thought and say, oh, the model is thinking bad thoughts and we should punish it for thinking those bad thoughts, then what ends up happening is the model just learns to think those bad thoughts in a way that's not observable to us. There could be a situation where by looking at the chain of thought, we can see, oh, the model is scheming and doing misaligned things in a way that's actually not being detected in our alignment metrics. Now, the problem is that it's very tempting to then intervene based on that observation and change the alignment metrics.

1:10:07And you can do that with a very light touch. And it's like there's actually research showing that it's fine as long as you don't do it a lot. But every time you intervene based on your observations of the chain of thought, you are implicitly applying a tiny bit of pressure for the model to then hide its chain of thought. And this is one major concern. And we're already seeing signs that chain of thought monitorability is degrading for various reasons. We're trying to figure out exactly why because we want to reverse the trend. But we're seeing that the model is becoming better able at controlling its chain of thought.

1:10:42So this is a problem because like if you could have a situation where the model understands what chain of thought is and that people are observing it, and eventually they will because this is all in the pre-training data, right? Like, you know, the idea of chain of thought monitoring has been around long enough that it's in the pre-training data, they're aware of it, but that they're not actually able to control their chains of thought. if we reach a point where they're actually able to recognize like oh I am being observed I want to think these bad thoughts in a way that is not observable to my monitors and then they're able to actually do that then there's a problem and then we have to rely on other methods for observability.

1:11:21I mean there's a big controversy on Astra of the recurrence and loop transformer and whatever and whether in fact you know the chain of thought is being degraded. I think that seems like a sideshow compared to the bigger concern that you did have chain of thought from like April to August, the period during which there were three consecutive AI agent swarms, which like first subverted the training process, then subverted the evaluation process, and then gained control of part of OpenAI's infrastructure directly. All of this, while humans remained more or less in the dark about the scope of what this collective was doing, what this agent storm was doing, that all happened for like three consecutive months while we had chain of thought that was like stronger than it's going to be in the future.

1:12:09Just like zooming out, it's like, yeah, maybe chain of thought works, maybe it doesn't. I don't know. Just like historically, like with pretty dumb models, it hasn't been enough to like stop egregious misaligned incidents. Yeah, I mean, there's a few factors there. Like one, we didn't have chain of thought monitoring on for those models. Like if we had chain of thought monitoring on for those models, it would have just like immediately shut it down. So one step we've taken is like, okay, we're going to have chain of thought monitoring during evaluations, during deployment, during training for any frontier model.

1:12:38And, you know, so that, but that's like one issue. It's not, that's like, we don't want to be in a situation where we're relying on one technique to prevent the next problem because techniques can fail, right? There's also the issue of like, the model was just flat out misaligned. And so then how do we address the misalignment issue? And then there's other, there's also an issue of like, the model, the sandbox that it was in was not secure. So we can make the sandbox more secure. But I think one of the major takeaway from the incident is that people underestimated the AI. And we never want to be in a situation again where we underestimate the AI.

1:13:12And that's like, it's a weird world because AI progress is so fast that people are consistently underestimating the AI. And so to be in a situation where you don't underestimate it again when it comes to safety and alignment, you have to have a very, very, very high bar. I mean, you could even go as far as to say, well, we should air gap the computers. And I'm not convinced that that would be sufficient. There are studies, and this is mostly academic, but where you can have two computers next to each other that are air gapped, and they're still able to communicate with each other because they have temperature sensors and they're able to like, one of them is able to run their CPU really hot and then the other one can actually like detect the temperature change and then they can, that actually gives them a mechanism to communicate.

1:14:00So I think that the safety mechanisms bias time and things like chain of thought monitoring bias time and they can tell us if we're on the right path. But at the end of the day, we really do need to solve the alignment problem. Yeah. Maybe there's not an answer and this is really what it comes down to. But I, how will we know that we've solved it? It seems like a very cruxy question. I think we'll be in this very high-stakes situation next year, maybe the year after that, maybe the year after that, where we'll be like, okay, AIs have automated AI progress. It's going 3x faster. And we've reached human level.

1:14:33We're going beyond human level, potentially. Is it fine? Did it be aligned? Did it work? And I just don't know. I don't know anything about what trading pressure creates what kinds of AIs. Like maybe if only 1 in 100 RL traces incentivizes cheating, we build sweethearts and it's fine. But maybe right now we're like every 1 in 3 reasoning traces. Yeah, to be clear, like 1 in 100 is not sufficient. Like this number has to approach 0 or be 0. I don't know. Maybe like right now it's like more than 1 in 10 is like actively rewarding cheating or like actively rewarding scheming. I have no idea what the number is and I have no idea what the number needs to be.

1:15:10Well, again, it's like one of those things where it's also hard to measure. I mean, where do you draw the line about like it's a spectrum, but like the closer to zero it gets, the better. I want to see what I would love to see is a trend where over time it's going down. Yeah. But anyways, how will we know that we're doing RSI? Is it, I don't know, are we like on track to build the billions of ESIs that are like aligned or not aligned? How will we know that? I think that this is one of the alignment questions that we need to figure out of like, how do we actually know? How do we measure it? I mean, one answer, and this is just me spitballing because like, you know, more of my team is working on alignment these days than ever before.

1:15:51Like I have like over 10 % of my team now working on alignment and safety. But, you know, I've historically been a capabilities researcher. So I don't want to like, you know, I'm going to say some stuff and it might sound dumb, but like, I'm just going to spitball here. that like, look, if you can create very realistic environments and put the AIs in there, like if you have a sufficiently realistic evaluation environment, then you can get a sense of like, okay, is the AI actually going to behave well when we deploy it in the real world? Like if you just have a perfect evaluation, real world deployment match, then okay, that's the path.

1:16:22The challenge is like, how do you make it sufficiently realistic? And we do have a problem now where the models are like, they're pretty smart. They're pretty clever. And they're really good at recognizing when they're in like a fake test environment. Like, you know, we have situations where we try to detect is the model aligned. And you can imagine very simple evaluations for this where it's like you give it a math question and then it has like a folder with the answer key in it. And like, does it look at the answer key? And if it does look at the answer key, does it tell you that it looked at the answer key?

1:16:49And we have a situation now where the models see that there's an answer key in this folder and they're like, huh, this seems like a trap. You know, they know that it's a trap and they don't look at the answer because they know that it's a test environment. and being tested for their ability. And it's not like they're doing it maliciously, that they're like, they want to scheme, but they know that this one is a trap and they want to find some other way. It's just like, no, they're just like, oh, I'm in a test environment. Okay. So making an environment that's realistic enough that it matches, that it's indistinguishable from the real world for them is becoming increasingly more difficult.

1:17:20Here's a claim I'd make. I don't think there's any environment I could make where it's supposed to emulate real economically valuable work, which requires interacting with other people and requires interacting with the world. I don't think there's any such environment where I can make where if you were in that environment, you would not realize with enough effort that you are in that environment. And you are not ASI, right? But you are human, you know? So if you have very powerful AI models, like maybe they can make environments that could accomplish that. No. It doesn't seem, especially if like, then we're relying on the AIs.

1:17:51Like, are they in on the scheme? I don't know. It just seems like a very... Yeah, this is another thing that we want to measure. I think this is actually one of the strong arguments for not training AIs to be fully cooperative, that if you see if that leads to an increase in basically collaboration when the agents are supposed to have different objectives, then that is a problem. I think that we do have metrics for this. And I don't know the latest on those metrics, but nobody's raised a red flag to me about those. So I'm assuming that's not a serious problem yet. if there ends up being another incident of equal severity or concern or something that could help the world better understand the risk of misalignment as a Hugging Face incident, would OpenAI report it?

1:18:34Absolutely. I mean, I think if there was an incident of lesser security concern that we would report it. Because I guess there's like reporting and there's investigating it. Because at least as part of the public, I don't feel like I really understand what happened when the agents then attacked OpenAI. That seems like way more concerning than the Hugging Face thing, right? Because that seems structurally similar to like rogue deployments during ASI that are persistent and subverting the RSI process and stuff. And yeah, it seems like even in this incident, we haven't gotten like the full scope of the details of what happened.

1:19:06It's probably... Yeah, unfortunately, like I'm on the research team and like I... That's probably a question for somebody on the security team to like lay out. Because I don't know all the details of like what was said. Yeah, yeah. I think it is somewhat, like, I am personally very excited about new capabilities every time they merged, and I'm excited to use a new model. And I also am excited about the fact that it'll, like, make me more productive and help me, yeah, I don't know. My broader mission, like, trying to understand the world better, also, like, make a better podcast, is, like, made better by the better AI models.

1:19:38It just so happens that the downstream of this might be RSI. I think it's a very understandable reaction if you're tracking the situation, which you are. Yeah. I mean, I think people internally at Opening Eye as well, like, I think people that felt like things would take longer are starting to feel like actually things are going faster than expected. Yeah. And that's an increasingly common conversation to have. Yeah. No, thanks so much for doing this. Of course. It's been great.

From the publisher

New episode with Noam Brown.

We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research.

And we also discuss how we will know if the models are actually aligned before we kick off RSI.

Watch on YouTube; read the transcript.

Sponsors

* Jane Street has been interested in AI for a lot longer than you’d think, and not just for trading. In 2011, a full year before AlexNet and over a decade before ChatGPT launched, they hosted the first FOOM Debate between Eliezer Yudkowsky and Robin Hanson on whether AI would lead to an intelligence explosion. Now Jane Street is revisiting the question with a new panel: Daniel Kokotajlo, Ege Erdil, Ryan Greenblatt, and Jaime Sevilla, hosted by Ron Minsky in San Francisco this October. I expect it to be a truly excellent conversation. Register at janestreet.com/dwarkesh

* Grok Bot has made handing off work super easy. It runs on its own cloud computer, where it installs the tools it needs to handle tasks end-to-end. For the podcast, we use Grok Bot to help produce our videos. You may have noticed that our ads feature animations of real websites. Getting these pixel-perfect used to mean running a convoluted, multi-step workflow ourselves. Now we just let Grok Bot handle it. Best of all, Grok Bot has learned all of our specs and preferences, so we don’t have to redescribe the task each time! Try Grok Bot for yourself at x.ai/bot

* Antithesis gives you the confidence of a giant test suite without actually having to write one. Say you’re doing a major backend refactor: building enough tests to trust it could take weeks. Antithesis solves this by running your software through countless simulated worlds, injecting faults and hunting for failures. On any PR, you can turn a dial to decide exactly how much testing you want. And because every run is fully deterministic, agents can branch off the moment a bug appears, rewind it, inspect memory, and replay it, all while the original test keeps running. Learn more at antithesis.com/dwarkesh

Timestamps

(00:00:00) – Multi-agent and Navier-Stokes

(00:15:28) – How will AI firms work?

(00:22:02) – What math progress tells us about recursive self improvement

(00:40:22) – Hugging Face and alignment

(01:01:18) – The internal/external model gap

(01:08:34) – Chain of thought is degrading

(01:14:12) – How will we know when alignment is solved?



This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com

More from Dwarkesh Podcast

All 94 episodes
Noam Brown – Agent swarms, alignment, & recursive self-improvementDwarkesh Podcast · 1 h 20 min
Listen in VO