AI researchers debate how close we are to recursive self-improvement

11 Sep 2026 · 1 h 37 min · 36 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode is a debate among three AI researchers about whether we’re close to recursive self-improvement (RSI)—AI systems that can iteratively improve their own research and training loops—versus a future where progress continues but doesn’t trigger a sudden “takeoff.” They discuss why, even by ~2036, we might not see “billions of superintelligences” transforming the world, focusing on technical bottlenecks: persistent limits in generalization (including sim-to-real), continual learning, and the difficulty of discovering new training paradigms rather than just scaling current ones. They also argue that capability growth may look like repeated cycles: models seem “AGI” at release, then feel “dumb” after use, due to bottlenecks in judgment and verification.

Key examples and claims include

chess ELO crossing human expert ranges as an analogy for linear-to-discontinuous capability shifts; the “Kaplan scaling laws” annealing mistake as an example of how AI “thinking” could catch errors earlier; and the idea that distillation and prompt-distribution (including router/proxy data) can reduce centralization advantages.

Guests

Beren Millidge (CTO, Zyphra; open-source models), John Schulman (chief scientist, Thinking Machines; previously OpenAI; RLHF/ChetGPT work), Charlie O’Neill (head of model training, BaseTent).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring the Future of AI Intelligence

0:45 to 4:25

Discussion on potential reasons for the lack of superintelligent AI by 2036.

“But what is the most likely technical reason that we don't, like 2036 isn't like a crazy alien super intelligence world?”

Cycles of AI Progress

4:25 to 8:05

Analysis of the cyclical nature of AI advancements and their implications.

“Sorry, but do you think discontinuity will be harder than anything that's come since 2012?”

Discontinuities and Breakthroughs in AI

8:05 to 11:00

Examination of discontinuities necessary for significant AI advancements.

“that nano GPT speed run is the thing to be optimizing for in 2014.”

Generalization Challenges in AI

11:00 to 13:20

Discussion on the challenges AI faces in generalizing from training to real-world applications.

“Then you do a century of thinking after the experiment is over where you're analyzing what happened and what the next experiment to run is.”

Conclusion and Future Considerations

13:20 to 14:03

Final thoughts on the challenges of achieving rapid recursive self-improvement in AI.

The Challenges of AI Recursive Self-Improvement

14:03 to 17:32

Discussion on the complexities of AI achieving recursive self-improvement and the roles humans play.

“And obviously evolution needs to create creatures that can survive on them by themselves for a long period of time.”

Model Providers and Centralization

18:38 to 23:28

Exploration of the dynamics affecting centralization in AI model providers and the role of distillation.

“What is the story for why there isn't huge consolidation in model providers?”

Training the Future AI for R&D Automation

23:28 to 28:00

Speculation on how future AIs might be trained to automate research and development tasks.

“are almost objectively worse models than GLM 5.3, Kimi K3, even though they've had access to not only distillation but logic distillation from Mythos.”

Training AI for R&D Automation

28:00 to 28:39

Explore how AI models will be trained to automate research and development.

“it means it's quite easy to like keep up really.”

AI Learning from Feedback and Environments

28:40 to 30:49

Discuss the role of feedback and simulated environments in training AI.

“by the point at which you have AI that are actually capable of automating AR &D, how are they probably trained?”
Show all 36 chapters

The Role of Goals in AI Research

30:50 to 33:04

Understand how AI research is guided by measurable goals and intuition.

“And that progress is being contributed to by AIZ, of course, but it also still has humans in the loop.”

Scaling AI Models for Diverse Tasks

33:05 to 36:59

Examine the strategy for scaling AI models across various domains and tasks.

“Yeah, I mean, presumably the models will be trained on some combination of all of these tasks, and some will be very easily verifiable, some will be like LMS judge, or just ask the human, does this look reasonable?”

Domain-Specific Training vs. General Learning

37:00 to 38:48

Evaluate the necessity of domain-specific training for AI models.

“how much I construe why there is so much task-specific knowledge in these models if the path is like this kind of generalization.”

Simulated Environments in AI Training

38:49 to 41:15

Discuss the effectiveness and limitations of sim-to-real training frameworks.

“So basically, you look at what the real world tasks are like, and then you try to create a bunch of environments that can be simulated, like in the data center, and you can do RL on them.”

Learning from Deployment Data

41:23 to 42:00

Analyze the potential for models to learn from deployment experiences and data.

“And right now, that data is just not in a meaningful sense helping the model get better.”

The Evolution of AI Learning

42:00 to 44:13

Explore how AI models are improving through deployment data and continual learning.

“And once they do, you would have something that almost feels like a widely deployed intelligence explosion because the model is assimilating so much information across all these deployed instances.”

Challenges of AI Learning and Interaction

44:14 to 46:41

Discuss the complexities of AI learning from real-world interactions and the concept of recursive self-improvement.

“And like eventually that loop will become like faster and faster.”

Cumulative vs Non-Stationary Tasks

46:42 to 48:59

Analyze the difference between cumulative tasks and those requiring continual adaptation in AI.

“the sample efficiency of weight updates, they just seem way far behind humans, right?”

Taste and Sample Efficiency in AI

49:00 to 51:32

Examine the concepts of taste in decision-making and sample efficiency in AI learning.

“So I think being less sample efficient in certain regimes might be one of the sources of weakness, but then I think there are other sources of weaknesses that are completely different than that.”

Hive Mind Learning Dynamics

51:33 to 54:29

Investigate the potential for AI models to learn from collective deployment experiences and market pressures.

“And so theoretically, it's possible to develop it like that.”

The Future of AI Model Updates

54:30 to 56:00

Discuss the evolution of AI model updates and the balance between efficiency and capacity in learning.

“Then you generate traces, you put that in your model.”

Challenges in Reinforcement Learning

56:00 to 58:00

Explore the limitations of RL in transferring knowledge and techniques.

“And you have to pour in a lot of compute to create the right environments to get the knowledge in.”

Continual Learning vs. Model Retraining

58:00 to 1:00:30

Discussion on the balance between continual learning and the need for retraining models.

“You just have a model that's already gone through so much training and then you distill some fork that's been further RL'd or something.”

Data's Role in AI Progress

1:00:50 to 1:05:20

Investigate the significance of data in AI advancements and its limitations.

“architectures on, would result in a super intelligence?”

Scaling Challenges in AI Training

1:05:20 to 1:10:01

Examine the implications of scaling in AI training and efficiency gains.

“It's really a question of where the signal is coming from.”

Scaling Parameters and Efficiency

1:10:01 to 1:12:09

Discussion on parameter scaling, RL regimes, and model efficiency.

“So like the inference efficiency, like having some form of compressed attention in the deep-seq models is not necessarily geared around, you know, this is fundamentally like a period of improvement.”

Data Efficiency vs. Compute Efficiency

1:12:10 to 1:14:43

Exploring the balance between data efficiency and compute efficiency in model training.

“It depends on how quickly, like McCore and then in-house, these guys can scale up the complexity of the aural environments they're training on.”

The Role of Reinforcement Learning

1:14:44 to 1:17:03

Insights into how reinforcement learning contributes to AI model performance improvements.

“And so right now, people are still using a lot of H100s and stuff.”

Understanding AI Progress and Generalization

1:17:04 to 1:24:00

Analyzing the qualitative advancements in AI models and their learning capabilities.

“It's like if you're super bottlenecked on data and not on compute, you should go bigger.”

Model Iterations and Task Performance

1:24:00 to 1:25:18

Explore how model iterations affect task performance and reasoning.

“I mean, I think a lot of this as well is just like, I think RL does generalize a bit.”

Reinforcement Learning and Creativity

1:25:18 to 1:26:18

Discuss the role of reinforcement learning in fostering creativity in AI.

“the super creative move, that because it was never initialized on human data, it can think in ways that humans are not even thinking and come up with extremely creative solutions.”

Diversity in AI Outputs

1:26:18 to 1:27:21

Examine the impact of reinforcement learning on diversity in AI-generated content.

“oh, it's like totally destroying their like entropy.”

Distillation and Monoculture in AI

1:27:21 to 1:28:31

Analyze concerns about AI output monoculture from distillation processes.

“So I think that like that kind of diversity has definitely been like cut down by RL a lot.”

Predictions on the Future of AI Workers

1:28:31 to 1:31:29

Forecast the capabilities of AI as remote workers and their impact on various jobs.

“Super rapid fire predictions about the future.”

AI Research and Productivity Uplift

1:31:29 to 1:34:05

Discuss the expected productivity increases for AI researchers over time.

“Because I mean, like, yeah, because obviously like an AI researcher can be a remote worker.”

Human-Level AI and Onboarding Challenges

1:34:05 to 1:35:20

Explore the challenges in training AI to match human experts across fields.

“that like you can delegate some of this to the ai so like the ai is becoming decent at like deciding, you know, it's run this experiment, it's got this result, it runs like the next experiment.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Beren Millidge:Today, I'm chatting with three of my AI researcher friends from whom I learn a lot every time we talk and who also happen to be at somewhat open-ish labs and companies. So you guys can actually say things on the record. I'm joined by Beren Millidge, who is the CTO of Zyphra, which is developing open source models. John Schulman, who is the chief scientist at Thinking Machines, previously the co-founder of OpenAI, led the RLHF work that led to ChetGPT. And Charlie O 'Neill, who is head of model training at BaseTent. The first question I have, if we're in 2036, it's been 10 years, and we don't have like billions of crazy super intelligences that are running around that have like radically transformed the world, what is the most likely reason that that doesn't end up being the case?

0:43Beren Millidge:Other than sort of exogenous political shocks or like there's a war or they banned AI or something. But what is the most likely technical reason that we don't, like 2036 isn't like a crazy alien super intelligence world?

0:55Charlie O'Neill:I mean, like, my reason would just be like, it's got to be the sort of, like, there's been a classic thing, almost like Marvek's paradox, right? Where like, we see like, you know, we think of the AI being like, if it can do this, it's going to be amazing, right? Like, if it can solve these hard math problems, if it can win a chess, blah, blah. And then it solves these things. And then it's like, not that impactful. Obviously, it's somewhat impactful, but like, not everything. It's like, if somehow that continues, and like, there's never like the true like spark of generalization that occurs. I think that could lead to like the AI is just being like extremely good at kind of everything that people like put into a benchmark, put into an environment, but like there are still some persistent like Sim2Real, which is somehow blocking everything.

1:31Charlie O'Neill:I think this is kind of unlikely. I think we do actually see this kind of generalization even from our own practice already. But like if it is just like ridiculously hard to like generalize meta learning, plus like we don't solve continual learning, it's just like super hard and impossible. Like this would be my like default scenario in that case.

1:47John Schulman:Yeah, I agree with that. Humans have a lot of advantages over models now, and each time a new model comes out, it'll sort of, it'll catch up in some of these areas, but like you end up getting bottlenecked by the places where the model is weaker and where it has worse judgment or the models can't check themselves well enough. Yeah, so there's this cycle that keeps repeating where people think where a new model comes out and people are blown away and they're like, this is it, this is AGI, but then they use it a bit and then it starts to feel dumb after a month or so. So that cycle just might keep going and it's hard to predict how many times it's going to repeat.

2:30John Schulman:And like right now, you don't get explosive growth in capabilities because you still get bottlenecked enough when you're trying to do research and engineering that even if the model can write way more code than a person, it doesn't make you like 100 times more productive. But yeah, so maybe there are just more of these cycles than we would expect.

2:52Charlie O'Neill:For me, it's like a question of how far off like this global optimum of a learner you could have on a chip is like the transformer plus like RL, basically like the current recipe. So like I think people imagine that even once, like once you have an agent which is better than all humans at AI research, even if it's like 0.1 % better than all humans, then the fact that you can run like, you know, hundreds of thousands, if not millions of these in parallel, you can run them much faster, like chips going to speed up, that's going to outweigh every other like bottleneck and like you're eventually just going to like hit this like very fast takeoff with recursive self-improvement.

3:27Charlie O'Neill:I could imagine that if we continue along the trajectory that we're currently on with that paradigm where, you know, it's basically just like self-attention, RL, scaling up RL environments, And I guess if you think about what happened with Moore's law, we had this very nice straight line and that held for a really, really long time. But there were so many discrete discontinuities and innovations that had to happen to keep that scaling law going. And the same thing has kind of happened with LLMs. Like we had this pre-training scaling law and then that was kind of like hitting the diminishing returns.

3:58Charlie O'Neill:And then we came up with RL and solved that. And then we got this new diminishing returns curve to hit that made it keep looking like a straight line going up. And so if it requires another one of those discontinuities to solve, I'm not sure that the current method of training LLens with these RL environments, even RSI-targeted RL environments, would be able to discover that discontinuity. And if not, we're probably going to hit this asymptotic curve.

4:25Beren Millidge:Sorry, but do you think discontinuity will be harder than anything that's come since 2012?

4:30Charlie O'Neill:If we had the answer to that, we'd kind of have the ability to implement it. but maybe we should distinguish between a discontinuity which adds to the current paradigm. Again, it's cumulative. There's something beyond the RL that we have to discover and maybe they're capable of connecting the dots in that straight line. Or again, how far off the global optimum are we? Do we have to go back and throw out gradient descent and neural nets in general? And I don't think if you continue to scale up the current paradigm, an LLM, no matter how many LLMs you're running, are capable of necessarily discovering that if it's too far away.

5:05Beren Millidge:Yeah, the only hope really is if deep learning just can't get us to an AI which is at least, can dominate human research and human development, including the human ability to come up with new paradigms and so forth. Or like, I don't know, maybe humans would also never have discovered the next learning architecture, but to the extent humans could have discovered it eventually. But it just seems like, I don't know, If you just look at the progress that's happened in 2012 till now, and you just continue that on, I mean, I know it's been powered by huge amounts of compute scaling and so forth, but it would be weird if it just didn't get to the point where it could dominate humans, at least in R &D.

5:48Beren Millidge:especially over the next few years, there's going to be, Ryan Greenblatt was on the podcast recently, and he made this point that you could imagine as the AIs get more and more capable and are capable of making progress on simulations, which incentivize getting better at not only AI R &D, but generally at science. So this is a thing that all the labs are targeting, many startups are targeting. Or another intuition pump is if you look at the ELO score of chess bots since the 80s, There's just like a very linear increase in ELO over time. But there's this huge discontinuity as they cross the human range of human experts always win against AIs to like human experts never win against AIs as this linear increase in ELO happened.

6:28Beren Millidge:And you could think, I agree with your point that so far, AI capabilities have not been that big of a deal in terms of their end economic impact in the world. But that's because like they're slowly rising in ELO relative to humans.

6:39Charlie O'Neill:Yeah, I mean, I agree. I mean, the only way for this to not happen is if, as you said, somehow asymptote just before, basically, because we're already pretty close, in my opinion, to where we'll start crossing the human ELO score. And so we'll need to asymptote before that. And that's the only way in this scenario you pose where somehow we're sitting here in 2035 and everything is normal for this to happen, I think. I mean, the only other way is there's some dramatic regulation on AI. This is kind of what I see as the most likely way for this scenario to happen, actually, rather than the technical thing.

7:07Charlie O'Neill:yeah I think there's different kinds of research there's like research where it's like the auto research style where the objective is already specified very cleanly and you're optimizing that objective and I think everyone is picturing like if we continue along this path like you know making pre-training loss go down making our own environments go up that's going to lead to like improvement but like yeah maybe what Ryan is talking about is like this much more open-ended type of science which is required for like paradigm shifts where we can't specify the objective and the AIs are definitely not able to specify that objective either like where you have to be really, really careful about how we specify objectives for any of these things.

7:39Beren Millidge:And maybe your point is that the nature of the breakthroughs that have happened since 2012 is that we have found, in 2012 people weren't saying, or I'm assuming, I don't know, you guys were there, or at least, John, you were there, but I was not. I was in primary school. Actually, John, I'm curious where you're sort of wisdom of the ages, or wisdom of being in the trenches way back when, but presumably a big breakthrough was realizing that next token prediction is the, like you wouldn't have thought that nano GPT speed run is the thing to be optimizing for in 2014. But now that we have come to this new paradigm, you wouldn't think to do a speed run on that and have AI's get really good at that.

8:17But maybe there's like a next inner loop to optimize

8:22Beren Millidge:that the AI's wouldn't anticipate. And there's an outer loop of like revenue or something that eventually should be strong, but it's a very slow outer loop.

8:29John Schulman:Yeah. In fact, I remember in the early open AI days having the intuition that actually just minimizing log loss wasn't going to get you to intelligence because the important bits are accounting for such a small fraction of the loss that it was going to be overwhelmed by noise. So just training a language model on Nextoken Prediction just wasn't going to learn the interesting things you wanted to learn. And we needed to craft better objectives that would put more emphasis on the important things. And like you can make all sorts of arguments for this and you could say, oh, humans probably don't learn how to like we don't learn how to model everything in our environment.

9:17John Schulman:We can't most like people can't create a photorealistic reproduction of some kind of scene they've looked at. So there must be, we must need a better objective. But then it turned out that it just worked anyway.

9:33Beren Millidge:As you were pointing out, the inner loop, even in current AI research of like post-training benchmarks or whatever, doesn't necessarily translate into what users like.

9:41John Schulman:Oh, yeah. I mean, the whole field relies a lot on generalization and it's very hard to predict when you're going to get generalization or when you're going to get some kind of out of distribution generalization. So we know that if you train on the task you care about, you're going to do better. But the most important advances are often the types of generalization that we have no right to expect. So, for example, from just pre-training on this very naive next token prediction objective to various tasks of interest that require understanding of the input in some deep way or learning some skill from pre-training that's very rare and not very heavily represented.

10:31John Schulman:And then also generalization from these verifiable tasks to less verifiable ones. This is also a type of generalization that there's no reason a priori to expect it.

10:43Beren Millidge:So this is an interesting question because one intuition pump that you could have for why you would see some sort of singularity very rapidly without even scaling up the inputs to AI progress that are not just AI labor is that before every single experiment you run that's, you know, like a seven figure experiment, you spend an equivalent amount of compute on AI labor. and so you just have automated versions of you guys spending a century thinking about what is the optimal experiment to run, doing small-scale ablations, developing literally a century's worth of theory, so going back even before deep learning, before you decide what experiment to run, doing extremely optimal setting up of the experiment.

11:27Beren Millidge:Then you do a century of thinking after the experiment is over where you're analyzing what happened and what the next experiment to run is.

11:33John Schulman:Well, I think if you think hard enough, you probably could have expected some of these things beforehand. Like there is probably some very clever way to do a small scale experiment that'll let you build the theory that then will generalize to the large scale experiment. So I would expect that like we're nowhere near the ceiling of how well you can do research. And I would imagine a future where AI is doing a lot of analysis and theory building, like spending a comparable amount of compute to the amount that you're spending on the experiments themselves, doing various kinds of analysis and building a theory around what we've seen so far.

12:15Charlie O'Neill:I think there's really concrete examples of this when the objective is well specified. So again, all thinking can do is update your posterior based on the bits that you've gotten since you formed your prior. You can't gain any new bits from just thinking. But when the objective is well-specified and there is this data sitting around, I imagine there will be this big speed-up in the current paradigm we're in. And a good example of this is if you've got an AI to think about the Kaplan scaling laws, an AI at this point would have noticed that they've just taken these intermediate checkpoints and didn't account for the annealing.

12:47Charlie O'Neill:And so this is wrong. And that would have caught that years earlier, we would have made progress, would have cut off a year or two of progress just from that observation from an AI. And again, once the objective is well-specified, which is lower pre-training loss or whatever, there's many, many good examples where if you just thought about it a bit more, you would have been able to cut down a significant on things that you've done. So like mu p and how learning rate scales with model size and realizing the model width is important in that as well. like I feel like you can you can really back out a lot of these things and cut off like a lot of like hanging fruits I would imagine like a 10 times speed up if our thing is just like maximize the objective we're currently on but I don't see that how that generalizes it all to you know come up with the right objective in the first place like just thinking doesn't necessarily buy you the right objective in the first place I mean yeah I think this is really the key question to like any kind of like very rapid RSI is like from current AIs it's like how well can AIs generalize to like learning their own objectives because to have any kind of like self-propelling automated loop, we need the AI to propose objectives, optimize them, figure that out, propose a new objective, and have this not go off the rails at any point for a long, long time.

13:53Charlie O'Neill:Coming back to Moravec's Paradox, there might be a case of Moravec's Paradox where we think there's autonomy and being self-encapsulated so we can think of what we should do ourselves and then go do it and have this loop as super easy because we always do this. And obviously evolution needs to create creatures that can survive on them by themselves for a long period of time. And this just might be something that for some reason is really hard for the AI in the same way that locomotion stuff is really hard. It's like math is super easy despite being super hard for us.

14:18Beren Millidge:I don't know. Doesn't the time horizon increasing suggest that that's...

14:21Charlie O'Neill:Yeah, exactly. I mean, this is another possibility, but I agree, there's no obvious evidence for this. In fact, the fact that RLH is not super persistent and it's quite easy to do this is kind of evidence against this. But this would be potentially one of the reasons why we just don't get this immediate takeoff is if this is hard.

14:38Beren Millidge:If you look back from 2012 till now, or maybe from when you started doing your research till now, what part of all the innovations that have happened since that time, including purely engineering ones, including purely conceptual ones, what seems like the thing that is the thing that would be the last things humans would have to do before AI is totally automated AR &D?

15:05Charlie O'Neill:Probably just like iteratively asking the right questions. Like if you can get the AI to like do any experiment, but like you need to decide what experiments to do and like right now I think AIs are not very good at this compared to coding the experiments at all. Yeah. Like whenever we talk about research they propose like a bunch of like miscellaneous things which are like very very tiny steps. Or even going from like you know DeepMind's approach of like we're going to solve intelligence by learning to play games at a superhuman level. That's going to be the approach to like one random researcher like Radford being like I'm going to try and just predict the next token off a very wide swath of data.

15:37Charlie O'Neill:Like and then even once Radford discover that right like it took a while before people decided to scale up because you had to come up with the idea of scaling walls and the fact that like you could very reliably predict these

15:46John Schulman:things yeah yeah i would say that the last um job for humans uh or the the role for humans that'll last the longest is like defining the objective uh and like deciding what we actually want um so So like in that vein, something like deciding what, how is it like the AI, the assistants should behave or what it means to be helpful or what's like the objective when we're doing our all from human feedback is one such thing. And then like, then later, like defining like constitutions and model specs is another one. And I think even if the AIs can do all the technical work, we'll have to still do a lot of that and decide what we actually want.

16:31Beren Millidge:Yeah, alignment is the final job.

16:33John Schulman:Yeah, alignment is sort of the answer, but it's also alignment itself can be kind of decomposed into like specification of the objective or figuring out what the right objective should be. And then like actually like achieving or optimizing the objective you've defined. And I think the first one is not going to go away anytime soon. And if I think about a post-training team and why you need a lot of people to be on the team, it's just because there are a lot of different areas where you have to figure out how the model should behave. And there's no way of, it would be very hard to automate the whole thing just because someone has to think about how should the model behave in this area.

17:23John Schulman:Yeah.

17:24Beren Millidge:Jane Street started using Antithesis to test their software in early 2025. And they were so impressed by the product that they decided to invest in the company. I recently caught up with Ron Minsky, who co-leads Jane Street's tech group, to ask about how Antithesis actually plugs in.

Read the full transcript

17:38Charlie O'Neill:The thing that I think is most impressive about Antithesis is we started using it in a team that was building high-assurance software and being really careful. And nonetheless, it was able to shake out bugs that were otherwise going to be really hard to find. And that's important both because it helps make those systems more reliable,

17:55Beren Millidge:but also because it helps the teams that build it to just move faster. This matters more and more as code production is increasingly automated. I think in general, as we've been using agents more and more,

18:05John Schulman:the key problem that you run into is the verification bottleneck. Just the time it takes from people to look at code and figure out,

18:12Charlie O'Neill:is that actually something you want to accept in your production software? And tools that make testing better are just incredibly helpful there.

18:19John Schulman:So they just ease the verification bottleneck and make it possible for you to get more stuff done and move faster because you can have more confidence

18:26Beren Millidge:that the code generated by the agent is actually not introducing new problems. To see how Antithesis fits into your development process, go to antithesis.com slash Dwarakash. What is the story for why there isn't huge consolidation in model providers? There's just so many things that point to centralization here. Is there, yeah, if you step back over the course of years, is there something that is going to prevent that?

18:54John Schulman:Yeah, I think distillation is the main thing that fights against the centralizing force. Because basically anything that can be learned through RL can be distilled very easily. Because, like, it's a small number of bits. It's something that you can learn from a small amount of data. So if you can get trajectories from the model that show a behavior, you can easily distill it. So I think distillation is one of the things that fights centralization. There's also, I mean, there is a possibility that there will be company-specific models, that it'll be possible to learn from deployment and have a company continually improving its own model.

19:42John Schulman:And such a system could be provided by the current oligopoly of model providers or some other currently smaller company. But I think that'll change the game a bit.

19:53Charlie O'Neill:Yeah, and I also want to point out that like continual learning and RSI doesn't stop distillation, right? Like even if your model is improving every day, like people could be distilling it every day. So it's like the loops could just operate at the same pace.

20:03Beren Millidge:Right, that makes sense. Okay, so copying model behavior, I guess you need to know yourself what the right distribution to prompt is in order to get the relevant model behavior?

20:16John Schulman:Yeah, for just distilling with supervised learning, the prompt distribution is extremely important. So it's very non-trivial to distill a model even if you have full access to it and have the cot, the chain of thought and everything. yeah, it's non-trivial to distill all of the useful capabilities from it because you need to prompt the model with something. You need to prompt it with realistic prompts. You need to have a really wide distribution of realistic prompts. So yeah, one thing that's been coming out recently is some of the Chinese companies are probably using these router services which are designed to allow people in China to use the U.S.

21:02John Schulman:frontier models, which would otherwise be blocked in China. But there are all these router or proxy services that allow people in China to use these models mostly for coding. And these router services are collecting and selling some of the data. So I think this is a very useful data set for distillation because it gives you the perfect prompt distribution.

21:25Charlie O'Neill:I think this is one of those things where AI has helped a lot here. If you actually look at the frontier pipelines, let's say the Chinese models that they actually put in their papers, it's a lot of humans or they get seed prompts from somewhere, which is some combination of humans, this kind of data. And then they synthesize a vast coverage from those seed prompts using their existing models or the other frontier models. And so it's like you can automate an awful lot of this prompt distribution gathering and environment creation. It's just like humans need to provide increasingly fewer amounts of bits.

21:51Charlie O'Neill:It's like the models get better.

21:52Beren Millidge:Right, but it still seems you're bottlenecked by having a service which has users or users are going through.

22:01Charlie O'Neill:Not necessarily. I mean, yeah, that's obviously very helpful. But theoretically, you can just think about what users want or a lot of tasks.

22:07Beren Millidge:The whole point is that the user says, make me an application like this. Oh, that didn't work. I actually want you to make this new feature. But actually, let's step back and do this other thing. And capturing that whole trace is the... or to the extent you could have done that anyways, then you just have like RSI.

22:23Charlie O'Neill:Yeah, I mean, like ultimately, like if you have this like fully automated loop, that is basically RSI, right? Like the AI is deciding, the data is deciding, the training, that is the loop. But yeah, I mean, like it depends how much human information you need. Like at some point, if you're just like, I want traces that look like this, you prompt that to the model, the model will be able to like come up with like a pretty good approximation.

22:40Beren Millidge:But what if you want to do like, make me a really good politician and then just like anticipate de novo, like how would a discussion in like the Senate halls go or something. I just feel like there's going to be a lot of things.

22:50Charlie O'Neill:Ironically, this is actually, I think, easier for the distillers than the frontier labs, right? Because the distillers just like, I want a good politician. They go to the frontier model. The frontier model already knows how to be a good politician, so it just generates those traces. Whereas if you actually want to build the first model that does this, you have to actually somehow get data on what politicians do every day and build that. So it's actually much easier to say, I want something like this and then get the AI to produce a billion variations than to actually create the thing like this to begin with.

23:15Charlie O'Neill:I think you can actually make a really concrete prediction based off this observation that the Chinese labs have this router data. So I think the thing that Jess did this originally was I was saying, isn't it weird how Sonnet 5 and Opus 5 are almost objectively worse models than GLM 5.3, Kimi K3, even though they've had access to not only distillation but logic distillation from Mythos. And so the counter here was that, okay, the prompt distribution really, really matters. You need to see what users are doing so that you can distill these behaviors and things in. I think the prediction from this is that the frontier labs don't necessarily have much of an advantage, if at all, in aural environments now.

23:55Charlie O'Neill:Because, yes, user distribution matters for general behavior and so on. But the best measure of a capability is the very, very hard aural environments you've made at the frontier. And so if you have access to those aural environments as Anthropic, and you have access to logic distillation, and you've still made a worse model, then maybe like...

24:13Beren Millidge:Then real-world deployment matters more than the environment. Yeah. That's really interesting. So, but they had to incentivize those capabilities in the first place in Fable or the Frontier model. And so it's weird that they can't incentivize them again with, like, a smaller model or something.

24:30Charlie O'Neill:Maybe, like, maybe we're just in this weird, like, uncanny valley where, you know, like, actually trying to copy that Frontier model too much. Like, the student-teacher gap or whatever it is is just, like, too large. and I think people made this point with Opus is it's like the difference between Opus 4.6 and Opus 5 is that Opus 5 really feels like it's got this AI as a judge checking every possible thing it's done. And it's like, that's why it uses so many tokens. It tries to think about all these things, but it doesn't necessarily have the big model smell of fable to know when to stop doing that or when's a good path to go down or whatever.

25:00Beren Millidge:The reach exceeds the grasp. Yeah.

25:02John Schulman:Yeah, it would offer a slightly different hypothesis. So I would say there are a couple of different axes for the environments you can create. And like one of them is difficulty and the other is realism. It's sort of easy to create, or it's comparatively easy to create a lot of difficult environments that are just like involve like doing a much more complicated task or doing something that requires a lot more cleverness. And you could say this is like the benchmarking distribution because a lot of the most prominent benchmarks just involve doing some very hard puzzle-like task that's easy to verify.

25:41John Schulman:And then there's sort of like the realism axis where you want the model to be good in the realistic coding agent setting where there's like multiple back and forth with the human and there's like multiple objectives. And like, I'd say like the people, like the labs who are crafting the model behavior for the first time need to push in both directions. And to get good model behavior, You need to really push on the realism axis and have like rubrics or some kind of human feedback that's informing the reward function you use there. But I think when if you try to do distillation naively, you end up just sort of matching the teacher on the bench maxing distribution.

26:23John Schulman:and but if you don't have enough of the environments that really exercise the capabilities in these like trickier realistic settings then you're not going to you're not going to get those into your student model and i think maybe one thing that's happening is the big models generalize better from the like the tricky narrow tasks to these sort of more realistic tasks so if you have a really good like realistic prompt distribution for distillation, you can match the big model really well. But if you only have this distribution of easily verifiable tasks, then you can match the big model on all the benchmarks, but you do worse on this broader distribution.

27:09John Schulman:So that might even explain something about the smaller anthropic models like Sonnet 5, though it's hard to predict exactly what they're doing to post-train those models. It could also be that they're always changing their post-training stack, and they just got a few things wrong in some of these models. So they turned up something too high and created some quirks that people really don't like. So it's really easy to screw up post-training in some way that doesn't show up in benchmarks.

27:44Charlie O'Neill:I mean, just one of the sort of very basic point is just like the frontier labs buy all their data from data companies. And like the Chinese can also just buy the same data from data companies. And they are, right? Exactly. There's a lot of people like, you know, being annoyed about this, but like if they have exactly the same data and like they can buy that, they can also distill, it's like, it's, it means it's quite easy to like keep up really.

28:05Beren Millidge:Yeah. Okay. The other question I had is how the first models that are capable of automating AI R &D will actually be trained. Because there's a toy version, which is this thing that Ryan was talking about, which is you just have GPT-8 try to build GPT-3 size models that are really good at like inner loop type challenges of beating video games that require continual learning or just getting to a certain loss with like the least amount of compute, et cetera. But John, I think you had an interesting point that maybe that's not the way it actually will happen in practice. So I'd be curious about, by the point at which you have AI that are actually capable of automating AR &D, how are they probably trained?

28:45John Schulman:Yeah, I think we'll probably do some combination of learning from human feedback to absorb the researcher's taste and just creating a lot of practice environments, which involve doing multi-step research projects. So I think, yeah, people will in practice do some combination of those two things and just each iteration, like patch whatever seems to be most broken. in the last iteration. So like researchers will be using the AIs a lot and and will notice that they have some consistent weaknesses and then those things will either be patched by like collecting human feedback or like creating environments.

29:28John Schulman:Yeah, maybe maybe useful way

29:30Charlie O'Neill:to think about this is like how much of the lineage we roll back and then let self play from there. Like I think in the limit like you're picturing like you know just giving them like a GPU and maybe neural nets or something and saying like, okay, figure out how to train a model to like do these particular tasks. Like the way it currently works is like we go up to the very like edge of the lineage and say, okay, like here are the bugs like, you know, Anthropic has found in their training stack in the last few months. We'll turn those into environments like you need to train to get better on the frontier.

29:59Charlie O'Neill:And so you obviously lock in all the previous history of the lineage, but you could imagine a world in which you roll back to like, you know, before GRPO or something and then you have environments which like trying to get it to discover like the best will form to like RL models on and then maybe roll further and further back but I think we will be still so compute bottlenecked that like people will just keep like staying at the frontier and like diffing essentially the bugs and whatever improvements they found since the last model

30:22Beren Millidge:version turning those into training environments which is also really good for having non-steal like new data between model generations it's just again this is basically continual learning within in the AI lab of distilling the last three months of AI research progress through environments and like RLHFs type stuff back into the model itself. And it is distilling, right?

30:43Charlie O'Neill:And that's maybe why some of us feel like it's asymptotic is like you're always like just trying to get the last three months of progress. And that progress is being contributed to by AIZ, of course, but it also still has humans in the loop. And it feels like, you know, you're just constantly inching closer and closer to what the human researchers are like finding and capable of doing. Yeah. I mean, the one thing I will say though, So it's like, obviously, if you're just distilling on trajectories, you can never go above it. But environments can go quite a far way above what the humans can do.

31:07Charlie O'Neill:It's very easy to design an environment that no human can solve, but the AI can always still try and solve it. And so that would be the path to go ahead of just what the human AI research is.

31:15Beren Millidge:Do you have an example in terms of RSI or what kind of training is that? But doing it even faster than a human speedrunner. Yeah.

31:23Charlie O'Neill:I mean, I feel like in AI research especially, it's very easy to define goals, which you could say the loss needs to be 1.3 or something. And no human can get that. now but like that's a very extremely measurable verifiable task and the if the AI gets there then then great.

31:36Beren Millidge:Right on a building like a hundred million parameter model that beats Minecraft that's maybe too easy but like it beats a much more complicated game or something.

31:45Charlie O'Neill:Isn't it crazy that a hundred million parameter models to beat Minecraft we're calling that too easy like imagine if you said that like five years ago.

31:52John Schulman:I would say a lot of research is not exactly like that though where it's like hill climbing on a well-defined goal it's sort of more like here's an intuition we have about some way models should be better. And then we also have some idea for an algorithm that seems to go a little bit in this direction. So let's come up with a task that is sort of designed to show signs of life on this approach and see if we get those signs of life. And then if we do, we can make successively more realistic versions of the

32:25Beren Millidge:task right it's like a lot more guided by intuition and then the the out the inner loop is to elicit the uh or make tests for that intuition rather than like the the the test itself leading to the

32:39John Schulman:insight right like you're not directly optimizing uh for the eventual objective you care about or the practical like production objective it's it's sort of uh you're um you're relaxing your objective a little bit you're saying yeah let's relax on the realism axis a little bit and find uh some methods that actually work and then like then try to get back to realism later after the method matures a little bit and then there's also like more uh there's research that's more oriented towards explaining things and like uh developing a theory or a sort of yeah often we don't have like mathematical theories in machine learning that are that um predictive but we have like a a lot of more informal theories for what's going on.

33:26Charlie O'Neill:Yeah, I mean, presumably the models will be trained on some combination of all of these tasks, and some will be very easily verifiable, some will be like LMS judge, or just ask the human, does this look reasonable? And then the hope would be that these would all generalize to these much harder, more vague, fuzzy kind of tasks, and it probably will to some extent, whether it generalizes enough that the loop can become self-sealing without humans being in the loop at all, it's unclear.

33:50Beren Millidge:Yeah, yeah, yeah. Maybe taking a step back, here's what it seems to me that the plan for AI research going forward is, and you tell me if you think it's going to work or if you agree with this characterization. So the bet is that we will scale up our LVR training across millions of diverse environments, across hundreds of different kinds of domains. And what will emerge at the other end is an agent which has learned these basic skills or less than basic skills around being persistent, being able to triage information in context, eventually having end-to-end optimization of working with other agents and things like that.

34:30Beren Millidge:and such an agent will be very sample efficient within the context. You've done research on how you actually scale up in context learning to make it arbitrarily long, but you just keep scaling it up. And so what comes out the other end will be something that basically functions like a drop in remote worker over the course of a week or a month. First of all, do you agree that that is a bet the labs are making? And second, is that enough? Like basically learning how to learn within the simulacra, within a data center, and then getting deployed into the real world but not actually like learning from real world deployment, only learning these meta skills from the simulated environments in the data center.

35:07Charlie O'Neill:Yeah, I think it's now hard to separate out like how much of the lab's effort is going towards like direct RSI versus like making generally intelligent models that they can continue to deploy to collect revenue to fund the next big training run. I think for the latter, like yes, that's probably just the bet they're making. Like, and it's very clear, like the pattern of like where these environments are going over the last few years. I mean, like, Anthropics lineage of environments is like a very clear example of this. Like, you know, first, like, we just focus on coding and, like, we're going to get really, really good at that.

35:38Charlie O'Neill:And then the task horizon that we've got from coding, which is probably the lowest hanging fruit in terms of, like, data available on the internet to create environments, like, their own internal stuff that they can turn into environments. Then we're going to generalize. We're going to go after finance next and, like, literally, like, just so much Excel data and all that sort of stuff in the RL training. And then, you know, it's PowerPoints. It's, like, this long tail of like the working economy. And like that seemed to work really well. And like a lot of the other labs and thing, even the open source labs have now realized that that was the correct bit.

36:06Beren Millidge:But what is the implication from that? When I had Dario on the podcast, the thing I asked him was, if you truly expect models, which will be human-like in their ability to learn on the job, why would you try to bake in all these skills of like working with PowerPoint or something? Wouldn't you just expect the model to be able to pick that up on while it's deployed? And so yeah, there's multiple different explanations. One is just that this is, we expect models to get there soon, but they're not there yet. So why not amortize these skills into the model training? Another is that we're not concentrated on making it really good at widely deployed work.

36:40Beren Millidge:We just want it really good at RSI. And this is just like a way for us to get revenue so that we can pour it back into a model that is actually really good at doing RSI development. And then once the singularity happens, the thing that comes out the other end will be really good at all the things which seem like bottlenecks to the current generation of models. Yeah, John, I don't know if you have a taste on like what, how much I construe why there is so much task-specific knowledge in these models if the path is like this kind of generalization.

37:07John Schulman:Yeah, I mean, if the models were good enough at learning in context, then in theory, you wouldn't need to train them on finance. They would just be able to figure out, read all the books on the fly and figure out how to do everything in the appropriate jurisdiction. Yeah, and you could argue that you need to do a lot of this domain-specific training just to make them more efficient. So even if they were smart enough to figure this out on the fly, you still might want to do a bunch of RL and bake all these intuitions into the weights so the model would be more efficient at runtime. Yeah, I'd say in practice, it does seem like model providers are going domain by domain and trying to strengthen the models in the highest value domain.

37:56John Schulman:And I'd say that that's one of the answers to why the models have gotten so much better. It's just because the model providers have covered a lot of the high value domains and the most common types of skills.

38:10Charlie O'Neill:I mean, I think another thing is just that it's not that expensive to do both at the same time, right? Because the models are massive. They can easily afford in terms of their parameters to learn everything. And there is likely some transfer. and sort of even just, even if like finance is not specifically like the information is important for like RSI, just the general like meta learning of like how to figure out what's important, how to have taste, how to like do long horizon work is potentially generalizable. And like, there's not that much RSI like data in the world as well. Like it's kind of hard to generate and like that requires a lot of effort to like, if you can sort of amortize in this other data, get some transfer from it, you already have masses of compute and massive parameters based to like, why not do that as well as like obviously the direct like commercial incentive of like selling a model.

38:47Beren Millidge:That makes sense.

38:47John Schulman:Oh, yeah, I'll add that, I mean, there's one question about whether this current paradigm of doing like sim to real will be the dominant one forever. So basically, you look at what the real world tasks are like, and then you try to create a bunch of environments that can be simulated, like in the data center, and you can do RL on them. And I think obviously this has been very successful, but it has a lot of weaknesses because a lot of things are just kind of hard to simulate, especially if they involve like interacting with a bunch of humans in real time. Yeah, so there's some question about like whether Sim2Real will be the dominant framework forever.

39:30Charlie O'Neill:I think Simtorial has to be the dominant framework where sample efficiency is kind of low. Because right now you need thousands and thousands of interactions with the humans. And no human is going to sit there and deal with this, basically be in the loop of RL training. And so we kind of have to simulate that now to get the samples you need. But obviously if sample efficiency improves a lot, you'd expect learning from deployment to become a much bigger part of it.

39:51John Schulman:Though there are also other things you could do. Like you can learn off policy so you can take all the traces. And even without re-simulating everything, you can potentially learn something from them.

40:01Beren Millidge:jane street just launched a new competition and it's their most ambitious one yet design a protocol emulator asic basically if you have a chip that you want to test you can connect it to this asic and then this asic will simulate realistic traffic that way you can see how the chip responds without having to plug it into a live system jane street is looking for flexible general purpose designs not single protocol emulators when i was chatting with them they suggested that i start off by trying to implement what are apparently three very common protocols, UART, SBI, and I2C. Jane Street also mentioned that they hoped that more ambitious designs will also tackle low-speed USB and Ethernet and any other protocols that flex your chip's specific architecture.

40:42Beren Millidge:Importantly, your design should be reprogrammable rather than smashing a bunch of specific protocols onto a chip. If a new protocol comes out after your ASIC is taped out, your chip still needs to be able to handle it. How exactly it does it is up to you. But there is one hard constraint. Your design must target an open source 130 nanometer process node. That's because Jane Street will pay to tape out the most novel submissions and send the physical copies to the winners. The competition is open till January 18th, 2027, and working in teams is highly encouraged. Go to janestreet.com slash thorkesh to download the template code and get started.

41:22Beren Millidge:I want to ask more about this because it's sort of weird that you have 50 % of compute that's spent on inference that is not directly helping the model become better. Like one of the key advantages you'd expect eventually digital minds to have is unlike a human who gets to have 50 years of like real world experience, a model will get to, through all its instances, will get to experience, I don't know, millions of years of deployment across all kinds of economically relevant work in the economy. And right now, that data is just not in a meaningful sense helping the model get better. It just seems so obvious that eventually models should be able to learn from this data.

42:00Beren Millidge:And once they do, you would have something that almost feels like a widely deployed intelligence explosion because the model is assimilating so much information across all these deployed instances. But when do you expect this kind of hive mind kind of crazy shit to start happening?

42:13Charlie O'Neill:I think broadly, at a very basic level, this is already happening, right? Just in the next generation of models. So like right now, you can obviously take your deployment data and put this in the pre-trained or the mid-trained of like future models, especially if you do like some kind of filtering or some kind of like judgment or annotation or like recent, you know, synthesization of that. How much do you think that explains the generation over generation improvement? I think it explains like quite a bit. I mean, especially like, I mean, this is, you know, I don't know whether the labs do this because, you know, theoretically they claim not to train on people's data.

42:39Charlie O'Neill:But like the Chinese 100 % do. And like they definitely get this advantage, both like obviously deploying. This is basically what distillation is. Like they take other models. They get some of their like deployment data. they get some fraction of that by like pinging the model and then they train their next generation of models on it and they can suddenly do it on their own models as well like there's no reason not to whatsoever i completely agree with this i think if you zoom out far enough this is like definitely happening like you're picturing this like and we're all picturing this is like what continual learning like the holy grail is is like this very very organic like live loop of like an individual model like getting an experience and like live updating on the spot and learning from that and like a lot of things break when you like zoom into that level of granularity but like yeah the big labs are doing this like the closed the closed models are doing this there's also early signs of life of like people using open source models doing this at a much faster cadence so like a good example is probably like composer um like you have some sort of model and you are able to or like you know harvey's doing the same thing with like legal legal agents like it is getting very specific environments from the data that you have for that particular task and things that like you know users are complaining about and like all the feedback that you're somehow extracting from like your specific deployments and a lot of these companies have the advantage over the big labs and that they can use this data really really well yeah and then they will create environments they will like you know do a big post train of Kimmy K3 um they will go deploy it they might do some online learning as well like Composer did online basically like reinforce for a long time um so yeah like there's still a human in lieu there's still a human saying okay these are the signals we care about here's how we're going to create environments from the data that we have and there's like still a longer cadence than maybe the one that you're thinking of but like it really is happening.

44:14Charlie O'Neill:And like eventually that loop will become like faster and faster.

44:16Beren Millidge:I mean, the composer thing is interesting because this is where the model, like in cursor people like press tab or they don't press tab on the next completion that the model suggests. And based on that, every single day composer gets better at like predicting the next.

44:29Charlie O'Neill:So that was, that was the old tab model. Like they actually did the same thing for the actual, not just like the tab model, but the actual like generative model.

44:36Beren Millidge:Oh, that's interesting.

44:36Charlie O'Neill:And they, it was, it's hard because when you do online reinforcement learning, you don't have groups, right? You just have one user saying one thing and then you get one rollout. And so you have a big variance reduction problem. And Cursor's kind of fuzzy answer to this was like, oh, we have very good heuristics which are able to estimate how much better than average this response was or how much worse than average this response was. And then they would do this big reinforce update. And then their solution to whether it got worse or not was if it improved on CursorBench, they would deploy the new model every five hours.

45:06Charlie O'Neill:And if it didn't, they would throw that version out.

45:08Beren Millidge:Interesting.

45:09John Schulman:Yeah, I think your biggest problem is actually just not knowing what the reward function should be from natural data. And if you use some kind of superficial signal, like did they accept the code, the edit, that might get reward hacked in some way.

45:24Beren Millidge:But this seems like a bigger issue with the Sim2Real thing, where the longer and longer horizon tasks get, the harder they are to simulate within a data center, right? It seems to me already, potentially, even in coding, we're getting to the point to where there's like not some year-long coding task that doesn't eventually require you to like talk to a client or interact with the company or interact with users. And if you think about the gamut of things we would want AI to be capable at, you want eventually super intelligence to be able to like run a business or like start a new business and make it profitable or like have a profitable day trading in the markets or win a court case.

46:02Beren Millidge:And these are all things which are very hard to simulate in a data center. Like an inherent part of the learning there is interacting with the real world. And so maybe they may have to learn how to get better at these things from the transfer between sim to real. But alternatively, maybe you do need weight updates from these kinds of interactions in order to get better at them. And then if that is the case, if transfer isn't strong enough and you do need weight updates, then the fact that the models are quite sample inefficient is maybe a deeper problem. And the reason I'm curious about this, I feel like by default, I don't see how you don't get some kind of crazy recursive self-improvement within the next 10 years.

46:39Beren Millidge:But the one reason why that might not happen is in terms of like weight updates, the sample efficiency of weight updates, they just seem way far behind humans, right? Like plausibly million fold behind humans in terms of how much data a human sees from birth to adulthood versus how much a model sees from, you know, like cold start to like finishing training. And so, yeah, this is all to say, first of all, is there going to be a good transfer between simulations and extremely long horizon and really complicated real shit that we want the EIs to do in the real world? And if not, does that really mean that the lack of sample efficiency in these models comes to bite us?

47:13Charlie O'Neill:I think maybe the way I'd break down the two types of tasks in which models get good and models will still continue to struggle is whether the task is cumulative or you have this non-stationary distribution you have to keep learning and relitigating a bunch of stuff. So maybe an example of a cumulative task might be RSI. It's theoretically possible that maybe have less than a million token Python file, which from scratch trains a model that is capable of recursive self-improvement. And every discovery that you make is kind of a line in the sand that you hold. If it's true that for RSI, we don't need to discover a new attention variant or whatever.

47:49Charlie O'Neill:Once you've discovered attention and then once you've discovered a mixture of experts and once you've discovered GRPO, you just add that to the training stack and that's there. And a good example of this is 5.6 Sol training, 5.6 Terrell, whichever one OpenAI told us to train. It didn't have to go back and discover retention. It basically probably would have called a bunch of scripts, which is like pre-training.sh and post-training.sh and just did that. So that's an example of a cumulative task. I think the real world and the reason people are thinking so much about continual learning is it's not really a cumulative task.

48:20Charlie O'Neill:Imagine in a law firm, you have an agent acting as a legal associate. That's a very non-stationary distribution. You have to be able to fit in your context all the relationships between all the important people in that company, which are also changing all the time. You have all these implicit ways about how things are done, where to find information, et cetera. And that's not as clean of an example of a cumulative task like RSI is. So I think that there will be this breakdown between tasks. But if the labs realize that and they do believe that RSI is cumulative in the sense that we don't need to go back and discover some brand new architecture or whatever, then maybe more and more effort and compute gets focused on that versus the RSI.

48:56Beren Millidge:It's so unfortunate that RSI happened to be easier than the apparently old. yeah um yeah i don't know if you guys have thoughts on this yeah i would say there's uh

49:06John Schulman:like models um are today's models are um weaker than humans in a lot of different ways and uh some of them might have to do with um sample efficiency in a certain regime uh where i mean in some regimes models are very sample efficient like learning in context uh but then there might be some like medium length regime where they're less sample efficient because humans can uh do some kind of weight update more efficiently than models. So I think being less sample efficient in certain regimes might be one of the sources of weakness, but then I think there are other sources of weaknesses that are completely different than that.

49:45John Schulman:For example, having lower diversity of thought than humans or being bad at certain kinds of long horizon judgments. I mean, I think a lot of what people call taste is something about behavior that works in the long run and that people have realized works in the long run. Not everything, but some aspect of taste, especially for something like software engineering. I think a lot of taste is what are the systems that are going to be maintainable and work well in the long run of this project. So yeah, I think the weaknesses of humans, which limit RSI along with other things, there's a variety of them and some of them are related to sample efficiency and some of them aren't.

50:37John Schulman:Maybe an interesting thought experiment is like

50:39Charlie O'Neill:if you were able to give a model like a context window of, I don't know, a trillion tokens or whatever you would have needed to fit in like your experience prior to like, let's say, RLHF. And like it's got all that experience in the context window and it has the same sample efficiency and in context learning ability as it does at a million tokens. Like, do you think taste is then solved? Like, would it be able to, like, make the same judgments that you did? Or is there, like, something fundamentally missing apart from just a longer context window with the same sample efficiency?

51:07John Schulman:Yeah, I mean, it would have to be trained to learn from that context. So I'm not sure. Yeah, either it would have to be trained to learn the right update to make from that context or you don't think you can just, like, dump it all in?

51:23Charlie O'Neill:like your whole like life like research experience I mean like you still need the data to train it long context right like even if you could theoretically get like a trillion context you would need a trillion lengths of data to train it like right now you have like I'm just asking if you had that in theory I think yes I mean this really just comes down to the question of like how meta learnable is taste from like shorter horizon episodes and like I feel like there's no obvious reason it's super long because like humans somehow develop taste with not having many long episodes like we don't live to be like 10 ,000 we have like we develop pretty quickly, right?

51:54Charlie O'Neill:And so if you think about even in a PhD, the difference between a first-year PhD student and a final postdoc or something, that's five years maybe, and they've only done maybe 10, 5, 30 research projects in total, but somehow they develop taste quite quickly from a relatively short succession of small things. And so theoretically, it's possible to develop it like that. The AI obviously will have vastly more experience in which to develop taste to meta-learn it, and then it's like how well does that generalize to really long-horizon things, is I think the question, which I think is really unsolved at this point.

52:23Charlie O'Neill:Like we don't know.

52:25Beren Millidge:Going back to this question, eventually there should be a regime where AIs are learning a ton from each individual instance of deployment that they have. Well, currently you could say there's a meta fuzzy process by which models do improve for deployment, but I feel like it's a very weak feedback loop. Do you see this around the horizon where there's this like hive mind kind of learning that's very rapid? and if so, how exactly does it happen?

52:51John Schulman:Actually, I would say that around, will we get a hive mind that learns from all of its deployment experience? I mean, a big part of that is actually about incentives rather than being a technical question. So like companies aren't gonna wanna have the model provider learn from all of their deployment because that might just reduce the advantage of their business.

53:15Charlie O'Neill:I think that maybe the economics of this will pressure not necessarily weight updates to one big common shared model, but kind of like modules that get subbed in. So a very obvious example, this is Allura, but it might be something else. There's been a lot of work to try and fit an arbitrary context length into a fixed size. This is all the linear tension stuff and all that sort of stuff. And cartridges, which are essentially KVCaches trained to be very, very compressed KVCaches to fit in a lot of information. That's another example of something that companies may be willing to sign up for if that gets subbed into the model, and it's not actually changing the base underlying model itself.

53:51Charlie O'Neill:So there's many different versions of learning from your data in real time, and the latter ones are not really helping the big labs because they are just these modules. But I think the economic pressure will force the labs to go down that path first before they can embark on this. So which economic pressure there? Because I feel like even if you have a bunch of cartridges or laws or whatnot, you can still just like take all these traces and just like distill this, dump this to the pre-training of like your next generation of bombers. So it may be a more indirect form of learning that the big labs are getting and that's obviously still really valuable to them.

54:24Charlie O'Neill:But I can't imagine a world in which we start off with like, you know, we're going to just like directly train this one big model on like all the exact data. No, I think it will definitely like go through stages because I mean, this is assuming there's like one discontinuous event where it's like suddenly we fix like weight updates continuously and like in practice, I think it's much more likely to be like the cartridges and stuff allow you to specialize and deployment. Then you generate traces, you put that in your model. Like three months later, you come out with a model which is better at this stuff.

54:47Charlie O'Neill:You specialize it again and you like consolidate it again. And then eventually, we'll just like make this loop faster and faster. So instead of like every three months, we release the model. Now it's like every week and then every like day and then every hour at which point we basically have obviously solved it. Yeah. And I think this is a good point as well because you asked like kind of how far off the current paradigm we are from being able to do this. I think we've done a bit of research to this and people have done a lot of research. Like at a really large scale, like when you wash out enough noise and you have large enough batches, like this outer loop process of like putting data into mid-training, creating our own environments, like it does work in like some sort of continual learning regime.

55:19Charlie O'Neill:But the problem is like when you're zooming close enough at like a micro level, it's like I've got one model and I'm trying to update it again for like a law firm or something. And I'm trying to do that very continuously like with a relatively small amount of data. Like all the methods kind of break down a bit. So like if I SFT the model on just like, you know, successful traces, off policy, on policy, like eventually in the very iterative regime like when you're doing like you know hundreds of these micro updates you see both catastrophic forgetting you see forgetting of like previous information I've learned on top of the base model that was much earlier on and I see degradation of general like use general capabilities you know on policy distillation seems to like push this horizon out a little bit but it still eventually succumbs to the same thing and RL is not very good at like it is good at like getting capabilities in but it's not as good as getting like knowledge in and like just this very explicit knowledge of like, oh, okay, like this person does this at this law firm and like this is a very specific process we find.

56:17Charlie O'Neill:And you have to pour in a lot of compute to create the right environments to get the knowledge in.

56:20Beren Millidge:And do you think that the fundamental issue here, why you get worse at these other skills or there's forgetting and stuff, do you think it's fundamentally an issue of capacity or it's an issue of techniques? A little bit of both.

56:31Charlie O'Neill:I think like SFT and even like distillate, like on-policy distillation can be like way too destructive. like the reason RL is so nice is because like yeah it changes a very very small amount about about the model and there's like a lot of evidence for why this is the case and so like it kind of just like tweaks it in this very very very small like loss valley to like get it into the right point but that also then limits what you can do with RL like how much you can actually change the model.

56:56Beren Millidge:You're saying like the reason this isn't the winner take all potentially is that it's just like very hard to distill that much information into the base model? Without ruining something internet or vision.

57:06Charlie O'Neill:Like it's easy to distill it into like a different base model. Like this is where I think it's mostly technique. It's not like, it's definitely not like just like there isn't capacity. Like if you had some modeling with all this data and you take like nearly the same size model and pre-trained it from scratch with like all of this stuff in mid training, it will be better. And I think that's a lot of what's happening today. Yeah. And so it's very much like there's, you know, a bottleneck that stops us from just keeping training the same model forever versus just like getting all the data from the old model and like training a new model from scratch.

57:29Charlie O'Neill:And this is exactly as Jolly was saying, like some combination like plasticity and like catastrophic forgetting. And that like, you know, if you just naively train on like non-stationary data because you're adding new data as you go, basically this is messing with the data distribution. So like the old stuff is just forgotten. And we don't really have good methods to like stop that from happening.

57:45Beren Millidge:And so maybe in the limit, you're like just bottlenecked by retraining the model from scratch with all this information. Yes.

57:51Charlie O'Neill:Which of course is like very expensive. Like training model from scratch is expensive. But you're going to do that anyways. I mean, not necessarily. I mean, like maybe eventually if you have continual learning, you never train a new model. You just like,

58:03Beren Millidge:might be like some deep technical reason why that's very difficult because of these like I mean that's the question that's the question yeah I think we have pushed back like how much

58:10Charlie O'Neill:from scratch we need to do like it is definitely possible now to take like the pre-trained base and like do very good mid-training on top of that like kind of continuously plus some RL from like different checkpoints that are later on in the training and like that's looking more like continual learning but certainly not the case of like you know take the most recent model apply a couple of very small updates and like iteratively like never lose it but isn't this like I'm a bit

58:31Beren Millidge:confused because isn't this literally what happens during training or during post-training or something? You just have a model that's already gone through so much training and then you distill some fork that's been further RL'd or something. Isn't that literally what happens? It's still at a large enough scale, I think, that you're washing out

58:46Charlie O'Neill:a lot of the noise. And you're not just focused on one distribution, which as Baron said, is like, you know, that is now a very if you're just focusing on one task, right? I mean, in the

58:57Beren Millidge:eventual regime, you'd be doing I don't know, there's billions of deployed instances you're like doing you're learning from all of them at once and so hopefully there's some washing out of noise and stuff from that right? Maybe that's good yeah.

59:10Charlie O'Neill:Yeah I mean I think like definitely as I was saying like you can do continual mid-training for like a long time and you can like roll back to a checkpoint give a new mid-training data but at the same time like you can't do this like indefinitely like if you just keep continual mid-training the same base forever it just like get it does it sort of asymptote at some point like you can't just learn new stuff in that base and this is why people end up training new bases like otherwise you would just keep mid-training the same pace forever.

59:32Beren Millidge:Whenever I finish recording an interview, I immediately brain dump all my thoughts into Slack. Things like what was most interesting and what should get cut. This ensures that my editors have all the context they need to start editing the episode. But it's not like these brain dumps have any clear timestamps and my unedited recordings are many hours long. It can take a ton of editor time to even find the exact moments that I was referencing. So we decided to try adding a GrokBot producer to our chat. Now, whenever one of my editors posts a rough cut of the episode, GrokBot opens a transcript on its own computer and starts working, usually before I've even seen the message.

1:00:06Beren Millidge:It takes the notes that I dropped into Slack, and it highlights the relevant snippets in the transcript. It also uses a big case file that I've compiled with all my preferences, so it can suggest potential edits. And when it's done, it sends me its top clip candidates so that I can review everything from my phone. This has worked really well. Being able to send informal messages, like I'm texting my editor, and then having the transcript immediately reflect my preferences has just been so helpful. Try GrokBot yourself at x.ai slash bot. Okay, let's talk a bit about data now. So I'm generally interested in this question of how much of AI progress is just explained by data progress.

1:00:42Beren Millidge:Doesn't mean it will be necessarily hard to automate, but that's a separate question. So is there some data distribution, which if you trained current architectures on, would result in a super intelligence? That totally dominates human experts across every single field.

1:00:58Charlie O'Neill:Are we talking about like pre-training plus post-training data like environments as well? Like I think the existence of this is obvious. It's just like whether we can create the right environment. Yeah, I mean in the trivial case we could just train it to output the Python file which like trains the actual super intelligence. Like just have them memorize in the way it's. Like yes, there's probably like a ladder of RL environments that is possible to construct. such that you would get an AI researcher which is at least as good as a human researcher. But the effort to climb each successive rung grows exponentially.

1:01:32Charlie O'Neill:And that's going to be the two things that you have to trade off against as to how fast we're going to hit that final rung where it's better. I think that's fairly clear. And I think we're still relatively early in our own environment creation. There's a lot of asymmetries that we exploit in order to create good environments. So one of those asymmetries which we've talked about before is like there's environments where it's easier to go backwards than forwards. And what I mean by that is it's very easy to define this complex data generating process. And this is like this kind of latent variable you keep hidden from the model.

1:02:04Charlie O'Neill:You can generate arbitrarily complex environments and the model has to do a lot of irreducible token spend and irreducible work to figure out what that data generating process was. There's asymmetries in terms of like you can inject information from the real world. So like Anthropic finds a bug. through tens of thousands of human and LMLs combined and turned that into a very, very neat environment, which in a single LML could theoretically find within a few million tokens. So there's all these asymmetries which we're cherry-picking and we're counting on this task horizon generalization. But I think, yeah, again, there's just going to hit diminishing returns at some point.

1:02:39Charlie O'Neill:At some point, there's diminishing returns and how hard it is to create these environments in the first place, like coming up with them, because you can't necessarily just have these processes where it's easier to go backwards than forwards. Like you actually have to sit down and construct like something that looks like with humans, like, you know, a long enough time horizon, like it's going to be a really complex task to create. And then there's also going to be like the compute and time bottlenecks for the agent to actually do those tasks. So like, I think you're just going to start seeing this like curve to fly up now.

1:03:06John Schulman:I saw something about how someone fine-tuned the Taki model, which is only trained on data up to 1930 on this like modern coding agent data. and it did better than Claude 3 Opus on Sweebench. So this model that has like no knowledge of code whatsoever can be fine-tuned on a moderate amount of data and like behave better as a coding agent than this much larger pre-trained model is pretty crazy. And it kind of shows you that like once you have an example of like the right expert behavior, it's actually surprisingly easy to like copy that into a like relatively weak model.

1:03:47Charlie O'Neill:Yeah. But a counter example to like that kind of is there was a paper recently where they trained it up to like fifth grade maths. And like also like primary school, like English and stuff. So it was like a decent language model. And they tried to RL it to do like, you know, late high school and college maths. And the gap was just too large. Like they couldn't get it to climb at all. Like, but if you did like successive wrongs of like, you know, year seven maths and then year eight maths and so on, like you could obviously climb to year 12. So again, it's just like, what is the distance between the rungs on those letters and how hard is it to create?

1:04:18Charlie O'Neill:Yeah, and this just comes back to the RL signal problem. RL is not very good at exploring right now. And so if the model can't get in 128 rollouts, it's very unlikely to get signal to progress. And this is why in RL we need curricula, whereas in pre-training we don't, because that's not a problem for pre-training at all. Yeah, and again, pre-training data is different to post-training data. and I imagine as we continue on like yeah humans will be involved less and less but that doesn't change the fact that you're bottlenecked on like how much signal you can extract from the real world so like there's a lot of signal in the world and that's true like you know there's people doing like spreadsheet tasks there's people doing like legal tasks and all this sort of stuff but you know the capability frontier of where the models are at now like how many bits in the world actually like really relevant to like improving the the model's capabilities like you know how many new math problems are being solved that couldn't just be on the reach or grasp of the current models?

1:05:11Charlie O'Neill:How many new coding problems are being created or solved that are beyond the reach of the current models? And I think that's why the diminishing returns kicks in because even the world as a whole is not giving you the bits that are useful for tipping you into the next basin of capability. Yeah, I totally agree with this. It's really a question of where the signal is coming from. And so the signal doesn't... In pre-training, the signal is already in common call, right? like for the tasks that you care about in pre-training, the problem is there's not just like getting signal at all, it's like filtering out all the noise that exists.

1:05:40Charlie O'Neill:And that's quite an automatable process. But as the models get better, as we enter mid-training and post-training, the signal just doesn't exist anywhere in the original data we have. No amount of filtering will get this. There's no hidden proof of the Millennium Prize problem sitting in common core. We can just filter until we see it. And so at that point, you have to get bits some other way, either from humans directly asking them to write out their reasoning or or by creating environments where humans decide what environment should be created, what the objectives of these environments are, or some kind of training on the human data that exists in deployment.

1:06:11Charlie O'Neill:You have to get the bits from somewhere.

1:06:13Beren Millidge:Yeah, yeah. There's a question of how much of the progress in pre-training is being driven by data. I did this investigation with Jerry Hahn, who's a student at Princeton, where we basically trained all the recipes from 2019 till now, pair-wise with all the data sets from 2019 to now. So you train like GPT-2 on the newest data set, like UltraFineWeb, and you train Delphi, which is the newest training recipe or the open source training recipe on like the pile or some old data set. And you do like the whole grid and you see the getting to some level of capabilities, how much less compute does it take across this grid?

1:06:51Beren Millidge:and you see that the data seems to explain like 9x of a compute efficiency gain, but the architecture improvements explains like a 3x compute efficiency gain at a very small scale. And so to the extent that that is true at large scale, that most of the pre-training compute efficiency gains are coming from better data, how much can that continue? Like, can you keep just filtering data more and more and building more and more synthetic data until, yeah, do you have a sense of how much this kind of pre-training progress can continue?

1:07:22Charlie O'Neill:I think my prior is that, again, the low-hanging fruit is somewhat exhausted. We got the internet as this big block, and it's not like the internet is necessarily growing at the same rate. All the useful stuff on the internet is growing at the same rate. We've probably got a bunch of 0.1 % loss drops to go, but definitely not as many as have currently occurred. That's also really interesting that you find this cumulative 27 times improvement across both. I think it was EPOC or someone who estimated three times a year since 2019, which would imply something like 3 to the 7, like over 2 ,000 times improvement.

1:08:01Charlie O'Neill:So where's that missing 100 times or whatever coming from? That probably gives you a good signal of how much of this is post-training.

1:08:08Beren Millidge:I think the explanation has to be that a lot of the computer efficiency gains are scale-dependent, and we're studying at extremely small scale. And that raises the question of do the data computer efficiency gains or the algorithmic computer efficiency gains have more scale dependence? I don't know if you guys were prior on that. We just didn't have enough compute to investigate that question. I mean, just naively, right?

1:08:32Charlie O'Neill:Theoretically, the scale dependence of the architecture is fairly well known. Yeah. And you can fit a straight line to it, whereas I would have no idea how to do that for combining pre-training plus post-training data and mid-training data. I feel like data is actually more important with scale. I feel like architecture is kind of like a one-time, I'm like, you know, an architect, I think like combining, like saying just like an X percent efficiency again is kind of misleading. Because like what an architecture does is like let you reach like a qualitatively new regime, which you couldn't reach with the old architecture.

1:09:01Charlie O'Neill:And then within that regime, obviously the data is like the primary thing determining it. But like, you know, if we say didn't have like, even like GQA, we're doing like full attentional day. We wouldn't be able to do like a million, it would be like ridiculous expenses to a million context. And then like, because of that, we couldn't, we can never use the data, which is like actually at a million context. context. And so we couldn't get these capabilities, even though like if you just do a naive, like how much does this do at like 2K context where the architecture isn't unlocking anything then like the data, you know, that will look much more important than in some sense it is, right?

1:09:27Charlie O'Neill:It's unclear to me that these things are like really just like multiplicative gains in this way. I see. Sorry, but then what does this take away for the scale dependence of data? So I mean, on scale dependence, I think like a lot of the like mid training and post training data we have now is like actually gets better with scale. Because like a lot of it, like the very long context horizon environment stuff really requires like big models to be able to like make use of them. And like this is not, you know, if you try and train like your 100 million parameter model on like three bench traces, it's not going to get anywhere.

1:09:55Charlie O'Neill:Like it's not going to show you the same kind of improvement that you would get if you train like an actual sensible size model on it. Yeah. And like it's hard as well now because so many of the architecture changes, like you look at like Kimi for instance, like, or DeepSeek, they're doing these architectural modifications with not just like dropping the pre-training loss in mind, but like, for instance, how the models are going to be used in the real world. So like the inference efficiency, like having some form of compressed attention in the deep-seq models is not necessarily geared around, you know, this is fundamentally like a period of improvement.

1:10:23Charlie O'Neill:It's just like, okay, we're considering how the models are going to be used. Right, right, right.

1:10:26Beren Millidge:One question I'm curious about to understand the future is how parameter scaling will go as we're getting into more of a RL heavy regime. Like, I don't know. I don't know how fast historically. Yeah, you can look at sort of open source architectures and see how fast parameters have been scaling. and maybe it's like roughly 2x every year for frontier open source models. And to the extent that like even frontier closed source models have like 100b or 200b active parameters. Do you think that like keeps 2xing year over year or another RL regime where you also want to conserve compute on rollouts?

1:10:59Beren Millidge:And also maybe there is like a threshold effect where you have enough capacity. And at that point, increasing parameters arbitrarily doesn't matter as much. Do you guys have a sense of in 2030, how many active parameters will a frontier model have?

1:11:10Charlie O'Neill:Yeah, I think for the next few years, we're going to be like, because we're so focused on doing longer and longer horizon rollouts for RL, where inference efficiency matters a lot, it feels like the moles aren't necessarily saturated on their ability to do that, where the bottleneck is still the environments. And so we might see a little bit of plateau. I have a feeling that Mythos and the GPT models are much smaller than the 10 trillion parameter range that people are talking about. even just naively comparing to open source models you can probably back out that conclusion. So yeah, probably for the next few years I wouldn't imagine a huge growth in the number of parameters but again, there's so many different things to trade off here.

1:11:51Charlie O'Neill:You decide the size of your model based on how much pre-trained data you have and then the difficulty of the RL environments that you've got to train on and you ideally want to get to the optimal point where you can get a decent pass at one or something on the hardest environments you have and it wouldn't make sense to make a bigger model past there because then you're just paying much more inference flops when you need to. So there's a lot of inputs to this. It depends on how quickly, like McCore and then in-house, these guys can scale up the complexity of the aural environments they're training on.

1:12:20John Schulman:I would expect the models to keep getting bigger just because people are scaling up compute and the GPUs are getting bigger. But I would say exactly how much they get bigger depends a bit on the scaling laws in non-obvious ways. So one thing is that I think like data efficiency is going to be a bigger driver than compute efficiency of like the exact architectures people use. Now that we're getting to the regime where we're sort of running low on like high quality pre-training data. So that might affect how sparse you want to make the model. And then I also think we don't understand sparsity that well.

1:12:57John Schulman:And it's like parameters are a different resource than active parameters, but it's, and like sparsity has definitely increased a bit, but it's not clear that it's going to keep increasing without bound. There might be some kind of sweet spot. There's an argument that sparsity should make data efficiency worse because you might have to learn the same thing on multiple experts. So that's debatable. So I think we don't, I don't think we have a good enough theory of scaling laws that we really understand why sparsity is helping and how much it'll help and if that'll like plateau at some point at a certain level of sparsity.

1:13:38Beren Millidge:Can you spell out exactly what the implication of data efficiency would be on, so it sounds like you'd say, well, there should be less sparsity, but what are the other implications on pre-ammer scaling?

1:13:49John Schulman:I guess just that the scaling law, you're not necessarily looking for the most, you're not trying to optimize compute efficiency. So you have all your choices you can make on the architecture and each of these gives you a different scaling law. And then like traditionally you would look at some kind of envelope based on compute. So you would look at performance versus compute and take the envelope of like the best models. but like if we're making that decision based on data so it's like yeah we're sort of assuming we can spend a lot of compute and like we're sort of data is on our x-axis instead of compute then we just get a different set of optima or a different set of models that are on that frontier.

1:14:38Charlie O'Neill:Yeah and I also don't think that we've necessarily like you know doubled like the size of the models every year for the last few years like the people have been training like one trillion parameter models for at least a few years like there was even an open source one called falcon but like liam from periodic labs like i think posted yesterday on twitter about how like an early experiment at open ai was like training a one trillion parameter

1:14:59John Schulman:model that was very very sparse oh yeah that was what they did before open ai that was like

1:15:03Charlie O'Neill:google the switch transformer yeah so like it was like um you know very very good at like knowledge but terrible at reasoning because it was so sparse and so like yeah it feels like we've been playing in this like 100 billion to you know up to 2 trillion parameter range for like at least a little bit and like it certainly hasn't been this is like nice linear increase yeah i mean i feel like there's two things so as charlie was saying like inference efficiency is super important for rl rollouts and so like this will really push down active parameters quite a lot and then i think the total parameters really depends a lot on the hardware as well so like you really need to get very high memory bandwidth and VRAM size to actually be able to serve multi-trillion parameter models.

1:15:41Charlie O'Neill:And so right now, people are still using a lot of H100s and stuff. And so as everyone moves to GBs and then the rubins will get more actual... The ability to scale and actually serve and do large RL inputs at larger scales. The data question I think is interesting because naively, larger models are much more sample efficient in the actual data points. And so even if you're not saturating the model, it's still better to go bigger because the larger models generalize better and get to a better loss for the same amount of data. And so right now I think we kind of have a lot of data and that's not the constraint rather than computing.

1:16:13Charlie O'Neill:So we're having small models, which are very inference efficient. But if computers are no longer at the bottom, it might come back to larger models, which are sort of undersaturated, but they have this generalization ability because they're much larger.

1:16:25Beren Millidge:If you just look at the basic Scentschilla scaling law and you just maximize out parameters, it actually decreases the amount of data you need to get to the same loss very little. If you go to infinity on parameters the amount of data you need I think goes down less than 10x just because of the nature of the power line.

1:16:43Charlie O'Neill:But we're now on the way too much data side of the Chinchilla laws, right? So right now we over-train models of the Chinchilla and so we could easily go back to a point to which as we're running out of data we move back to the Chinchilla optimum point or even a bit on the over-training, under-training model side. But surely even with these new chips that come online and stuff like we're just going to be so compute bottlenecked for the next few years that that won't necessarily be okay. I mean, this could well, yeah, this depends on like the ratio you have like training and inference to compute really.

1:17:07Charlie O'Neill:It's like if you're super bottlenecked on data and not on compute, you should go bigger. If you're super bottlenecked on compute, you should always go smaller. And then like, yeah, but you can also use compute to generate synthetic data. So it's like one of these very hard things to predict.

1:17:19John Schulman:Yeah, I think part of the reason it took people so long to figure out the scaling laws in the first place was that if you don't get all these things right, then you don't get such a clean relationship. And the beautiful stray lines on graphs hide a lot of complexity on how you have to make sure to scale every hyperparameter the right way or parametrize your optimizer in a way that scales and where you don't have to change your hyperparameters as you change the model size.

1:17:46Charlie O'Neill:And bugs have their own clean scaling laws as well, right? Like, you know, like with Kaplan forgetting the cosine annealing thing or like even just like not considering embedding parameters, I think. And so that messed up the estimate at smaller models because embedding parameters are a decent size of the model.

1:18:03Beren Millidge:A bit on RL. So I feel like a year ago, a lot of people were making this argument that RL will not be super successful at scaling for models. I think, John, you wrote a research paper where you were pointing out that models learn one bit per episode. When you RL, they basically learn, did I get the answer right or did I get it wrong? And I wrote some blog posts earlier this year. I was like, it's even worse than that because when the pass rate is low and the model is very unlikely to get the answer right, it learns almost nothing at all from an RL episode. But I look at the models today, and they seem pretty smart, and it seems to be the result of scaling up RL.

1:18:41Beren Millidge:Baron, you had a post, I think, a few weeks ago where you're trying to explain what's going on. But why has RL been more successful than one would have naively thought?

1:18:49Charlie O'Neill:I mean, so I think the success of RL comes down to a bunch of different things. So first, I think what is slightly underestimated is actually the mid-training. So an awful lot of what we see as successes of RL actually comes from very, very good mid-training data, which is basically where we're essentially doing pre-training, but on synthetic reasoning data and the kind of environments that get the model warm started for RL. And so this actually takes the model almost 80 % of the way to the final RL checkpoint often. And then what RL does on top of that is it does a lot of essentially tweaking to the policy.

1:19:20Charlie O'Neill:And so this is one of the reasons why it doesn't need as many bits as you would naively think. It doesn't have to learn all of these behaviors from scratch. It needs just a few bits from these episodes, which you do get. And then the other thing that I really point out in my blog is that these bits are actually extremely high signal compared to regular pre-training, which is why you need RL at all versus just like SFTing on like successful reasoning traces.

1:19:40Beren Millidge:Because it's exactly the bits about how to get the answer right.

1:19:44Charlie O'Neill:Well, there's two things. So yes, one, it's exactly the bits about how to get the answer right. But like this is not exactly how you think of it because in SFT, you have a trace, right? You have like a bunch of math reasoning and then the answer at the end. The bit is still there. Like you still SFT on the answer token. So that bit is still there. What's important is that the objective ignores all the other bits. So in SFT, you like have like, you know, to try and match like the exact reasoning tokens that the model produces. So you're essentially getting like too many bits about like the exact way this other model you're training on reasons.

1:20:11Charlie O'Neill:For RL, you only get the one bit. And that means that like this, that signal is not drowned out in the noise of like all the other bits the model has. And so that's what really like, it's really a super dramatic like increase into the signal to noise ratio during training, which is why like RL is like so dramatically efficient in terms of steps.

1:20:28Beren Millidge:I don't know if you guys have thoughts on that.

1:20:30Charlie O'Neill:Yeah, I like, there's been so much debate about like what RL does to the model versus like, you know, mid training or SFT or whatever. And like, you know, everyone talks about how, you know, parser one will go up, but parser 256 will go down. Like very rare, correct reasoning traces will be like downweighted and kind of like outweighed by a gradient signal from like easier kind of reasoning traces. And I, I think the simple like way to view RL now is that if you have a large enough, like a large enough amount of compute to sample a large enough group size, such that your probability of getting a bunch of correct answers is like past some, like not insignificant probability, then like it will be up weighted and like to to veron's point like basically mid-training and you know more pre-training like the the parser one the starting point for rl like scales in a long

1:21:15Beren Millidge:number of pre-training tokens you can answer very basic questions um i i guess that answer makes sense and maybe there's empirical research which shows that this is what's happening but then i just look at the models themselves and i don't know what's happened so like maybe you can give me a sense of what is the basis of the AI progress over the last year. But if it's, yeah, maybe it's just up weighting the policies, which we're going to do the correct thinking anyways. But it just seems like qualitatively, the models have gotten so much more capable. And anyways, maybe there's nothing to, there's no inherent contradiction there.

1:21:56Beren Millidge:But how do we square like the relatively small impact this take would imply that RL would have from the actual qualitative capabilities the model seemed to be gaining.

1:22:06Charlie O'Neill:So like one thing I want to point out here is that like it doesn't necessarily imply that RL has a small like effect right even if you have a few bits and like you only change the parameters a small amount like the actual impact on like function space the model and like the input to output mapping can still be like super dramatic you know even if it's like even like one bit can change like your function space a lot and it can like will add like half the hypothesis space which is huge so like I don't think it's necessarily the case it's like small amounts of bits, small amounts of RL, once you're starting from a really good point, means that you don't have dramatic impacts in behavior.

1:22:33Charlie O'Neill:At least not necessarily. I think it comes down to two things. I think the first thing is that everyone was hoping that RL would generalize this reasoning across all these different domains. And I don't think we necessarily got this horizontal generalization. Just training on math doesn't necessarily make you the greatest coder. You do have to do RL on code environments. I think what we did get, though, is horizon generalization. like the models just learned how to use more tokens for longer and still make progress on some sort of task and so like you can train on environments where they get longer and longer and longer and then put them into a completely new environment and yes like they may not have generalized the reasoning patterns which allow them to do well in that environment but they've at least generalized the ability to like continue on that task for longer which is correlated with like success I think there's a paper called Edge Bench which show that the rate at which models can work for longer is like doubling every three months and so that's a clear evidence of generalization And I think the final way to think about it is in pre-training, there's this idea of quanta.

1:23:31Charlie O'Neill:So you have this very smooth pre-training loss curve. And when you actually look at what's happening in the model, the model is learning all these very discrete tasks and there's all these emergent points. There's kind of a phase transition. It didn't have induction heads, now it has induction heads. And there's tens of thousands, millions, probably hundreds of millions of these things. And you average them all together and you get this very smooth loss curve. I think like to an extent like a similar thing is happening for RL like there is this very slow outer loop as Baron mentioned of you know we will train a model and then RL it and then like the next kind of model iteration of training we will dump a bunch of these synthetic reasoning traces into the mid-training data like we're kind of hitting all these quanta for all these different tasks and like on an individual task level it may look like a phase transition and like you're suddenly going from like a 0.5 % pass rate to a 90 % pass rate on a particular finance task or Excel task or whatever, but you average all these things together and plus the horizon generalization, you kind of go, wow, we've got qualitatively better models.

1:24:29Charlie O'Neill:Yeah. I mean, I think a lot of this as well is just like, I think RL does generalize a bit. Suddenly you get some transfer between math and code or puzzles and math and this kind of stuff. Also, just the sheer amount of environments I think that people are targeting is just vastly greater. So before, when you tried to do some task, which you do in your daily life, like two years ago like the labs wouldn't really care about this they wouldn't like train the model for it and now like it's just so much poorer they have a lot of environments

1:24:53Beren Millidge:targeting this specific thing earlier in the conversation we were talking about RL in the context of causing this entropy collapse or just you know concentrating probability on solutions the base model were already done and causing relatively sparse updates in the policy but when I think like when I think about I think there's also another story about RL which is going back to the Atari game and then AlphaGo coming up with Move 37, the super creative move, that because it was never initialized on human data, it can think in ways that humans are not even thinking and come up with extremely creative solutions.

1:25:28Beren Millidge:Yeah, do you have a sense on when we should expect or if we should expect RL on LLMs to result in things like Move 37, just extreme creativity, even beyond human creativity because there's just de novo initialization of intelligence?

1:25:44Charlie O'Neill:I mean, so a couple of things here. Like, first off, I think that the alpha goers using MCTS, which obviously does like more exploration and like stuff than regular policy gradients. But I kind of also think that like RL doesn't necessarily like reduce the creativity. And like, I mean, even if we look, I think, you know, this is obviously qualitative, but if we look at like the, you know, the open area hugging face incident, like these models were coming up with like multiple zero days at a time to like break out of the sandbox. And like, this is clearly like some level of like move 37 creativity, I think already, which we just get from just like the general generalization properties of the LMs.

1:26:16Charlie O'Neill:Like, I don't think it's definitely not the case of like, oh, it's like totally destroying their like entropy. Yeah. Especially on long horizons. Yeah.

1:26:23John Schulman:I mean, one thing that people call creativity is just solving hard search problems. So, and so that's like, like Move 37 is obviously an example of that, or like writing some kind of poem that satisfies a ton of different constraints. So that's something AI is obviously going to be extremely good at if trained for it. And then there's another way in which the models, like the diversity of their outputs is a lot lower after RL and they sort of develop these ticks. And like even though the models seem like they're good at writing, when you do some kind of like distributional analysis, you find that like they're reusing certain themes like all the time and they're using the same character names all the time.

1:27:12John Schulman:So there's actually, it's not like you're getting the same kind of diversity that you get when you, like from human authors, you're sort of getting one really good like style. So I think that like that kind of diversity has definitely been like cut down by RL a lot. And in fact, since we were talking about distillation earlier, that's sort of something. yeah one thing that's happening is that so many people are distilling uh mostly from clod that like uh like all the open weight models write the same way as clod and use the same like have the same ticks so this uh seems kind of concerning to me that we're having this uh like this monoculture

1:27:53Charlie O'Neill:emerge yeah again i don't think this is like fundamental to rl as like a method though and same with distillation like even with distillation like you're just training on the data it's like just because your data is not like super broad that doesn't mean like the training method itself is somehow wrong. It's like a problem with the data. And I think a lot of, for instance, like the RL, like entropy collapses basically due to like exploitation of fairly simple, like verifiers when you don't have like a huge diversity of environments. Because like, for instance, like the writing, I think the writing is presumably graded by some judge and like the judge has some specific ticks and like the model is learning to award hack the judge.

1:28:24Charlie O'Neill:And that's why like it collapses. But like, this is really a problem with the judge. It's not a problem with like RL in general.

1:28:30Beren Millidge:Okay. Super rapid fire predictions about the future. So I want timelines on the following couple questions. By when do we have models which you can, here's what it feels like to a user. You basically hire them as a drop-in remote worker for all kinds of white-collar work, not just coding, but I don't know, video editing, law, paralegal, et cetera. Like it's like literally an actual remote worker. with like full computer use with like literally a month of seamless learning and operation and executing on like complex projects and it required interacting with other people, et cetera, et cetera. It's like everything a human worker could do over a month.

1:29:15Charlie O'Neill:If you like mandated to use like a browser or whatever, rather than like these, again, the firm setting up the information to be like programmatically accessible, like maybe a couple of years. But if it's not like browser-based, like it can send Slack messages, it can do all this stuff. I'd still probably say around a year. Yeah, I mean, I would say maybe for the full generality, maybe three years. But I think, to Charlie's point, we will end up with a lot of people making their organizations easier for the AIs to use. And so you get 80%, 90 % of the way there before that.

1:29:45Beren Millidge:Sorry, but the diff between one year and three years there is just literally like...

1:29:49Charlie O'Neill:I think there's going to be a long tail of miscellaneous stuff, which some human can do, which will take them all quite a while to do.

1:29:56Beren Millidge:Are you thinking of computer-year stuff or basic cognitive capabilities?

1:29:58Charlie O'Neill:I mean I think this really comes down to a question of like how quickly can we solve this kind of like online learning and like whether we can like get like 80-90 % of the way there with like compaction and like writing files to yourself and stuff and like that's my big uncertainty I really don't know and another like maybe an example of something that I wouldn't be good at is like you know if I have to like yell at someone to get something at work or like really push someone to get something done like the model isn't just going to do that it's just going to be too nice yeah

1:30:24John Schulman:I'd say there's a wide variation in quality of human remote workers So if you try to hire someone like off of Upwork to do a software engineering project, there's going to be a huge variation. It's like often quite hard to get them to do like to do a good job or like pay attention to all the feedback you're getting. And like I would guess that in some cases it like it'll be worse. Like the pre-AI version of this was worse than what you can get now from existing AI. uh so i think it might end up being a little complicated uh because maybe to some extent we already have this uh like for some like not so high quality of work but then like uh then it's obviously like we're not yeah we're not matching human level in certain like higher quality like um forms of work so but i basically agree with charlie and baron that maybe yeah we'll yeah we'll have some version of this in a year or so that's like okay and it will be able to do maybe we'll have that form factor and it'll be able to do some things really well some things not so well and yeah things will be improving from there like we ship the goalpost based on

1:31:34Charlie O'Neill:a very long tail all the time like I think I feel like you've used this example before of like doing your taxes or something like this year I literally just like told Codex to like go get everything I needed to do and send it to the accountant and like there was this massive list of stuff it had to use computers to click through and like download some stuff and it didn't it was it was like fine it was perfect so like i don't know a lot of this stuff it can already yeah okay um give you 10x total productivity uplift basically if you if it takes you a year to make a breakthrough now

1:32:03John Schulman:you make a breakthrough every month i think i would just refuse to give you a scaler on this uh like like we might already be past that in some like types of work uh like let's say you're just trying to prove uh yeah you're trying to do um like certain types of math uh oh sorry but for

1:32:21Beren Millidge:you as ai researchers trying to make um yeah like advance you know the state of ai research is how much are like ai researchers sped up yeah we're uplifted somewhere between five and ten years

1:32:33Charlie O'Neill:oh really okay that's far away well you think it's longer than like for general remote worker

1:32:37Beren Millidge:yeah interesting i think you're right i think i'm realizing you probably have very different definitions of fully general remote worker. I could have specified that.

1:32:44Charlie O'Neill:Yeah, this is true. Because I mean, like, yeah, because obviously like an AI researcher can be a remote worker. And so like... Yeah, no, I'm picturing like, you know, normal white collar work over the period of a month. Yeah. I think it starts to diverge

1:32:55Beren Millidge:a little bit past the month. A very competent white collar worker, but not necessarily like a super creative researcher.

1:33:00John Schulman:I would say like two years. Two years? Yeah.

1:33:03Beren Millidge:10x?

1:33:03John Schulman:Okay.

1:33:04Beren Millidge:How are you, Bernd?

1:33:05Charlie O'Neill:I can kind of see that actually. Because like, it really is just like, right now it's already like definitely more than 10x of coding stuff. And so it's like, if it can do even like one or two loops of like experimental feedback, that would actually be massive already.

1:33:17Beren Millidge:So 10x uplift of AI researchers within two years, if you just plug it into like a very naive model of like AI progress and how much is coming from AI researchers and there's like a 10x increase in their productivity. Yeah, you have like radically accelerated pace of AI progress starting two years from now.

1:33:36Charlie O'Neill:Yeah, I mean, I think like this will mean that AI progress doesn't get bottlenecked on like AI researchers ability to run like small experiments it gets bottlenecked and other things of course of course but it just like happens 10x faster for sure yeah which is a huge deal and that also like helps the

1:33:48Beren Millidge:next thing which makes gives you 100x speed up happen sooner etc yeah i'm happy to stick with

1:33:54Charlie O'Neill:longer on that one and what's what's like the crux uh like my capacity to absorb information and make the like beige and optimal decision on the next makes sense yeah i mean i'm assuming that like you can delegate some of this to the ai so like the ai is becoming decent at like deciding, you know, it's run this experiment, it's got this result, it runs like the next experiment. And then if it can run like two or three experiments in a row without like crashing, then like that is actually a big update, like uplift.

1:34:18Beren Millidge:And okay, final question. An AI which is, which dominates top human experts across every single field of work that can be done over a computer. So not only AI research, but all cognitive work. And not just like short horizon work, but like literally if it takes like three years or something. they also would do better than humans.

1:34:41Charlie O'Neill:This is basically just like ASI. Okay.

1:34:43John Schulman:I would say like three or four years. The fuck?

1:34:48Beren Millidge:I mean, that doesn't seem wrong.

1:34:50John Schulman:I mean, I would say like, like AI is obviously being more, getting more attention. So it's like one of the harder things, but it's like a lot of energy is being put into it. And it's also like not one of the hardest things for AI because it's like involves a lot of code and math, which models are really good at. Maybe for things that involve 3D and spatial stuff and physical stuff, I think that will take a little longer. Especially if it's mechanical engineering or something and it's not getting the most attention right now, that might take a little longer.

1:35:27Beren Millidge:But it also does include fields where there is relatively little data because of the nature of the field. And it has to learn that data on the fly. so for example has to become superhuman at like being an engineer at TSMC or something

1:35:41John Schulman:so you would have to assume that like the onboarding yeah you can give the AI the same onboarding material and then there's some like something has to be solved about like sort of longer horizon learning

1:35:57Beren Millidge:I'd say 5 to 10 basically yeah you think automating AI research is like ASI complete or something. Yeah, I think so.

1:36:07Charlie O'Neill:Yeah, I think there's so many things in the world which like, even if you have some sort of memory system external to the model and even if like context length grows a little bit, like there are just fundamentally things like even if you could research the information or write notes to yourself, like you'd need more than a minute and take the context to be able to do. Yeah, I mean, I kind of agree in like the five-year range, at least for like the stuff that like labs are focusing on, but I think like there's going to be a long tail of stuff which like the AI could theoretically go out and learn about but like no one has bothered to do it and like the compute isn't being allocated to that so that might take longer for like literally every single human expert.

1:36:41Beren Millidge:That's right but by this I also included like the ability to learn as fast as a human a new domain.

1:36:45Charlie O'Neill:I mean I think that's not necessarily necessary actually because like the AI will have vastly greater experience than like any human. Right.

1:36:51Beren Millidge:Thanks for doing this guys. I feel like this was a great format for getting different experts to disagree and debate and discuss things together. It was very productive.

1:36:58Charlie O'Neill:Cool. Thanks for having us. Thanks. Thanks for having us.

From the publisher

New episode with John Schulman, Beren Millidge and Charlie O’Neill. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.

Watch on YouTube; read the transcript.

Sponsors

* Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street’s tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to antithesis.com/dwarkesh

* Grok Bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at x.ai/bot

* Jane Street just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy! The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at janestreet.com/dwarkesh

Timestamps

(00:00:00) – Steelmanning the case against RSI

(00:18:39) – What’s driving the Chinese labs’ progress

(00:28:06) – How will automated AI researchers be trained

(00:33:51) – Will long-horizon RL elicit AGI?

(00:45:24) – The sim-to-real gap

(01:00:33) – How much progress is explained by data?

(01:18:03) – Why is RL working so well?

(01:24:54) – Move 37 and entropy collapse

(01:28:32) – Rapid-fire timelines



This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit www.dwarkesh.com

More from Dwarkesh Podcast

All 94 episodes
AI researchers debate how close we are to recursive self-improvementDwarkesh Podcast · 1 h 37 min
Listen in VO