Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)

23 Oct 2025 · 1 h 10 min · 30 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that AI progress is still accelerating in a sustained, exponential way (not slowing), and that this implies major economic impact soon. It connects “task length” improvements to agentic autonomy, discusses benchmarks like MetroEval and OpenAI’s GDPVal, and predicts milestones for 2026–2027. It also addresses whether AI can produce “Move 37”-style novelty, potentially enabling novel science and eventually AI-driven research.

Guest backgrounds

Julian Schrittwieser is a core contributor to DeepMind’s AlphaGo Zero and MuZero, and is now a key researcher at Anthropic.

Key claims

Exponential trends are hard to intuit; frontier labs show consistent benchmark gains (e.g., tasks getting about twice as long every few months) with no slowdown. A “wider ecosystem” bubble could coexist with frontier-model revenue growth. By mid-2026 agents can work all day autonomously; late 2026 models match industry experts; by 2027 they often outperform experts. Pre-training plus RL is likely the right practical paradigm; discontinuities are unlikely.

Notable examples

Move 37 (AlphaGo’s unexpected unconventional move); AlphaGo (deep nets + Monte Carlo Tree Search, trained from human games); AlphaGo Zero (self-play from rules only); AlphaZero (generalized to chess/shogi); MuZero (learns implicit dynamics for planning).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI Progress

0:00 to 0:33

Exploration of the rapid advancements in AI and their implications.

“The talk about AI bubbles seemed very divorced from what was happening in Frontier Labs and what we were seeing.”

Exponential Trends in AI

1:06 to 2:26

Julian discusses the perception of AI progress and its future trajectory.

“A couple of weeks ago, you wrote an incredible blog post that broke the internet entitled Failing to Understand the Exponential Again.”

Predictions for AI's Future

2:26 to 4:12

Julian shares his thoughts on AI advancements and potential economic impacts.

“But it's very hard for us to intuitively understand these exponential trends because it's just not what we're used to in our normal environment.”

Forecasting AI Capabilities

4:12 to 5:36

Discussion on the capabilities of AI models in the coming years.

“And so it's possible that there may simultaneously be like some sort of bubble in the wider ecosystem, while at the same time, the frontier labs on a very solid trajectory, having a lot of revenue, making a lot of money.”

Evaluating AI Performance

5:36 to 7:42

Julian explains benchmarks used to evaluate AI's economic impact.

“if you look at other benchmarks think we would have something like next year, maybe the models will be able to work on their own for a whole day worth of tasks.”

Assessing AI's Real-World Impact

7:42 to 9:07

Exploration of real-world adoption and productivity improvements from AI.

“So I think that's a super cool evaluation.”

Can AI Surpass Human Intelligence?

9:07 to 10:32

Discussion on the potential for AI to exceed human capabilities in various tasks.

“How do new runs go compared to past runs?”

Novel Discoveries by AI

10:32 to 14:00

Julian elaborates on AI's ability to create new concepts and its implications.

“A key question of the moment is, can it be, to which extent can it become better than humans?”

AI Discovering Novel Science

14:00 to 14:59

Explore how AI is making significant scientific discoveries.

“So not just one move, but like a whole new idea, new concept.”

The Future of AI Breakthroughs

15:00 to 18:12

Discuss the timeline and potential for AI to win prestigious awards.

“But yeah, I'm not very worried because I see this process continuing.”
Show all 30 chapters

Sustainability of AI Progress

18:13 to 22:40

Examine the balance between AI productivity and research challenges.

“that we improve in productivity so much that we can actually accelerate.”

Personal Journey to AI Research

22:41 to 23:28

Learn about the guest's path to becoming an AI researcher.

“I think I'm often quite pragmatic about this.”

The Evolution of AlphaGo and Its Impact

23:29 to 28:00

Understand the development and significance of AlphaGo and its iterations.

“Yeah, actually, when I was a kid, I didn't have any expectations of becoming an AI researcher.”

Understanding Tree Search in AI

28:00 to 28:48

Learn how tree search relates to decision-making in AI games.

“and then use the tree search to really make a big plan of what are all the possibilities in the game.”

AlphaGo's Journey and Its Surprises

28:48 to 30:28

Explore the historical context and challenges faced during AlphaGo's development.

“And by the way just for the lore of it did you guys have any sense that AlphaGo was going to crush Lee Seedal, so the famous goal player that you mentioned earlier in the conversation.”

From AlphaGo to AlphaZero: Evolution in AI

30:28 to 32:09

Discover the key differences between AlphaGo and AlphaZero in AI development.

“So basically, you know, it would play and we would tell it, you know, who won, who lost or, you know, cannot make this move.”

The Innovations of MuZero

32:09 to 33:33

Understand how MuZero revolutionized AI by predicting environmental outcomes.

“And also, we as a human, we don't do this, right?”

World Models in AI and Their Implications

33:33 to 36:29

Learn about the concept of world models in AI and their importance.

“What did you learn about the general power of search and learning that is today relevant in modern agentic AI systems?”

Pre-training and Reinforcement Learning in AI

36:29 to 38:39

Explore the relationship and challenges between pre-training and RL in AI systems.

“If you think about super high resolution video, audio signals, it's a very large amount of data that probably you don't actually need.”

Scaling Reinforcement Learning: Challenges and Solutions

38:39 to 41:29

Discuss the complexities involved in scaling RL and the necessary strategies.

“I think the main challenge or the main thing you need to watch out for is that you don't over-encode or you don't restrict your search space too much.”

Rewards in AI: Progress and Trends

41:29 to 42:00

Examine the latest trends and state-of-the-art approaches to rewards in AI.

“It's really useful to be able to isolate the component and say, I have non-good data over here, I have a non-good target there.”

Research Trends in Reinforcement Learning

42:00 to 44:40

Exploration of compute trade-offs and latest developments in reinforcement learning rewards.

“But I think if you look at all the RL literature over time, we see very similar returns on compute in pre-training and in RL where we can invest exponentially more compute in RL and keep getting benefits.”

Data Generation and Quality in RL

44:40 to 47:44

Discussion on how reinforcement learning generates data and the importance of data quality.

“Again, following the evolution from like AlphaGo, it used to be human data and then like self-play.”

The Role of RL in Agent Development

47:44 to 50:52

The significance of reinforcement learning in creating autonomous AI agents and their learning processes.

“more stable is by improving this, by, for example, putting more reasoning into your language model to generate much more high quality training data.”

Future Directions for AI Development

50:52 to 53:35

Insights on improving AI capabilities and the challenges that lie ahead in AI development.

“building an AI app, and I build it on top of Anthropic, Anthropic is whatever model is going to come with some of this sort of batteries included.”

Evaluating AI Models Effectively

53:35 to 56:01

Exploration of Goodhart's Law and strategies for creating unbiased benchmarks for AI evaluation.

“And then how should labs compare results so that it doesn't end up with this kind of leaderboard theater that we've seen a little bit in the last couple of years?”

Measuring Model Performance and Evals

56:01 to 1:00:01

Understand the challenges and approaches in evaluating AI models.

“Make your own internal benchmark that really represents what you care about and then measure on that.”

Safety and Alignment in AI Development

1:00:01 to 1:06:19

Learn about the rigorous processes for ensuring AI safety and alignment.

“For the last part of this conversation, I'd love to zoom out and talk about the impact of AI.”

The Impact of AI on Jobs and Society

1:06:19 to 1:08:17

Explore the implications of AI advancements on jobs and economic structures.

“And it's much less a technological problem.”

Optimism for Future AI Developments

1:08:17 to 1:09:24

Discover the potential breakthroughs in various fields driven by AI.

“To get more wealthy, we really need to grow the pie.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The talk about AI bubbles seemed very divorced from what was happening in Frontier Labs and what we were seeing. We are not seeing any slowdown of progress. We are seeing this very consistent improvement over many, many years where every, say, like, you know, three, four months is able to, like, do a task that is twice as long as before, completely on its own. It's very hard for us to intuitively understand these exponential trends. If you manage to make everybody in society 10 times more productive, you know, what kind of abundance can we achieve? What will we be able to unlock in the next five years?

0:31I think we can go extremely far. Welcome to the Mad Podcast. I'm Matt Turk from FirstMark. Today, my guest is Julian Schritweiser, one of the world's most impressive AI researchers. Julian was a core contributor to DeepMind's legendary AlphaGo Zero and MuZero projects, and he is now a key researcher at Anthropic. We covered the exponential trajectory of AI and his predictions for 2026 and 2027, the frontier in reinforcement learning and AI agents, and the science behind AI creativity and the famous Move 37 from AlphaGo. Please enjoy this fantastic conversation with Julian. Hey Julian, welcome.

1:07Hey Matt, thanks for having me. A couple of weeks ago, you wrote an incredible blog post that broke the internet entitled Failing to Understand the Exponential Again. And what is it that so many people are missing about the current trajectory of AI? Yeah, it's funny that you bring up that blog post. I really didn't expect it to blow up that much. I actually had the idea when I was on holiday in Kyrgyzstan a few weeks ago on a very long car ride. And then I started thinking about this and like all the talk about AI bubbles I've seen on X and this discussion. And it seemed very divorced from what was happening in Frontier Labs and what we were seeing.

1:48and that made me start to wonder a bit like is it that things are moving so fast that people maybe struggle a bit to extrapolate and understand intuitively oh you know maybe it's far away now but you know it's doubling every so many months which means that once it gets close to us it's going to move past and become really good very quickly and that reminded me a lot in like a different way but like what happened during early covid where we had a similar situation where you know at the beginning it's like very few cases it's like wow you know it's never going to happen and it's only a few hundred people who cares.

2:19But if you understand the math and if you look at it, it's like, oh, it's going to double every week, two weeks. Clearly, it's going to be a massive scale. But it's very hard for us to intuitively understand these exponential trends because it's just not what we're used to in our normal environment. And so that's what got me thinking, oh, is something similar happening here with AI, right? We are clearly, if you're looking at many benchmarks, we have many evaluations we have, we are seeing this very consistent improvement over many many years where every say like you know three four months is able to like do a task that is twice as long as before completely on its own and so we we can extrapolate this right and we see that in a year from now maybe two years from now it's the top models are going to be able to work completely on their own for like a whole day or more combined with this combined with the fact that there's a huge number of knowledge-based jobs in the economy, knowledge-based tasks, and combine that in the frontier labs, we are not seeing any slowdown of progress.

3:22Just extrapolating those things together over a very short time, like half a year, one year, already that is enough to know that there is going to be massive economic impact. That means if you look at current, if you look at OpenAI, if you look at Anthropic, if you look at Google, those evaluations, those revenue numbers are actually fairly conservative. I think some more thoughts, some things I have seen more recently is that it's maybe actually even more interesting and more complex that, you know, while those frontier labs and frontier models are clearly very capable and on an extreme trajectory, there are like a lot of other companies, right, that are trying to follow into the same AI sphere that may also have very high evaluation, but not necessarily the revenues to support it.

4:12And so it's possible that there may simultaneously be like some sort of bubble in the wider ecosystem, while at the same time, the frontier labs on a very solid trajectory, having a lot of revenue, making a lot of money. I think that may be quite an unusual situation that in the past, you know, maybe in the dotcom bubble, maybe people were talking about, you know, the railroad rush and stuff like this. We did not see this bifurcation. So I think, yeah, I've been thinking about it more and I guess I think it's getting more and more interesting the situation. Fascinating. So you alluded to some of your predictions or extrapolations for 26 and 27.

4:53Do you want to unpack that? You had three of those. Maybe, you know, calling it my prediction is giving myself too much credit, right? I will just say this. If you look, for example, at MetroEval and you very naively extrapolate the linear fit, that's what you would expect to happen. and so I'm just going to be humble and say like, oh, most of the time I'm not going to be smarter than statistical models, statistical extrapolation of past trends that have been very consistent so I'm just going to be very humble despite what I might know all about research and what's happening, probably the most likely, the best prediction I can make is actually just follow that data, that extrapolation and see where it's going to take us and yeah, in that case, if you roll this out if you look at other benchmarks think we would have something like next year, maybe the models will be able to work on their own for a whole day worth of tasks.

5:45If you think of software, you might say like, oh, implement this entire feature, build out this entire set of the app. If you think of knowledge work, maybe do like a whole research report, this kind of scale. The reason I think why task length specifically is interesting is because that's what allows you to delegate more and more work to language models, to agents. Even if you have a very clever model, but if it needs feedback or interaction with you very often, then it really limits what you can delegate to it. If you need to talk to it every 10 minutes, right? Versus if you have something that can go for hours at a time, obviously, right, then you can not just have one copy of it.

6:24You can have a whole team that you delegate tasks to and manage them. And so I think that's why it's really critical that the models are actually smart enough, The agents are smart enough to work on their own, to correct their own errors, to iterate. Because that's really what allows you to delegate. Indeed, task lengths and time to complete as the metric for progress. So by mid-2026, you mentioned agents can work all day autonomously. Late 2026, at least one model matches industry experts across many occupations. And then by 2027, models frequently outperform experts on many tasks. So it's more time running and then generalization across the economy.

7:06And you mentioned GDPVAL, the OpenAI metric, as a benchmark to already see the progress towards multiple professions. Yeah, I think that GDPVAL is like a super cool evaluation from OpenAI, where they collected a lot of real world tasks from real domain experts to make sure it is actually representative of what you might do in the economy. and then they evaluated a lot of models on those tasks. They compared them against real experts' performance to give us a really good indication of how close, how far are we from having significant economic impact. So I think that's a super cool evaluation. The sort of obvious question is that GDB Val and Meter are carefully designed benchmarks.

7:52How do they predict production value once you add compliance, liability, messy data, the messy world, tool friction, and all the things. So I think like messiness and task length, like, you know, time duration, you're able to work independently are very similar or very correlated. So I think that's why it's interesting that Meter tries to measure how long can the model go on its own? Because, you know, if you think about like, you know, how do you come up with a task that, you know, takes a human eight hours, 16 hours, right? you will have to include all these messiness and all this real-world mess to even be able to measure it.

8:31But I think ultimately to go further, we really want benchmarks. We really want evaluations that come from the actual users, whether it's the industry, whether it's private users, because that's what ultimately matters, right? Is the model helpful to you? Do you get something out of it? Does your paperwork help you write something, fixes your code, helps you study right i think that's the real proof if you release a new model right do people start using it more do they really enjoy it is there anything that would change your mind any kind of signal whether that's real world adoption or benchmark performance is something that would make you more um cautious about um that exponential is anything that would change your mind i mean many things just i think like you know many of these things are like you know internal only right I might look at our model pre-training.

9:22I might look at our fine-tuning. I might look at our rel thing. How do new runs go compared to past runs? Do they match our expectations? Does scaling continue? Then I might look at more public things of like, are people actually able to use those models to be more productive, for example? At the beginning, there's always some adaptation period of, oh, you have a new tool like Cloud Code. It takes you some time to figure out how to use it. But then in the medium term, in the long term, do people keep using it? Are they getting more and more productive using it? I think that's one of the things I look at.

9:55Many, many signals. I think when you do RL, when you do research, I think you get very much in the habit of looking for signals to prove yourself wrong. Because you often have ideas that you get attached to, but that's not a good way to do research. Most of your ideas are not good and they're not going to work. So you really want to figure out as quickly as possible whether this idea is any good or whether it's actually wrong. So you really get into this habit of like, finding the fastest thing that will show that, oh no, this is actually not true. So by 2026, 2027, in your extrapolation framework, AI becomes as good as humans.

10:34A key question of the moment is, can it be, to which extent can it become better than humans? There's some chatter these days around Move 37 and whether AI can create those like alien new path to think and solve hard problems. So first of all, maybe remind the audience what Move 37 is. And then do you think that AI in its current state is going to be increasingly able to provide Move 37 type thinking? Yes. So I guess, yeah, to give background, Move 37, that was when we were building AlphaGo, an AI program to play the game of Go. So it was like in the year 2016, I think. And we were playing one of the best players in the world at the time.

11:26Because at that time, you know, no AI program, no computer program had ever beaten the top human players at Go. And it was considered to be, you know, one of the most difficult board games, sort of like, you know, a real test of intelligence. Move 37 happened during the second game of the five-game match, where AlphaGo played like a really unexpected, unconventional move that surprised many professional Go players. I think the commentator said it was, you know, really creative, unexpected. And then ultimately AlphaGo ended up winning that game. And so I think that was for many people an early sign that AI is not just purely calculating following an optimal path, but you can also do something that is truly novel and creative that you might not expect just from imitating its training data.

12:16Yeah, I think that's very relevant in the modern context as well, right? Because as you alluded to, there is a lot of discussion of, are LLM just parroting the training data? Can they actually do novel things? For me, as somebody who has been doing research a long time, I think it's pretty clear that these models can do novel things. And that's why they are so useful to many people. Whether it's writing code for you, because obviously you're not just writing code that you already have. That wouldn't be very interesting. Or helping you write a paper. The way those models are trained, they're literally trained to generate the whole probability distribution.

12:53which means that when we sample from them we can generate infinite amount of novel sequences from them. For the question of something like move 37 I think there it really comes down to is it something that is sufficiently creative and impressive that we can easily recognize it in the game of Go, that was pretty ideal conditions because it's very clean, very abstract you can really each move is very impactful so you can really see it clearly. I think to have the equivalent for our modern models, you need the combination of a task that is sufficiently difficult and interesting and a model that is both able to create sufficiently diverse and creative ideas and also able to evaluate accurately how good they are so that it can go down increasingly novel paths while making sure that this novel path is actually interesting and useful.

13:47Creating novel things is actually very easy with language models. The hard part is creating novel things that are useful and interesting. Extrapolating this further, there's the idea of creating novel science. So not just one move, but like a whole new idea, new concept. What's your current take on this? So I think AlphaCode and AlphaTensor proved that you can discover novel programs and algorithms. Very recently, I think last week, there was some news of Google DeepMind and Yale in the biomedical field coming up with brand new things as well. So do you think that's accelerating and that AI is in the process of discovering novel science?

14:35So I think we're absolutely at the stage where it is discovering novel things. And we're just moving up the scale of how impressive, how interesting are the things that it is able to discover on its own. And so I think it's highly likely that sometime next year, we're going to have some discoveries that people pretty unanonymously agree that this is super impressive. I think at the moment, we're more at a stage of, oh, it came up with something, but there's a beta button. But yeah, I'm not very worried because I see this process continuing. And then once it gets clear enough, there's less need to argue about it.

15:13How far do you think we are from an AI winning the Nobel Prize? Yeah, I think that's a really interesting question, right? Because we had a Nobel Prize for AI with AlphaFold, of course. And so I think the next very interesting point is going to be when can AI on its own make a breakthrough that is so interesting that it would win a Nobel Prize? I think my guess for that level of capability might be maybe 2027. I think we're probably not going to find out for quite some time afterwards because of the delay in getting prices. But I think by 2027, 2028, I think extremely likely that the models will be smart enough and capable enough to actually have that level of insight and that level of discovery.

16:01Amazing. Yeah, not just a Nobel Prize, right? It's like the Fields Medal for Math and all these kind of advances. I think that's what I'm truly excited about, actually. It's like AI that can help us advance science and really unlock both all the mysteries of the universe and all the improvements in living standards and abilities for us that we could have if we understood the world better. All right. So extrapolating this even further, then we get into the AI 2027 thing that you probably saw. So this general idea of if AI can create novel science, then AI can create AI researchers. And basically, AI can create itself, which effectively leads to a discontinuity moment.

16:52So I don't know if that in the blog post, that's the singularity or whatever. but does that strike you as somebody who's as deep in the field as possible as something that is possible in the short term or are they counterbalancing forces that makes that path to discontinuity harder as you get closer? Yeah, I think a true discontinuity is extremely unlikely from, you know, obviously AI researchers are already using AI to accelerate themselves and so what's already happening and what is likely to continue to happen is that we see a smooth improvement of productivity. And then the main open question is, how does the difficulty of improving AI keep scaling?

17:37Because a very common issue in many scientific fields is that we find all the easy problems first. And then as we continue exploring the field, it gets more and more difficult to make advances. So in my mind, the main question is, do these two trends balance each other out? so that the AI makes us increasingly more productive so that as it gets more difficult to make advances, we just sort of about stay on trend and we keep improving roughly linearly. Or is it still too difficult and eventually after some time we still see a slowdown? But it seems quite unlikely to me that we improve in productivity so much that we can actually accelerate.

18:20That would be very unlike any other scientific field. The normal course in many scientific fields is that we actually need to exponentially increase the research effort just to keep making progress and find new insights. For example, if you look at pharmacology, discovering new drugs, it's nowadays in the range of billions of dollars to discover a new drug versus maybe 100 years ago, a single scientist could discover the first antibiotic by accident. It's not that we will be surprised by a sudden takeoff in progress or we're just doing our research and suddenly our model is like 10x better. Maybe we'll be seeing advanced signs of, oh, we're making faster progress every single week.

19:02We can see something's happening. Maybe we decide to pause if we don't understand what's happening. Do you think the current approach to modern AI systems, which effectively is pre-training plus RL, does that take us to where we want to be, whether we call it AGI, SI, it's unclear what any of the things mean. But do you feel that this paradigm is the right one? Or do we need to come up with a different architecture altogether, post-transformers or otherwise? I think that's a great question. And I think it hugely depends on what you mean by where do we want to be? So I think if you're thinking of, oh, we want some kind of system that can perform at roughly human level in basically all tasks that we care about, productivity-wise, then I think it's extremely likely that the current approach, pre-training, RL, transformers, is going to get us there.

20:01If what you care about is, oh, we want to have a model of intelligence that is conscious in the same way we are, or more abstract qualities like this, I think that's maybe more uncertain, right? And I think this is where a lot of the confusion and disagreement comes from, as you alluded to, AGI, ASI, people talk about very different things and they have very different things in mind when they say, oh, the current paradigm is going to get there, it's not going to get there. I often like to not use the term AGI or ASI and just talk very concretely about what problem are we solving, what task are we solving, what quality are we interested in because I find that often makes the actual disagreement much more obvious.

20:42But yeah, I think if you're just thinking of in terms of like, is this going to help us be massively more productive? Is this going to massively accelerate scientific progress? Then I think definitely the current approach will get there. And given how extremely deep you are in RL, I cannot resist asking you sort of the trendy question du jour based on Richard Sutton's recent appearance on Dorkish podcast. do you think that the model of the future will be trained in RL from scratch and that actually having pre-training in addition to RL is the wrong way to go? Personally, I think that's unlikely.

21:28Not because pre-training is strictly necessary. I think we may well be able to train something completely from scratch as we've been able to do in other domains. but more because pre-training on this vast data sets that we have just brings us so much value that we would, from a practical point of view not want to give it up so we might well do some agents that are trained from scratch out of scientific interest it could be very interesting to learn about what would a non-human intelligence look like but from a programmatic point of view I I definitely think we would keep using pre-training data, not just from an efficiency point of view as well, but also I think there is interesting safety angles because by pre-training on all this human knowledge, we're implicitly creating an agent that has similar values as we do.

22:24And I think that is quite valuable for aligning a highly intelligent agent. If you already start out by caring about sort of the same rough set of values, that makes things much easier than you create an arbitrary alien intelligence that may have completely different values. Despite having done a bunch of from scratch RL in the past, I think I'm often quite pragmatic about this. I'd love to put a pin in that specific discussion about alignment and what we do to ensure safety to later in the conversation, because I think that's a super interesting vein. But maybe to switch tags for a minute, But I'd love to go into a little bit of your story and then the monumental body of work that you've done at Google DeepMind before joining Anthropic around AlphaGo, AlphaZero, MuZero.

23:21So maybe just the three, four minute version of your personal story from when you were a kid. What was the path that led you to become a world-class AI researcher? Yeah, actually, when I was a kid, I didn't have any expectations of becoming an AI researcher. I was always very interested in computers. I grew up in the Austrian countryside in a small village. So, you know, it's not like there was a huge amount of things happening. But computers were always very interesting to me. You know, it's like this connection to the wider world. So these other interesting things. and I was very interested in computer games as well.

23:59And I think that's the first time I became interested in programming because I wanted to make my own games, which I think is very common in people who get into programming. But I somehow, I always got distracted by the technical aspect of, I'm going to build a very general game engine that can run any kind of game. And so I never actually ended up making any game. I learned a lot about making game engines and different technologies. and that's how I ended up studying computer science eventually in Vienna. Yeah, there was like a classical computer science degree. And then by chance, after my first year and my first summer holidays, I had an internship with Google.

24:42And that's when I realized, oh, wow, these guys are doing really interesting things. That's where the big clusters, the tens of thousands of machines are. that's first time I radically changed my plans from wanting to stay in academia and you know I had originally thought oh maybe I do a PhD that's when I changed oh no actually I just want to join these guys at Google and I will finish my degree as quickly as possible and so that's actually yeah when I got my full-time position at Google finished my degree the next year and then moved to London so I was just working as a normal software engineer at Google working actually like an advertising which i wasn't super excited or interested in so like the technology was interesting right it's like these huge systems and google has you know famously great technology but actually after you know a year ish of this i was pretty done and bored of advertising and so i was actually planning to leave google and thinking of maybe drawing a headband going into finance when by chance I saw an email in my work inbox that this guy, Demis, was going to come to the office and give some talk about Atari and video games and AI.

25:55And it was actually a day off because I was visiting a friend somewhere else in England. But that email looked so intriguing that I was like, oh no, I'm going to have to take the train back to the office right now and see this talk. And yeah, I'm really glad that I saw this email and I did go back because that's the moment where I decided, oh no like no i'm not going to go into finance i'm going to move to deep mind i'm going to join these guys because this looks clearly you know super interesting super amazing they are doing really interesting research all right tell us a story of alpha go alpha go zero alpha zero mu zero what those are because it's it feels like it's fundamental ai knowledge that everybody who has an interest in the space should know about shouldn't understand the progression in in particular so starting with the beginning of AlphaGo you alluded to it a second ago but like what did it do what how was it trained and then how that how did that evolve with each version AlphaGo I think at that moment in time Go in the machine learning community was this really big target where everybody felt like oh you know it's this big unsolved challenge ImageNet had just happened before so clearly you know the models were starting to do something with images and being able to recognize them and predict them and if you look at the go board you know the right way it looks a lot like one of those images that you classify so there was a lot of momentum around using neural networks to somehow play go and then at the time david silver and actually one deep mind had been working on go seeing like both of them had been working on go for quite a while had published some very interesting papers.

27:42And that's when the idea of using Monte Carlo Tree Search with deep networks came together. So the idea was to train a deep null network to predict which moves you might want to play, whether you're winning or losing the game, and then use the tree search to really make a big plan of what are all the possibilities in the game. How would it go for you if you chose a certain move or a different move? How would the opponent respond? And to explain this in super plain English, the term search in this case is, as you said, it's tree search. It's not what people normally think of search, which is searching a corpus.

28:21This is searching a series of options, effectively. Is that the right way to think about it? Yes, it's quite literally what you might do when you play a game of chess, when you play any board game. it's quite literally thinking of you know what move am i going to do what move is my opponent going to do in return and then thinking about many possible moves like that and mapping out all the possibilities in the future so deep learning plus search what was alpha go trained on initial training phases of alpha go were on some human amateur games if i remember correctly so basically just predicting if you have humans playing many games of Go try to predict at each turn in the game what move would they have played and it turns out that if you train a deep network to do that you can get something pretty decent like amateur Go level but not good enough to actually beat a really strong player.

29:17And by the way just for the lore of it did you guys have any sense that AlphaGo was going to crush Lee Seedal, so the famous goal player that you mentioned earlier in the conversation. Was it obvious before? Was it a surprise? We thought we had a pretty good chance, but we were very nervous about, like, you know, are we going to win or are we not going to win or are we going to lose? Yeah, we actually had some bets beforehand of, like, how many games are we going to win or lose? I think it was very ambitious to put the match as early as we did. If we had wanted to be a bit more safe, we may have like tried to do a few months later and i think if we had done it a few months earlier we would have probably lost so it was very knife edge of i guess which also made it much more interesting for us right because it really means that each game is like a nail biter oh what's gonna happen are we gonna win are we gonna play you know dumb move what's gonna happen so that was very exciting alpha goes zero which was i believe the year after how was it different where was the progression main change between alpha go and alpha go zero was to remove all the human goal knowledge so instead of starting by imitating human goal games we were training it just from scratch playing only against itself and rediscovering basically all gone completely figuring out from scratch how to play did you give it the rules of the game We didn't give the rules of the game to the network per se, but we used the rules of the game to score the result.

30:55So basically, you know, it would play and we would tell it, you know, who won, who lost or, you know, cannot make this move. So the next hop was AlphaZero, which was a year or two later. Yes. How is that different? So AlphaZero, the idea was, well, obviously, it's a really beautiful game, but ultimately we would like to do something more general, right? Can we remove anything that is Go specific and verify that the algorithm can actually solve more problems? And in that case, we did that by trying to solve both chess, Go and Shogi, which is a Japanese chess, basically, with the same algorithm, same network structure, just by running it in different games and also making it much simpler, elegant, faster.

31:41So basically, that was really laying the groundwork for applying the algorithms to solve real problems. And then the next stop in the journey was MuZero. And just to bring it home for people, you were, I believe, second author on AlphaGo Zero, and you were the lead author on MuZero, which in the world of AI, I'm sure you're going to be very humble about it, but in the world of AI, it's as big a deal as it gets. So I'll say it so you don't have to say it. so mu zero what was the next um what was how was that different so the main motivation i had for making mu zero was that if you want to solve many real world tasks you have no way of perfectly simulating what's going to happen you know if you play a board game obviously you know if you make this move you know what's going to happen it's like the piece is going to go there it's going to take a piece whatever right but if you actually want to solve something like a robotics task or anything more complicated, it's impossible for you to simulate what's going to happen accurately.

Read the full transcript

32:43And also, we as a human, we don't do this, right? We just imagine in our head of, oh, I'm going to say this, then he's probably going to respond in that way. This meant that alpha zero, as it was, could not be applied to such problems because it required some way of, you know, simulating the game, scoring the outcomes. And the idea with mu zero was that, well we already have a deep neural network right these networks can learn a lot of things so why not let it why not teach it to predict the future of the environment the future of the world why not make the model be able to learn for itself what is going to happen after each action it takes after that you also applied this to code and math so that was alpha code and alpha tensor So zooming out a little bit, that evolution of reinforcement learning in games and then code and then math.

33:42What did you learn about the general power of search and learning that is today relevant in modern agentic AI systems? How did that whole body of work translate to what you are doing today? So games are a really good sandbox to learn very quickly about a lot of the reinforcement learning science. But, you know, the algorithms that work well, the kind of problems that we encounter, even from a technical point of view, how do we build a learning system that spans, you know, many data centers, uses tens of thousands of machines? Because games are very clean sandbox, very clean environments, so we can make many good experiments.

34:24And then now that we have a much more general model, right, the language models can do almost any task, but they're much more complicated and much slower to experiment with. We can apply those same lessons of, you know, we know how to build a really robust reinforcement learning infrastructure. And now we can build the same one for language models. You know, we know if you do this kind of RL, then the model will learn how to exploit the reward. and so we can apply the same lessons, the same mitigation techniques to the language models. If I understand correctly, I think MuZero had a learn world model.

35:06So basically it rehearsed the future for a better expression. So do modern LM agents have anything like that? Do they have an internal world model that lets them preview actions before they commit? So I think, yes, I would say that language models have not an explicit world model, but they do have an implicit model of the world. Because to be able to predict, you know, what is the next likely word in this sentence? How is this paragraph going to continue? They need to internally model, you know, what is the state of the world? that makes this person say that thing. And so it's actually somewhat similar to mu0 in the sense that mu0 also only had an implicit world model.

35:55It was never trained to predict what does the screen actually look like if you take an action. It was also only trained to implicitly predict if I take this action, what is the next action I should take? Or is it going to be good or bad for me? So in both those cases, you have an implicit representation of the world in your model that you can use to make predictions, but you're not actually reconstructing the full state of the world. Because reconstructing the full state of the world, that can be very expensive and complex. If you think about super high resolution video, audio signals, it's a very large amount of data that probably you don't actually need.

36:39If you think of human attention, we are only aware of a very small subset. of what's actually going on all around us all the time. Because that's the most relevant information that we actually need to make decisions. And that goes back to the prior discussion about pre-training. So the reason why pre-training and RL work well together is that you have that world model that's implicitly embedded into the corpus. Although the argument against it is that it's what humans think the world model is as embodied by language versus what the world model actually is. And that's my understanding of the debate.

37:24I mean, for the debate, I think different people have different points of view, so I don't want to speak for anybody. Yes, yes, yes. But yes, I think pre-training on this rich knowledge gives you some representation of the world already so that when you actually start to act and interact with the world, you can very quickly make meaningful decisions, meaningful actions. I like to think of it in a similar way if you look at many animals when they are born they very quickly know how to move how to run even if you look at gazelles for example in the savannah clearly they did not have time to really learn this from scratch a few minutes or hours in their cases they did not do pre-training but they have some evolutionary encoded structure in their brain because clearly it is very beneficial to have some sort of knowledge to make your learning more efficient yeah just rl in nature would uh would lead to not so good results like you're a gazelle and like you have to a b test whether to run towards the lion or away from the plane it's like the you know thousands of generations of gazelles acquired this knowledge over time yeah it was encoded in their genes and their brain structure in some way right and then you get to start on top of that.

38:41I think the main challenge or the main thing you need to watch out for is that you don't over-encode or you don't restrict your search space too much. If your pre-training, if your prior knowledge prevents you from exploring something that might be the correct course of action, that will be bad. So there is some danger there you have to be aware of. So this general idea of pre-training and RL work together in modern AI systems seems to be the big idea or topic of 2025, although, of course, I know it's been years in the making. Why did it take so long? It feels like RL, you know, progressed in its own direction and then pre-training worked in its own direction and those were slightly separate.

39:28Why did it take so long to put them together? Is it just purely practical and economic or anything else? scaling up the language models to the massive degree that we scaled them up took a lot of effort on its own and from a science point of view from an engineering point of view retraining and supervised training is more stable and sort of easier to debug because you don't have this feedback cycle you basically you know have a fixed target and we're trying to learn this target and so then you can focus on like you know is my training working and like is my infrastructure working and then you know this scales fall over versus if you compare to RL in RL you have this feedback cycle of oh I learned something and then I use that to generate my new training data and then I learn from that training data and now if you have you know something is not working it's very hard to figure out where in this cycle your problem is coming from you know maybe your training update was bad and that's why you suddenly started behaving badly.

40:31Or maybe the way you select actions to behave is not correct. And so you generate the bad training data. That's what messed up everything. So it's just much more complicated to get working correctly. And so I think it makes a lot of sense to first scale up the pre-training, the architectures, figure out something that works pretty well, especially if you can already get pretty far by some fine-tuning, some prompting. And then when it's clear that these models are really general, they are really useful, and we have them in a pretty stable state, then you can ramp up RL and take them even further.

41:08Even in our own work, if you look at AlphaGo, AlphaZero, we always follow the similar split as well, where we first set up the architecture of the network, the training using fixed supervised data. And only when we have that working really reliable, only then do we do the full RL loop and the full training. Just because debugging all of it at the same time, you're just setting yourself up for failure. It's really useful to be able to isolate the component and say, I have non-good data over here, I have a non-good target there. If the thing in between is not working, I can isolate it. And then we can isolate all parts of the system.

41:43How compute intensive is it to scale RL? and are there scaling laws for RL the same way you do in pre-training? There's less published literature about it. But I think if you look at all the RL literature over time, we see very similar returns on compute in pre-training and in RL where we can invest exponentially more compute in RL and keep getting benefits. There's going to be some interesting research to come to figure out what are the trade-offs between pre-training and RL compute. We know what should be the split for a big model, for example. It could be 50-50, it should be like 1-10, which way should it be 1-10?

42:31So I think that's going to be extremely interesting. But so far, yeah, we definitely see good returns on both. What's the latest state-of-the-art or thinking in the field of rewards? so in what you described for AlphaZero AlphaGo that was basically win-loss as a reward then it sort of feels like we went into kind of like fuzzy human matching this is good this is not good and now that we expand as per the above into more general fields where it's sort of unclear whether you win or lose how does that work? What parts of the evolution are you working on? Are you excited about? Personally, I don't work that much on reward modeling.

43:18I mostly work on sort of reasoning, planning, search compute, ways of making the model smarter by spending more computation. Yeah, thinking about rewards, I think the reinforcement learning process per se doesn't really care where the reward comes from. The algorithms are very happy to use any source of reward, whether it has a human feedback signal, It's like some automated signal from like, you know, winning, losing the game or passing a test. Whether it's something more model generated. For example, Anthropic, we have this paper about constitutional AI to, you know, have the model itself score whether you're following some guidelines.

44:02So it can be very flexible to what kind of reward you follow. The RL, VR, all the things, those are at this stage stuff that you see commonly used. Any thoughts? Yeah, I think we're seeing a huge mix of rewards and environments. And I think it's very much people are working very hard in figuring out what are the best reward sources and how do we scale it up and how do we get more rewards, more reliable rewards. That will be one of the key ingredients in scaling up RL further. And so switching from rewards, what is the latest thinking in terms of data, training data for RL? Again, following the evolution from like AlphaGo, it used to be human data and then like self-play.

44:55How does that work? Where does the data come from and what kind of data works best to train modern RL? Yeah, I guess a great thing about RL is that the data is generated by your model itself. So the smarter our models become, the better RL data we can generate, the more interesting and complex tasks they can solve, which then gives us more and more data that we can train on. Because the more complex a task, the longer it takes to solve the task, the more data it generates that we can then use for training. I think part of the challenge is to find tasks that are really representative of what people actually want to do with the model.

45:36Because now language models are so general, people are using them for so many different things. There's more and more of a challenge of, you know, we need to cover as many of those as possible in NARA to make sure that, you know, the model is actually able to do this diverse set of tasks. What matters more for training data? Is that quality? Is that quantity? Is that recency? I think that's a very interesting question that maybe doesn't have a super clear answer yet, or is maybe still interesting research to be done. I think we've seen papers arguing for different things, or we've seen different benefits, right?

46:11Clearly, we see improved training as we scale up the data, we can keep improving. But we've also seen very interesting fine-tuning results, papers published with a very small amount of examples. You can teach the model how to do an interesting skill. And I think we don't have any good scaling laws yet that tell us the trade-off, especially I think because it's very hard to measure what is the quality of a data point, right? Like how good is this example compared to this other example? Without being able to measure this, it's very hard to quantify the trade-off in any way. I think intuitively it's definitely true that if you have bad data, RL doesn't work that well.

46:51and if you have very high quality data it becomes much more stable for example I think that was like a it's very clear in like alpha zero days where you we spend alpha zero spends a lot of computation it does a lot of planning and search to decide which move to take and so that generates very high quality data to train on which then resulted in RL training that was incredibly stable so you can run it across continents take a long time to generate the data and then train on it. It's very robust, versus in modern RL with language models, the difference in how good is the model and what data it generates that we then train on is not so large because we more directly sample from the model and then train it, which then results in reinforcement learning that is less stable.

47:43As a one direction of scaling RL and making it more stable is by improving this, by, for example, putting more reasoning into your language model to generate much more high quality training data. That can then give us training that is much more stable and that we can scale up much more easily. I'd love to spend a little bit of time now on the general topic of RL and agents. So the famous agentic AI that everybody's been talking about, you know, breastlessly for the last year. so for people listening and you know as often in an effort to make this broadly accessible by you know a group of people in tech could you drive home the sort of intersection and overlap between RIL and agents does RIL power agents how does that work yeah so I guess like maybe first let's take a step on what do we actually mean by agent yes as compared to like a general language The second most debated question after a GIS is what is an agent?

48:47I guess, yeah, for our purposes, let's just say that an agent is an AI that can act on its own. Maybe take some actions on a computer, save some files, edit some files, send an email, whatever you want. But the main characteristic is that it doesn't have to interact with the user all the time. It can do things on its own. The reason why RL is very important for this actually connects back to pre-training because our pre-training data is not very agent-like. If you think of the pre-training data, there is websites and books and all kinds of recent text that has a lot of information, but it doesn't have a lot of actions.

49:33It doesn't really capture how do humans actually interact with the world. So if we take a raw pre-trained model, it's not a very good agent. Maybe you can prompt it a bit and sort of push it in the right direction, but it's not going to be very good at interacting and especially it's not going to be very good at correcting for its own errors because the pre-training data has no examples at all of how is our agent going to fail. And that's exactly where reinforcement learning comes in because in RL, we can take our agent, let it interact with the environment and then directly train on that interaction.

50:12So for example, if the agent did well we can reinforce those actions and if the agent did badly we can push it away from those actions and if the agent sort of did badly at the beginning but then recovered and managed to do well then we can also reinforce that recovery. And so that's super important because it allows the agent to actually learn from its own distribution of behavior. And that just makes it much more robust because now it doesn't have to generalize to something it has never seen before. It can actually learn on the actual problem that it's trying to solve. And that's why RL is really unlocking so much authentic capabilities now.

50:51If I'm an AI builder today, building an AI app, and I build it on top of Anthropic, Anthropic is whatever model is going to come with some of this sort of batteries included. But as a builder on top, do I need to do my own RL? There's this emerging space of like RL as a service where, you know, for this task or that task that I build on top of a general model that sort of like offers the ability to do RL or can I do a lot of damage just through prompt or like maybe like supervise fine tuning first? I think nowadays with the capabilities of top Anthropic Cloud models, top OpenAI GPT models, you don't need to do any fine tuning.

51:41You can take the model as is, ride your own tools, your own harness, and benefit from that Argentic training. Because doing good Argentic fine tuning is actually very hard. and so it's it's quite hard to do better than the top frontier models that you might get but on the contrary coming up with good tools and a good representation representation of your task makes a huge difference so like you know depending on how you express your problem for the model can make it way harder or way easier and so you can get a lot of mileage out of that What's currently missing to achieve the big dream of urgent AI?

52:24Is it model capabilities at the core or is it sort of like boring, quote, end of quote, engineering around reliability, tool use, safety? What needs to happen? I think we have sort of just basically improvements needed around the whole space. make the model better able to correct its own errors, make the model better able to continue going for long times without getting distracted, making the model just smarter in general, maybe making the model faster. There's basically a whole set of things that we know that we can improve. There's probably not one individual blocker, and that's why we will continue to see smooth incremental progress over model releases but sort of given how many things we know there are that we can do better on and improve yeah i'm quite excited about where models are going to end up i think that's actually you know one of the reasons why ai is a very fun field is that there are so many low-hanging fruits that you know you can do much better on but already the current models are so good that it's very fun to work on it it's like oh i can fix this thing it'll be even better versus you know if you're in a place where everything has already been solved and it's really hard to figure out how to make it better it's a very different story i spent a minute on on evals uh and um we we touch upon this a little bit but just to give it some some some proper space so there was um you know in your blog post that we talked to at the very beginning of this conversation this this this concept of external benchmark and then you quoted in your piece, Goodhart Law.

54:10So first of all, what is Goodhart Law? And then how should labs compare results so that it doesn't end up with this kind of leaderboard theater that we've seen a little bit in the last couple of years? Yeah, so Goodhart Law basically says that any measure that becomes a target stops being a good measure. And you can think of that intuitively that if you start paying, for example, programmers based on how many lines of code they write, while suddenly they will discover many ways to add more lines of comments, which is completely useless. And this is a very general effect that obviously, if you give people an incentive that they should optimize, they will try very hard.

54:54And we also see this with language model benchmarks. Of course, people want to get promoted, they want to launch their model. So any benchmark that is too easily measured or that has a lot of attention on it, people will optimize very hard for it, which means that probably the model will look very good at that benchmark, but if you then use it for your own task, you might get different performance. You asked about what do we do about this. It's very hard to prevent people from optimizing on the benchmark. One possibility is just periodically create completely new held-out benchmarks that nobody has seen before.

55:33and that gives you a fairly good estimate of model performance. I know, for example, a lot of researchers have their own toy problems that they use to test all the models precisely for that reason. So this is a set of problems that nobody has seen. You have a pretty good guess that it's going to give you an unbiased estimate. If you're an individual of your company trying to decide which model to use, it's probably something similar. Make your own internal benchmark that really represents what you care about and then measure on that. I think that's likely to be the most objective, most accurate way of measuring.

56:13Internally, what does that look like at a place like Anthropic or previously DeepMind? I know there are teams that are focused on evals. How do you think about what works, what doesn't in terms of internal evals? It definitely used to be easier to have good evals. Five years ago, the tasks we were doing, I think it was easier to measure model performance. I think nowadays it's much more difficult. And I think we try to not over-rely on evals so much because it's quite hard, for example, to measure how good is this model really at writing code. Yeah, I think it's one of the big unsolved or very important problems in the field of making really good evals that are both cheap to run, reliable, and accurate, because it's easy-ish to make an eval that ticks one of those, but to get all three is quite hard.

57:14For example, at the beginning, we were talking about OpenAI's GDP eval, no, GDP eval, and that one is very accurate and unbiased, but it's very expensive to run, because what it actually involves is taking human experts, having them do the task. And then compare the model task to the experts and rate it with multiple people. So it's very accurate, but it's extremely expensive to do. And related to that topic of evals, what's the latest in terms of our ability, or I should say your ability, to truly understand how models work, so the general field of mechanistic interpretability? You alluded to the fact earlier that RL, if I understood correctly, sometimes make it a bit harder because it does things occasionally in a more inscrutable way.

58:08My words, maybe not yours. So what is the latest? And indeed, does RL make things harder or easier? Oh, so what I meant before is that debugging RL in general, completely unrelated to interpretability, is harder because they're more moving parts. But it is also true that if you're not careful with RL, you can make interpretability harder. For example, one common thing with modern models is they do reasoning with the chain of thought. You could look at the chain of thoughts to see what are the model internal thoughts. And then you could also have a thought that, oh, maybe I should use that as a reward signal in RL and punish the model if it thinks the wrong thing.

58:47But then suddenly you completely destroyed your interpretability angle. So you sort of have to be careful that, yeah, you don't do RL on the signals that you actually want to use to interpret what the model is thinking of doing. That said, I think, yeah, there are some extremely exciting interpretability things happening, including mechanistic interpretability. I think actually last year, I think before John Anthropic, maybe even, there was a super cool golden gate cloud model where they found the neurons in cloud that were responsible for the golden gate concept. And then modified them to make a version of cloud that really loved the golden gate bridge in San Francisco.

59:24And so that's like a really vivid example of, oh, you know, we really understand what's happening in this model. And I can know what better way is there to verify that understanding than actually changing the behavior of the model. And so I think that's a super important direction for safety. As the models get smarter, we really need to be able to understand what is the model thinking internally? What is the values it has? Is it lying to us? Is it actually generally following the instructions? And so I think definitely an extremely important area to invest in and work in. Especially if people are interested in working AI or doing AI research, I think interpretability is a great area to get into.

1:00:03Yeah, perfect segue. For the last part of this conversation, I'd love to zoom out and talk about the impact of AI. So if we think that we are on the exponential and things are going to only accelerate from here, what does that mean? and certainly safety and alignment, which is a core value at Anthropic, hopefully in other parts of the field as well, but like Anthropic is particularly vocal about safety and alignment, let's say. How does that actually manifest? So we just talked about interpretability. What, for people who are concerned that this is going too fast and that we collectively are creating a monster, quote, end of quote, Can you give us a glimpse into the kind of work that is done for alignment and safety at a place like Anthropic?

1:00:56Yeah, I think the focus on safety alignment pervades all of Anthropic. And there's very rigorous processes where we train a model, whenever we want to release a model, both to analyze the capabilities of the model, verify the alignment of the model, ensure that it does not do harmful things on its own ensure that it does not enable malicious users to do harmful things and to the point where if we are unsure about the safety model we will delay the launch until we're sufficiently sure that it is actually harmless we will not launch and release a model which shows that people clearly take the safety much more seriously than any financial return or revenue.

1:01:49I think also in terms of research and resources, the teams working on safety and interpretability are a big focus of the company, which gives me a lot of confidence that we actually care about this and put a lot of effort into it. And at a more technical level, and to tie back an earlier part of the conversation around when we're discussing the pre-training and safety. So is safety and alignment an RRL problem? And by that, I mean the beauty of having pre-training is that you import that world model as we're discussing. but arguably you also import into your brain a lot of bad stuff if you collect data from the internet as we know there's good things but also a lot of toxic content so is alignment largely using rl to get rid of the bad stuff that is built into the pre-training we can definitely use rl to like shape the model behavior and ensure that for example given adversarial given bad input it sort of behaves safely or knows that it can refuse or is you know robust who attempts to prompt hack the model yeah I wouldn't view it alignment adjust like an REL problem I think it sort of it goes throughout the whole stack you might you know for example filter the pre-training data in some way you might after training you might have classifiers that you know look at the model monitor the model behavior to ensure that it is actually aligned you might when you write a system prompt for the model that you use you might put safety guidelines in there so i think safety alignment it really pervades the whole of research and the whole of you know product and deployment it's not just isolated into any one part and then another super interesting topic in the same vein of like the impact of ai is obviously the discussion around jobs so if as per the gdp of our discussion the agents are becoming just as good or better than humans obviously what does that mean for all of us in terms of our jobs what have you learned after the experience of alpha zero alpha go that that could give us a glimpse into what may happen once we all have super powerful agents do our jobs so i think the first thing that we didn't talk about yet so far is that artificial intelligence is quite, I mean, this may sound a bit simplistic, but it's quite different than human intelligence.

1:04:32So we can see that, right? The model may be much better at us on some tasks, like, you know, calculation, obviously, and like much worse than us at other tasks. So it is not, I don't think it is at all going to be any like one-for-one replacement. It's going to be much more complementary of a model is really good at something that maybe I really don't like doing or I'm not interested in or I'm very bad at. And then I'm much better than the model that's a mother part. And so I think it's going to be like a gradual process of we're all going to incrementally start using models more and more to improve our own productivity rather than have a model that one for one is able to do exactly the set of things we can do.

1:05:15For example, I use cloud all the time to, for example, refactor code or maybe write some front-end code that I don't want to write. At the same time, there's other parts where I'm clearly much better at putting than cloud still. So there is a synergy of use the best, most productive skills. I guess economists call it a comparative advantage. But there is this long process of will both improve our productivity incrementally. I think that process is going to give us some time of figure out politically and figure out economically, how do we want to benefit from this massive productivity increase?

1:05:55Even independently from AI, the promise of technology has long been that we're going to be all so productive, so wealthy, that we need to work much less. Yet mysteriously, we all have like 40 hours working week for decades. And so, you know, I think it's much more like a political, social problem of like figuring out how do we actually benefit from all these improvements and like, you know, bring the increases in wealth and productivity to everybody. And it's much less a technological problem. Which also means that we can't really solve it with technology. We have to solve it at like a sort of a democratic political level.

1:06:31How do we spread these benefits? Do you think that that increases inequality? quality so uh as you think about the impact of alpha go and and and mu zero what happened to the top go players and what happened to the top chess players did they did they disappear or did they get enhanced and better yeah i think at least in the case of chess and go there has been like more interest and it has become much easier for people to study how to play go how to play chess because now you don't need to find an expert tutor. Anybody can practice on their own, right? Spend a lot of time. I guess chess streamers are very popular on Twitch right here.

1:07:14And similarly, a lot of students are using language models to study. I think also for coding, right? Cloud code, these agents, they raise the bar of what anybody who has an idea can accomplish on their own. I think the larger picture, whether it increases or decreases inequality is quite hard to forecast. It both sort of raises the floor of what any person can accomplish, but it also gives very productive people an ability to be even more productive. It's possible that we see quite a difference between countries depending on the taxation, social redistributive system that they have in whether inequality increases or decreases, for example.

1:07:57Overall, I'm quite excited that it is very much non-zero-sum. It's very much, you know, increases the total wealth available in society. I think if you think about progress, if you think about prosperity, that is the most important thing. Like redistributing the pie is kind of a loser's game. To get more wealthy, we really need to grow the pie. You know, if you think of the agricultural revolution, the industrial revolution, the reason why we have much better lives nowadays is because we're so much more productive. We have so much more wealth. And so that's the key step we want to unlock. If you manage to make everybody in society 10 times more productive, what kind of abundance can we achieve?

1:08:44I think that's a good key question, right? What advances does that unlock in medicine? Curing diseases, halting aging, what does it unlock in terms of energy? Obviously, we have climate crisis. We need more energy to sustain our lifestyle. What advances in material science can we have? All of those are basically bottlenecked on how much intelligence we have access to and how can we apply it. So I'm incredibly optimistic about what will we be able to unlock in the next five years. I think we can go extremely far. Well, that feels like a wonderful place to live it. Thank you so much, Julian. This was absolutely fantastic.

1:09:30Thank you for spending time with us. Yeah, thank you for all the exciting questions and giving me the time. Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

From the publisher

Are we failing to understand the exponential, again?

My guest is Julian Schrittwieser (top AI researcher at Anthropic; previously Google DeepMind on AlphaGo Zero & MuZero). We unpack his viral post (“Failing to Understand the Exponential, again”) and what it looks like when task length doubles every 3–4 months—pointing to AI agents that can work a full day autonomously by 2026 and expert-level breadth by 2027. We talk about the original Move 37 moment and whether today’s AI models can spark alien insights in code, math, and science—including Julian’s timeline for when AI could produce Nobel-level breakthroughs.


We go deep on the recipe of the moment—pre-training + RL—why it took time to combine them, what “RL from scratch” gets right and wrong, and how implicit world models show up in LLM agents. Julian explains the current rewards frontier (human prefs, rubrics, RLVR, process rewards), what we know about compute & scaling for RL, and why most builders should start with tools + prompts before considering RL-as-a-service. We also cover evals & Goodhart’s law (e.g., GDP-Val vs real usage), the latest in mechanistic interpretability (think “Golden Gate Claude”), and how safety & alignment actually surface in Anthropic’s launch process.


Finally, we zoom out: what 10× knowledge-work productivity could unlock across medicine, energy, and materials, how jobs adapt (complementarity over 1-for-1 replacement), and why the near term is likely a smooth ramp—fast, but not a discontinuity.


Julian Schrittwieser

Blog - https://www.julian.ac

X/Twitter - https://x.com/mononofu

Viral post: Failing to understand the exponential, again (9/27/2025)


Anthropic

Website - https://www.anthropic.com

X/Twitter - https://x.com/anthropicai


Matt Turck (Managing Director)

Blog - https://www.mattturck.com

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://twitter.com/mattturck


FIRSTMARK

Website - https://firstmark.com

X/Twitter - https://twitter.com/FirstMarkCap


(00:00) Cold open — “We’re not seeing any slowdown.”

(00:32) Intro — who Julian is & what we cover

(01:09) The “exponential” from inside frontier labs

(04:46) 2026–2027: agents that work a full day; expert-level breadth

(08:58) Benchmarks vs reality: long-horizon work, GDP-Val, user value

(10:26) Move 37 — what actually happened and why it mattered

(13:55) Novel science: AlphaCode/AlphaTensor → when does AI earn a Nobel?

(16:25) Discontinuity vs smooth progress (and warning signs)

(19:08) Does pre-training + RL get us there? (AGI debates aside)

(20:55) Sutton’s “RL from scratch”? Julian’s take

(23:03) Julian’s path: Google → DeepMind → Anthropic

(26:45) AlphaGo (learn + search) in plain English

(30:16) AlphaGo Zero (no human data)

(31:00) AlphaZero (one algorithm: Go, chess, shogi)

(31:46) MuZero (planning with a learned world model)

(33:23) Lessons for today’s agents: search + learning at scale

(34:57) Do LLMs already have implicit world models?

(39:02) Why RL on LLMs took time (stability, feedback loops)

(41:43) Compute & scaling for RL — what we see so far

(42:35) Rewards frontier: human prefs, rubrics, RLVR, process rewards

(44:36) RL training data & the “flywheel” (and why quality matters)

(48:02) RL & Agents 101 — why RL unlocks robustness

(50:51) Should builders use RL-as-a-service? Or just tools + prompts?

(52:18) What’s missing for dependable agents (capability vs engineering)

(53:51) Evals & Goodhart — internal vs external benchmarks

(57:35) Mechanistic interpretability & “Golden Gate Claude”

(1:00:03) Safety & alignment at Anthropic — how it shows up in practice

(1:03:48) Jobs: human–AI complementarity (comparative advantage)

(1:06:33) Inequality, policy, and the case for 10× productivity → abundance

(1:09:24) Closing thoughts

More from The MAD Podcast with Matt Turck

All 44 episodes
Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)The MAD Podcast with Matt Turck · 1 h 10 min
Listen in VO