In short
Claude Sonnet 4.5 release and Anthropic’s view of accelerating AI progress, including the “AI plateau myth,” the role of reinforcement learning (RL) and test-time compute, and why coding agents can run for ~30 hours via computer use, tool use, memory, and self-correction.
Guest
Sholto Douglas, leading AI researcher at Anthropic. Background: grew up in Australia; competitive fencing (top 50 in the world); early robotics/computer science research; read Gwern’s scaling hypothesis essay; built robotics foundation-model work with simulators and teleoperation data; received TPU support from Google; joined Google around the Gemini/ChatGPT era, where he helped build an LLM inference stack and led RL infrastructure for reasoning/reasoning-strike efforts; moved to Anthropic in February.
Key claims
progress is accelerating due to compute availability (“compute super cycle”) and a two-paradigm regime (pretraining + RL); Sonnet can be smarter than Opus because mid-tier models are cheaper to train and can be upgraded with RL; SWE-bench is near-saturated but Sonnet 4.5 is best-in-world for coding; 30-hour agents work by maintaining long-term coherency, memory, and taste/context to avoid losing the global plan.
Notable examples
SWE-bench jump to ~78; Cursor PMF with Sonnet 3.5; Cognition/Devon rebuilding architecture around Sonnet; 30-hour Slack-like app built by an agent; Anthropic cloud demo replicating Claude.ai with “artifacts.”
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Beginning of the Compute Super Cycle
0:00 to 0:26
Exploration of the current state of AI compute advancements.
“So finally, this year is where the compute super cycle is beginning properly.”
Pace of AI Model Releases
1:08 to 2:07
Discussion on the rapid release cycle of AI models at Anthropic.
“Congratulations on the release of Sonnet 4.5, which is the big news of this week.”
Understanding Model Categories
2:07 to 4:14
Sholto explains the distinctions between Opus, Sonnet, and Haiku models.
“I think it's also a reflection of the fact that this is now two-ish years after ChatGPT, ChatGPT, two and a half years after ChatGPT.”
Sholto's Journey to Anthropic
4:14 to 6:39
Sholto shares his personal journey and motivations leading to Anthropic.
“All right, so before we go into all of this in greater detail, I was curious about your story, your journey to Anthropic and then what you currently do at Anthropic, how you would describe your role.”
Influence of YouTube and Learning
6:39 to 8:37
Discussion on the impact of YouTube on skill development across generations.
“Or in other ways, like an introduction to, you can watch these people on YouTube and analyze what they are doing to become who they have been, who they are.”
The Importance of Independent Work
8:37 to 10:46
Sholto discusses how independent projects can signal research potential.
“I was at the time working on robotic manipulation stuff.”
Growth of Talent in AI Research
10:46 to 12:40
Insights into the growing pool of talent in AI research environments.
“I actually think that right now, a lot of the signals we look for aren't traditional PhD or anything like this.”
Early Days at Google
12:40 to 14:00
Sholto describes his initial experiences and challenges at Google.
“I started at Google, I think, like a month before ChatGPT or something like this.”
Building the Inference Stack
14:00 to 15:20
Learn about the challenges and strategies in developing a new inference stack.
“And so we had to notice that, design one from scratch.”
Transition to Anthropic
15:21 to 16:54
Discover the motivations behind the speaker's move to Anthropic.
“So the move to Anthropic was in February of this year.”
Show all 35 chapters
Comparing AI Research Cultures
16:55 to 18:28
Explore the differences in culture and focus among major AI research labs.
“Like, I think that DeepMind will directly contribute to more scientific discoveries from AI than anything else.”
Understanding Taste in AI
18:29 to 21:06
Delve into the concept of 'taste' in AI research and its implications.
“But before we do that, you mentioned a couple of times the word taste, which is one of those important words in 2025.”
Scaling and Mechanistic Understanding
21:07 to 24:34
Learn about the challenges of scaling in machine learning and the importance of understanding mechanisms.
“Methods that can take advantage of compute, and this is like in particular search and learning, will wash away all tweaks.”
Experimentation Culture in AI
24:35 to 27:20
Discuss the balance between experimentation and immediate results in AI.
“How often does an ID fail in a company like Anthropic?”
Focus on Coding at Anthropic
27:21 to 28:00
Understand the strategic reasons for Anthropic's emphasis on coding.
“I think this is another way in which Anthropik and DeepMind differ a little bit.”
Focus of Anthropic on Coding
28:00 to 28:50
Explore why Anthropic prioritizes coding in AI research and its economic impact.
“DeepMind has a much broader scientific culture because it has the resources to do so.”
The Importance of Coding for AI
28:50 to 31:00
Discusses how coding assists AI research and the advantages it provides in development.
“We're really focused on coding for two reasons.”
SWE Bench Benchmarking
31:00 to 34:20
Analyzes the SWE Bench benchmark and its significance in measuring coding progress.
“So there's this tractability, there's this replayability that doesn't exist in other fields that touch the real world in some ways.”
Evolution of Coding Models
34:20 to 36:30
Examines the evolution of coding models and their increasing complexity and capabilities.
“So let's roll back a year and we look at 3.5 Sonnet, which is the first really strong agentic coding model.”
Long-Term Operation of AI Agents
36:30 to 38:40
Discusses the breakthrough in AI agents' ability to operate over extended periods.
“cognition, I think, has always bet on a longer running, more independent, like, you know, agentic SWE, and maybe this is the moment that, like, really hits PMF for them, for example.”
Capabilities of Extended AI Tasks
38:40 to 41:00
Explores the tasks AI can now handle over longer durations and the implications.
“one being raw intelligence and the other one being how long an agent can operate.”
Future of AI in Software Development
41:00 to 42:00
Looks at the future potential of AI in developing complex software autonomously.
“What are some examples of tasks that you can do with 30 hours that you could not do with shorter runs?”
Progress in AI Autonomy
42:00 to 43:10
Explore how AI models are evolving to create complex websites autonomously.
“Which, for people who haven't seen it, like shows the progression of the models and how replicating the website went from basically...”
Memory and Context in AI
43:10 to 44:18
Learn about the importance of memory management and context understanding in AI models.
“Let's double click on the breakthrough part of this.”
Taste in Code and AI
44:18 to 46:41
Discuss the concept of 'taste' in programming and how AI can learn this skill.
“and they sort of lose themselves in the context.”
Advancements in Reinforcement Learning
46:41 to 48:01
Examine the role of reinforcement learning in improving language models.
“Yeah, I mean, I think it's important to recognize that it's not one individual breakthrough, really.”
Understanding the Role of Feedback in AI
48:01 to 49:34
Delve into how feedback mechanisms enhance an AI model's learning process.
“For those listening, a good way to understand at a high level of pre-training in RL, pre-training is like skim reading every textbook in existence.”
Test Time Compute and Reasoning
49:34 to 51:19
Discover how test time compute enables models to tackle complex problems.
“There's a whole bunch of things you can't otherwise know.”
Evolving Strategies in Reinforcement Learning
51:19 to 54:06
Analyze the transition and breakthroughs in applying reinforcement learning to LLMs.
“It's not like from a field that you really know, whatever heuristic that you've already done.”
Long-Term Coherency in AI Models
54:06 to 55:55
Learn about the significance of long-term coherency in language models and its implications.
“So there was this real phase shift in, oh, language models are smart enough, underlying priors, that they can solve sensibly difficult questions.”
Understanding AGI and Its Implications
56:00 to 59:58
Explore definitions of AGI and the potential cognitive capabilities of AI models.
“sentiment that the combination of ever more powerful LLMs plus RL gets us there with a side obvious question of what there actually means and what AGI means today Yeah, there's a few definitions that one could use.”
Debating AI Learning Efficiency
59:58 to 1:02:23
Discuss the efficiency of AI learning versus human learning and the debate around it.
“anywhere near as efficiently as humans do, right?”
The Future of AI: Benchmarks and Evaluations
1:02:23 to 1:05:28
Examine the importance of benchmarks in measuring AI capabilities and progress.
“But when I look at an LLM training pipeline, it is two and a half years of best effort, last minute, desperate effort.”
Preparing for a Transformed World
1:05:28 to 1:07:24
Learn how to adapt to the changes brought by AI in the digital and physical realms.
“and really make an effort in investing and figuring out whether we are on track for what I've been claiming we're on track for.”
Robotics and AI Feedback Loops
1:07:24 to 1:09:32
Discover the relationship between robotics, AI, and feedback mechanisms for improvement.
“But maybe things which we find are hard, like reasoning through mathematical problems, are easy.”
Transcript
Automatic transcript. May contain errors.0:00So finally, this year is where the compute super cycle is beginning properly. People have said that we're hitting a plateau every month for the last three years. I look at how these models are produced and every part of it could be improved so much. It is a primitive pipeline held together by duct tape and the best efforts and elbow grease and late nights. And there's just so much room to grow on every part of it. I think worth crying from the rooftops. Anything that we can measure seems to be improving really rapidly. Bet on the exponential. Hi, I'm Matt Turck from FirstMark. Welcome to a special episode of the Mad Podcast for the release of Claude Sonnet 4.5 this week with the incredible Sholto Douglas, a leading AI researcher at Anthropic.
0:37In this conversation, we go behind the scenes of how Sonnet 4.5 became the best coding model in the world and what happens when you enable AI agents to work for 30 hours straight. Beyond the launch, we talked a bunch about Frontier AI, how big AI labs operate, and how we are well on our way to AGI. At my request, Sholto made this conversation very approachable by breaking down a lot of key concepts, such as reinforcement learning, computer use, and AI benchmarks in plain English without the jargon. Please enjoy this great chat with Sholto. Sholto, welcome. How are you doing? Great to be here. Congratulations on the release of Sonnet 4.5, which is the big news of this week.
1:15I was just looking back as I was prepping for this, and I was struck by the pace of releases at Anthropic, in particular, Sonnet 3.7, which was like this huge deal. At the time, in my mind, if you had asked me, I would have said, oh, no, that was last year. But in fact, it was just in February of this year. What's the right way to think about that pace of releases? Is that a proxy for progress accelerating? Yeah, I think it's a proxy for a couple of things. One is that there's now this two-paradigm regime, where previously you did pre-training scaling and reinforcement learning scaling, and now we're in a mix of the two, basically.
1:55And so I think that gives you more opportunities to update models, because it means that you can make advancements along multiple frontiers, and then that means that you just end up shipping more frequently. I think it's also a reflection of the fact that this is now two-ish years after ChatGPT, ChatGPT, two and a half years after ChatGPT. And so the post-ChatGPT investment cycle is finally hitting, where like compute availability is increasing and all of this. And so it means that you should expect actually the pace of progress to be, because there's lead times in commissioning chips, basically.
2:32So even if you, as much as you wanted chips last year, it would have been impossible to get them because TSMC was booked out and so forth. So finally, this year is where the compute super cycle is beginning properly. In effect, yeah. Okay, great. Maybe for situational awareness for people listening to this, this Sonnet, this Opus, is this your Haiku somewhere? Maybe walk us through the differences between those models. Yeah, so we release models along three categories, three tiers. So Opus, which is the smartest model, Sonnet, which is the mid-tier model, and Haiku, which is the fastest, cheapest model.
3:12One of the interesting things about this most recent release is actually Sonnet is smarter than Opus. And this has happened before. In fact, this happens last year. It's a reflection of fast progress because it is cheaper to train mid-tier models than large models. And so what happens is that you end up doing a lot of progress on smaller models. Eventually, you need to choose when to scale up and sort of get the benefits of scale in a model, often you make progress fast enough that your mid-tier model is great anyway. And it's actually better than the large scale up model that you did previously.
3:52And I think this is also a little bit of a reflection of the reinforcement learning paradigm where you can take a model and you can train it. and extend it with reinforcement learning, basically. So that allows you to take a mid-tier model and make it as good as a larger-tier model of six months ago or three months ago. All right, so before we go into all of this in greater detail, I was curious about your story, your journey to Anthropic and then what you currently do at Anthropic, how you would describe your role. Yeah, so I think how far back do you want me to start? From the beginning, yes.
4:35From the beginning, yeah. So a couple of things. One is that growing up in Australia, there's a very traditional set of paths you can take. You can become a lawyer, you can become a doctor, or you can go into finance. Australia is like a wonderful country in so many ways. In particular, the quality of life is so high that it means that people just like choose these default paths, have a fantastic life. And I was very lucky in some ways. my mom was actually she was frustrated in her ambitions and so this meant that I had the perfect mentor throughout my entire life she studied medicine went on to do emergency medicine in South Africa but wasn't ever quite able to break into public health in the way that she wanted to do like she wanted to do like systemic change in public health and at the time that was just very difficult for one so instead I had her full attention so this is great But, you know, like growing up, when I did exchange in China, I got this dossier this thick of like China's political economy and different actors in the current startup ecosystem and this kind of stuff.
5:40So I had this like wonderful, constantly driving education and really supportive and wonderful way. I also was lucky enough to get into fencing. And through fencing, I had the experience of becoming one of the best in the world at something via repeated effort. I became the top 50 in the world at my best 43rd and it was partially a consequence of well I think in large part due to having a coach and perfect mentorship that was one of the best in the world he moved to Australia because his wife was Romanian he just coached he had just coached Italy to the gold medal in Olympics moved to Australia because his wife was Romanian she was facing discrimination in Italy.
6:25And so I had, on the one hand, perfect academic mentorship, and on the other hand, perfect athletic mentorship and a proving ground to, like, watch, you know, grow up watching these people on YouTube and then become one of the best in the world at something. Early introduction to reinforcement learning, like do this, don't do that. In some ways, yes. Or in other ways, like an introduction to, you can watch these people on YouTube and analyze what they are doing to become who they have been, who they are. and replicate that. Yeah. And you could be part of that world. All it just takes is intense amounts of effort.
7:00It's a thing that I find fascinating that across any field, the fundamental impact of YouTube and the fact that regardless of the field you look at, like every kid seems to be just much better than the prior generation. Right. I don't know if it's been studied, but at least that's your experience. Yeah. And I think we should see the same thing with AI, right? Like in the same respect, everyone will now get a perfect tutor. I then actually had that experience again with AI. fencing wasn't something I wanted to do ultra long term. I wanted to take a shot at the Olympics and then try and progress into like working in technology, basically.
7:31I was very lucky to read a Gwern essay on scaling where he basically details the scaling hypothesis. After reading that, I was like, oh my God, this is absolutely, clearly AGI progress over the next decade is going to be one of the most meaningful things to work on in the world. It's the largest level we have to have meaningfully advanced the world. and so I started doing my own research on nights and weekends and as like part of undergrad. How old were you then? So this is like last year of undergrad, like in the year after. And in undergrad you did? I did computer science and robotics. And I like sort of vaguely, I grew up like, you're looking up like Elon Musk and this kind of stuff.
8:15I wanted to like build rockets and Tesla, But I didn't have like a concrete idea of what actual problem I wanted to solve. Reading that essay was the critical hinge of, okay, AGI is possible this decade. It seems like the most meaningful thing in the world to work on. And I need to figure out how I can demonstrate that I should be working on this. I was at the time working on robotic manipulation stuff. And so I started working on like scaling up robotic manipulation. trying to train general foundation models for robotics from the bedroom, which is now a big thing. There's a lot of general foundation models for robotics companies.
8:53It was a little bit early then, but I rigged up my own simulator, collected a lot of teleoperation data, trained models, got a loan of TPUs from Google. Eventually, some people at Google noticed the work I was doing and said, hey, this is great work. Would you like to come and work with us? This is actually very fortuitous because, for example, I didn't get into the PhD programs that I wanted to. I applied to a couple of PhD programs here after undergrad and didn't get in. But I was very lucky that the work that I was doing really resonated with Google. And so they reached out. Which is a fascinating concept that at some point you could have had an academic roadblock, but still succeed to the extent that you're currently succeeding.
9:38You know, for people maybe outside of the AI research world, all this sort of feels like whoever is the smartest academically wins. Right. But does that suggest that being great academically and being a great anthropic researcher are two different things, so you need slightly different qualities? I think they're very highly correlated, but I think the signals that are usually used to gate academia are, like there are dramatically more people that satisfy the criteria of being really effective than there are that have the correct signals that would then enable them to progress to the next stage in academic career.
10:12For example, if you're here in the US, you end up doing as an undergrad research that can get you a NeurIPS or ICLR paper, whereas in Australia, it just isn't the case, right? I remember Peter Abiel actually once visited our lab in Australia and asked people to put their hands up if they were going to NeurIPS, and no one put their hands up, not even the PhD students. So it means you don't have, again, that mentorship aspect that is so important, and so you don't get a chance to develop problem taste on the things that mattered. And therefore, you don't have the correct signals that indicate you would have high potential for academia.
10:47I actually think that right now, a lot of the signals we look for aren't traditional PhD or anything like this. I mean, this is obviously very useful. But the fastest route or the most immediate one is whenever we see a really good blog post where people have done an incredible amount of work in an independent fashion, it's one of the highest signal things there is. One of the examples I love to use here is this guy called Simon Boehm, who's one of the leads on the performance team at Anthropic. And he's published, to date, the best guide on how to optimize a CUDA matmol on a GPU. It is simply the belt's best CUDA matmol guide.
11:26No one has done this retention, right? If someone did this retention, then I mean, we would reach out with a job interview offer the next day, right? And in fact, someone did it for TPU and for some kind of retention the other day. and we were like, this guy. Immediately, let's send out a request for an interview. So I think there is actually an absence of agency. There's an absence of taste. And there are still many ways to demonstrate this, usually by producing a world-class artifact in some independent fashion. Yeah. And a little bit to this conversation and the YouTube discussion, Do you see the pool of talent in AI research, whether academically sanctioned or sort of like more indie?
12:14Is that growing? Yeah, I think it's growing quite dramatically. I also think we have done a pretty good job of growing people. And I mean, I think like Anthropic has taken many, many junior people and grown them into really fantastic researchers and engineers in quite a deliberate way. So I think it's definitely been growing. So Google noticed you, and then what happened? So Google noticed me. I started at Google, I think, like a month before ChatGPT or something like this. So it was actually a fantastic time to start at Google because the entire company was suddenly forced to react instantaneously and compete with the Gemini program.
12:55um so uh it meant that there was this gap of the i guess like the typical uh command structures and everything were not well suited for that particular battle you know wasn't a pre-existing org you know gemini was sort of forged out of the foundation out of the merging of uh of brain and deep mind um it it meant that there was just a huge gap in terms of agency really of figuring out what we needed to do doing it as fast as possible organizing people together to to work on important things um and so i ended up uh one getting the chance to develop a lot of taste by working closely with people um in those like early months of gemini um but two also quickly got the opportunity to step up and to lead various parts of this.
13:50So one example of this is we just didn't have an inference stack that was at all sensible for the modern world of LLMs. And so we had to notice that, design one from scratch. A lot of the things you now see in the sort of SGLangs and stuff of the world are things that we had to derive from first principles at that point in time. and we wrote the inference stack. This ended up saving several hundred million dollars, I think, even over the first six months and meant that I was then trusted to solve both hard technical and sociopolitical problems. Because one of the interesting things about the inference stack as a problem was that it was both a really large technical challenge and a large sociopolitical one because the ownership of the preexisting stack was distributed across five or six different teams.
14:45And it meant that actually enacting change was quite hard. And so we had to... It was a challenge along multiple dimensions. That then led to me having the trust to solve problems of this form across other parts of the ML stack. And so, for example, later on, for when the thinking strike was started, the reasoning strike, I was in charge of research infrastructure for that to get us an RL code base that could actually allow us to do large scale RL and reasoning and this kind of stuff. Great. And then the move to Anthropic? And then the move to Anthropic. So the move to Anthropic was in February of this year.
15:27And I think it was motivated by a couple of reasons. The number one one is I'm just really excited by how deeply every single person in the company cares about how the future goes. And I think that's one thing that really struck me is everyone at Anthropic has an articulated theory of why what they're working on contributes towards a better future. Whether that is AI that is better in ways that can help people improve their lives or whether it's because it's AI that's more safe and controllable and aligned with our civilization's interests or even just more deeply understanding what's actually going on inside this AI and trying to better forecast the progress curves and where we think we're actually headed or policy.
16:13And I'm probably such a strong advocate for policy in many ways. It's fascinating as a thought. Again, like seen from the outside of the, you know, Big AI Research Labs, a little bit of a question is, you know, how different are those? It seems that everybody is incredibly smart. Everybody has access to the same resources. Directionally, more or less, people seem to be focusing on the same problems. and then you see, you know, one model comes out and it's better. And then, you know, a week later, there's another model that comes out from another lab and is better than the prior one. But from your experience, you see real differences in terms of like culture and goals.
16:53Yeah, I think there are, for example, DeepMind, if you wanted to solve science, is the best place in the world. Like, I think that DeepMind will directly contribute to more scientific discoveries from AI than anything else. Right? Like, absolutely. And I think it's just so well set up to do this across every aspect, right? You've got both the direct scientific efforts like AlphaFold and the material science work and this kind of thing. And also generally large efforts to do AI, like make AI scientists and all that. Whereas I think Anthropic has been laser focused on two things. One is like long-term AI alignment and two is near-term economic impact.
17:34so Anthropik has been laser focused on coding and computer use and things that we think will make a direct impact to the economy within the next six months you know one thing that Anthropik noticeably hasn't focused on compared to DeepMind and to OpenAI is mathematical reasoning DeepMind and OpenAI have been pursuing mathematical reasoning because of the implications for science and for scientific progress and because I think so many people there just love math so deeply and would love to see the field progress. We've had to reluctantly sacrifice a focus on that because we want to focus on, well, it's partly for many reasons, but we want to focus on near-term economic impact with the models and then much of our research along other dimensions.
18:29Let's double click on this in a minute. But before we do that, you mentioned a couple of times the word taste, which is one of those important words in 2025. What does taste mean when it comes to AI research? Yeah. I had a really interesting discussion about this with a biology friend. We were comparing taste across like biological research and animal. I think one of the most important things is mechanistically understanding exactly what you're trying to do and having an important simplicity regularize. When you think about taste in ML, it's often the crucial ingredient that allows you to decide what goes into your large training run when you have imperfect information.
19:15Because we can study very deeply what the impact of an architectural change is. But past a point, past a certain level of scale, you have to guess whether or not the impact of that change will compound with other ones, whether it will conflict. Because you can't test your full-scale run, right, end times. You only have one shot at that. And so a lot of taste comes from being able to make good inferences about do we think that this ultimately will sort of deliver increasing returns to scale. It also comes down to, do I think this direction of research is worth pursuing? Because often our baselines in ML are so well-tuned that it's very hard to beat them, even with what is theoretically a better method, because there are so many small tricks that are required to make a machine learning method work.
20:13And they can fail for any number of reasons. It's not like building a bridge where you actually have a pretty good idea why a particular shear was introduced. it can be all these quirks and so knowing whether it's right to push along that direction or to give it up and try something else is another question in taste and I think it always comes back to the simplicity regularization of people love to be clever we all do and that's sort of the bitter lesson I think is like maybe the best expression of this where you know generations of people have developed clever methods of encoding priors about how they think an artificial intelligence should reason and encoding it into the model and all of this gets wiped out by scale and uh and and through planning like basically like search and uh and learning um and scales applies to those two things yes and the the beta lesson being the richard sutton essay which uh everybody in ai uh knows about, but people may not, not everyone may have heard it, which is exactly what you described, this idea that generalization and compute will win over time.
21:24Yes, exactly. Methods that can take advantage of compute, and this is like in particular search and learning, will wash away all tweaks. I think I can offer a couple of examples of this in some ways to make it more concrete. one way that one of the reasons that convolutional neural networks were more effective is they encode a prior in convolutional neural networks in many ways you can think of it as like a little square like being you know drawn across an image so to speak that nearby pixels are related to each other this is very sensible prior right because if you throw a picture at an AI model and you don't tell at anything about the world, it has to learn that nearby pixels form curves that then form other things.
22:14And so there's this hierarchy of abstractions. But ultimately, that is not true of all images. And so convolutional neural networks will be better than a more general, like a vision transformer for the vast majority of images and up to a certain amount of scale. But past a point, actually, you need to be able to flexibly integrate information across the entire image. And a similar example in language might be, well, we know a lot about grammar. And so you might actually want to decompose a sentence into the constituent, you know, the verbs, the nouns, and how they relate to each other and so forth, and provide that explicit structure to your AI algorithm.
22:55But then what happens when you want the model to write poetry or to write code? All of a sudden, these assumptions have to be thrown away. And so you can't generalize across poetry and code and writing. And to the taste discussion, the art versus science part of this. So are you saying that at least in terms of anticipating how the training run may go, it's more intuition than actual numbers? So you can do actual numbers up to a point. The way to sort of illustrate or think about this is you are testing a system at multiple levels of scale. And actually the analog to biology was, if you think about it, you might test a new therapeutic in a cell and in mice and in model organisms.
23:45But that's no guarantee that it will work in a human, right? So you test across multiple different scales and multiple different model organisms. and it seems to work in basic single-cell bacteria, seems to work in mice, maybe works in monkeys, that gives you a lot of indication it's going to work in humans, but it's not a guarantee. So at that point, you need to understand the underlying mechanisms of how does this thing work, like what receptors is it binding to and so forth. In ML, it's exactly the same, right? You have your different model scales and you figure out, well, okay, it's delivering benefits of these model scales, and I think it should work, because mechanistically, I understand what this is doing to the learning dynamics of the model.
24:27And then you can have confidence that it's going to work. But if it's like, oh, it's a hack, we don't really understand how it works, and it's really complicated and introduces all this stuff in the code, then... How often does an ID fail in a company like Anthropic? Or in general? Yeah, or in general. I mean, I think a good example here is I once asked this question of Noam Shizia. And he was like, yeah, maybe like 10 % of my ideas work. And that's Noam, right? You know, one of the, you know, an absolute genius, one of the best in the field. So if only 10 % of his ideas work, then I think that establishes a bound on the percentage of ideas work.
25:08Most don't. And it's part of the success, again, of a place like Anthropico DeepMind to just encourage people to just experiment again and again. And then, I mean, those are very expensive runs, right? Yes. Just to say the obvious big reason behind the massive amounts of capital going into those companies is that the compute is expensive. And I'm curious about culturally the tension between, you know, you need to deliver because there's so much money at stake versus, no, you should have just a free open mind and just go for it. It's one of the things that at both Anthropic and at DeepMind we really tried to build, which is like a culture of safe experimentation where people were trusted to explore ideas for a long time out in the wild.
25:53Because you often need months of independent research to really prove out a novel research direction that is like a substantially different one. This is, it's hard. particularly I think it's hard actually less from the the compute cost of the experiments and more from a cost of time and focus because there are so many like remaining wins I guess and even in like the current architectures and paradigms and everything in such a way that well this maybe there's so many low-hanging fruit, right? Like a really high ROI use on your time would probably be to just go and look at the data and think hard about what the model is learning or doing and make some tweaks.
26:44You could even just... The simplest things in the world will still deliver massive gains. And so asking people or giving people the time and space to breathe and say, well, we know that there are short-term things you could be doing, but actually we want to try and develop a more general or fundamental technique that allows you to scalably do this in future is important. There's this tension between doing things that scale and doing things that don't scale. I think people at Anthropico are the places that are deeply researching completely different avenues, so non-transformers, non-RL. Yeah. I think this is another way in which Anthropik and DeepMind differ a little bit.
27:30Anthropik is a very focused bet. We think that, you know, AGI is within reach in the next couple of years. We think that, you know, it's the current paradigms or something. You're not crazy dissimilar to them. Maybe there's something new, but it's not like we think it's like some crazy out there research program, right? Really, for the last five or six years, Anthropics ethos has been scaling compute with broadly the current set of techniques. AGI is tractable within those bounds. DeepMind has a much broader scientific culture because it has the resources to do so. Anthropics has to be a focus bet.
28:12DeepMind has the time and space to be like, well, we're happy to bet on something that's really far outside the current paradigm. and I think depending on which kind of question you want to ask whether you think the really focused bet or the wide exploration of different and novel architectures is better that's like one of the sort of research ethos differences not to say that Gemini itself is a very focused bet but if you look Gemini, it's like a thousand people then there's still like 10 ,000 plus people doing all kinds of really long-term foundational research at DMI yep, got it So closing the loop on something that you mentioned earlier, why is Anthropic so focused on coding?
28:53Yeah. We're really focused on coding for two reasons. The first one is that we think it is the, how should I say? It's the thing that will allow us to assist ourselves in AI research faster. So there's this notion of automating AI research, right? and that work. We think that one of the most important signals of whether or not we are, basically the speed of takeoff, the speed of progress, is driven by how much AI is able to assist AI research. And so pre-fetching this is really important, we think. Secondly, we think it's the nearest term tractable problem domain in terms of economic impact. For Anthropik to be a viable research program that can research the things that we think are important requires economic return.
Read the full transcript
29:54And coding is a huge market full of people who are really, really, really keen early adopters who love trying and switching things, who are really excited to play with new tools. There's massive, massive demand. There's dramatically more demand for software in the world than there is good software. We've seen that in every previous iteration of compilers and general web abstractions and so forth. There's just a booming demand for software. And so, I mean, basically, the models are better at coding earlier than anything else because coding is a uniquely tractable problem in some respects for the techniques that we have.
30:40The data exists in many ways. you can containerize and run things in parallel, you can run unit tests, and so you can verify. Something that you know when it works and know when it doesn't work. Self-driving is uniquely hard, right? You need the car to work first time, kind of. Whereas coding, the models can fail 100 times. As long as it succeeds once, then that's fine. So there's this tractability, there's this replayability that doesn't exist in other fields that touch the real world in some ways. You don't want a lawyer arguing your case that is an AI, right? because if what if it gets the case wrong sorry so as techniques develop coding is uniquely tractable it is and you can see that right like already people are dramatically I myself am dramatically like higher productivity when I'm using the AI tools to write code and I have a friend who like manages nine clawed codes which is just like a crazy number I don't know how he does that but I can only handle two.
31:43So it's maybe like a skill issue on my behalf. All right. So Sonnet 4.5 is presented as the best coding agent in the world. So maybe unpack that for us, including performance on the SWE Bench benchmark. What are the numbers? What are the facts? And then we'll go into how that works. So SWE Bench is the current benchmark for how we measure coding progress. in the outside world, which all the companies use to evaluate against each other, it's an imperfect benchmark in many ways, right? It is like 50 % one particular web framework and this kind of thing. But what it does do is it takes real-world scenarios of work that people have done, just submitting a pull request, so a change to a code base.
32:33That's stuff that's on GitHub. Stuff that's on GitHub. Right. And it checks whether or not the model is able to do that same pull request and pass the same tests. And this ends up being a pretty decent proxy for a couple hours of work from a software engineer. These changes aren't incredibly complicated, but they're a reasonable complexity, a couple hours of work. We moved recently from roughly 72 to roughly 78 in SweetBench, which is a pretty substantial step up. I think it's worth pointing out that as recently as a year ago, I think we were under 20 % or something like that as a field. So there's been dramatic progress on the ability of models to do this unit of work that a software engineer does.
33:20I think that Sweebench is imperfect in a lot of ways. And it's probably pretty close to what we call saturated. One interesting thing to look at, like AI benchmarks, is you see these, they lose their utility past the point because they no longer disambiguate the differences between different models of high capability.
33:44But the models, one, are SOTA. They're the best in the world on SweetBench. It's a decent proxy. We're also, I think, more excited by the fact that a lot of our customers and partners are really excited by the model. So one example of this is the Cognition folks in Devon found the model so useful. They had to rebuild their architecture around it. Yeah, yeah. They had a great blog post on this. A great blog post on this, right? I think that's the real measure of whether or not a model is good, is whether or not it enables people to do things that they couldn't do before. And really coding as a whole has been transformed in this way over the last year.
34:20So let's roll back a year and we look at 3.5 Sonnet, which is the first really strong agentic coding model. The first model that you could ask to do something in front of you and it was able to interact with your code base in your computer and do it. In many ways, this model is what caused the PMF for Cursor. Cursor took off like a rocket with 3.5 Sonnet because they were in the right place and they were able to capitalize on that model as offering a coding experience that didn't previously exist. And then actually Cognition and Windsurf went for an even more ambitious target. So basically there's like a spectrum of agency here where either you can ask it to do 30 seconds of work or a couple of minutes of work.
35:08Windsurf was made as a company in part by betting more aggressively on the agentic abilities of 3.5 Sonnet. Then roll into this year. Which just as a quick aside is one of the key lessons for anybody in the startup world in 2025, which is bet on what the models will be able to do in six months from now. Right, exactly. Bet on the exponential. So what I think something that a lot of coding startups are now asking themselves is what can they now do with models that are capable of independently pursuing goals for substantially longer than previous goals? You know, before you had to supervise the models every 30 seconds.
35:49Over time, over the next couple of months, you're probably going to end up in a situation where you only need to supervise the models every, you know, 10 minutes, 20 minutes or so. That's a pretty dramatic change depending on the complexity of the task. even we have a couple of examples I think it was mentioned in the blog post where we asked it to build something that looks roughly like a chat app you know something like Slack or you know and it was the model just worked for 30 hours like it was just spinning there on a computer for 30 hours and came out with a really good working Slack like you know Teams like app it's pretty incredible yeah that is nowhere near built into any of like the existing products maybe maybe cognition is cognition, I think, has always bet on a longer running, more independent, like, you know, agentic SWE, and maybe this is the moment that, like, really hits PMF for them, for example.
36:40Yeah. Let's unpack the 30-hour aspect, which is fascinating. So, first of all, to just ground it for people, so this is a computer use... Just coding, I think. Just coding. So what does the agent do for 30 hours? Is it clicking on stuff? Yeah, it is there. It's reading files and it's writing code and running tests. So in exactly the same way that a human would.
37:13Basically, you can think of the model as running in a loop where it can constantly decide what to do. People often mention something called tool use. And tool use is the ability to, well, I mean, it's in the name, But in this case, it can use things like tools, like read file, write file, et cetera, or run code in the terminal. And it is sitting there in a terminal on a computer in a loop, just constantly looking at the current code, deciding, oh, well, it can't quite do this yet, so I'm going to work on that next. It's often making plans, particularly to run for 30 hours. One of the things that we're pretty happy with about the recent launches, we finally taught the models to use what's called memory.
37:55And we've built that into the agentic harness. So it's able to create a markdown file of to-dos and things that it thinks are important to do, check them off and work on them and check whether they've been completed. There's almost this like self-verification loop. One of the things that people were worried about with language models over like I think a year ago or so was that they would fall off track. Like they wouldn't be able to self-correct. and that this would basically ruin their utility. I think one of the things that may be remarkable about the current generation of agents is that they can self-correct.
38:30In fact, they're astonishingly good at self-correcting. And this, like, emergent ability has been pretty helpful. Yep, so much to unpack on this. So this, I think I heard you speak about two axes in the past, one being raw intelligence and the other one being how long an agent can operate. So is the fundamental breakthrough in very simple terms that if you can do it longer, you basically have a very smart AI that can just work longer? Yeah, exactly. If you can maintain long-term coherency, then the model is able to do things that it couldn't possibly have done. If I asked you to just, in a single stream of thought, write a Slack or Microsoft Teams working version of one of those.
39:20You wouldn't be able to do it, right? You have to sit there and take notes, do this closed loop feedback system. So long-term coherency is really important. And it's something that we think is just really critical for this. I think a good way of measuring this is to look at the meter evals. They're probably my favorite eval at the moment. And what this eval is, is they've taken a bunch of tasks, which humans do, particularly in the machine learning or programming context. And they've annotated how long it takes a human to achieve strong performance of those tasks. And then they ask AI models to do them.
39:59And what they found is that there's this really strong relationship between progress and the time horizon over which the AI is able to complete tasks. And so I think it's like every couple of months, the time horizon that the AIs are capable of doing is doubling or something crazy. Maybe every six months, the time horizon doubles, which is just, it's like, it's utterly insane. Yeah. Now, again, like all benchmarks, this one is imperfect, right? It only measures pretty simple tasks. It only measures, I think, 50 % success rate or something like this at the task, not like 99 % success rate. But it's a good directional measure.
40:48And it certainly resonates with my own experiences of, you know, as I've been using the recent models, I start to feel, if I just set everything up right, I feel like I could leave this overnight. And I could just churn away and it would probably have something pretty useful for me in the morning. Yeah. What are some examples of tasks that you can do with 30 hours that you could not do with shorter runs? Yeah. I think in this case, the Slack-like thing is a pretty good example, where it's a significant piece of software. It's really like an end-to-end working piece of software, which often takes a bit of time, like not an MVP demo.
41:27Other things I think are interesting is your machine learning experiments and stuff like this are pretty interesting. you want something that's able to propose an experiment and write a bit of code, run some initial tests, come back later, et cetera. Really, it opens up the world pretty dramatically. Basically, working software rather than demos, I think, is the key thing. Right. Fascinating. Now, I'm not saying that the models will spin you up a full working software right now, right? It's not going to spin you up a Slack competitor this weekend. Yeah, although the cloud AI demo that you guys produced was pretty impressive.
42:04Yes, right? I think we're seeing... Which, for people who haven't seen it, like shows the progression of the models and how replicating the website went from basically... Impossible. Yes, caricaturing, but like doing wireframes to now doing a fully functional website built autonomously by the AI. Yeah, and it got some pretty complex features. Like it got artifacts. Artifacts is a feature where the model is able to write code and then the results of that code are displayed in the web browser. And in this case, the model replicated Claude.ai with artifacts, with everything else. I can't quite remember how long that one took, maybe a couple hours to do.
42:49But basically regard this as the first halting steps of this. It kind of works. Sometimes it won't work. Sometimes it will. over the next six months, over the next year, expect dramatic progress here. And look at where we are now versus where we were a year ago. And the difference is I expect the same jump, basically. Let's double click on the breakthrough part of this. So, Sunnet 4.1, I think, was able to run up to seven hours. In this case, it's 30 hours, which I realize is not across all tasks, but that's the upper limit. You alluded to some of this memory evolution, this context, this ability to self-correct.
43:34Maybe explain in greater detail the advances that enable that jump to 30 hours. I mean, I think that the biggest things here are, or the question that we often ask ourselves is what is preventing the models from like working for a long go, basically? Or when do you need to intervene? And I quite like the model of interventions in a Tesla sense. as an example. Because right now you need to intervene quite frequently. But it's usually on questions of taste rather than it is questions of raw programming ability. It's not like the model is unable to, when it's decided to do the right thing, to do it.
44:12But sometimes the models take shortcuts and sometimes the models forget the overall structure of what they're doing and they sort of lose themselves in the context. They're doing a locally sensible change, but it doesn't actually make sense in the global context of what they're trying to achieve. And so I think a lot of the improvements, both that we've made and that are still to go, are on this taste and context, basically. It's on making the model better able to decide smart things about the overall structure of the program that it's going to do.
44:48And not take shortcuts and write sensible and good code. What about memory? Memory is also very important because, you know, the models do eventually run out of context. And so being able to manage memory over time and sort of like, I suppose even learning from experiences is something which would probably help this a lot. You don't want the model to be constantly rediscovering facts about how a particular system or code base works. And now, this is actually one of those areas where, like, the question of taste or, like, the bit of lesson comes up. Because you can imagine us going and launching a massive effort to teach the models coding taste.
45:31And that could be one way that you solve taste. You have heaps of human software engineers decide, well, no, this is good or this is bad or whatever. Where does taste come from in software engineering or what do we regard as taste? It's typically that it's able to like, it easily sets you up to make changes later on or so on and so forth. Or it's easy for maybe multiple agents to communicate with each other and collaborate. Like, often, good abstractions are something that, you know, you or I could work together on a code base and not conflict with each other, right? And so, there is this question of how much do you focus on teaching the model coding taste via getting software engineers to decide what is good or bad, you know, or should you be creating, like, a society of models that all, like, have to together code a giant monolithic code base and, like, you know, if they're arguing that it's bad.
46:20You can sort of imagine the spectrum of potential strategies and picking the right one there is like a difficult thing. Going back to the jump in performance from last year's models or even this year's models or even actually Sonnet 4.1 to like 4.5, again, to the point about the pace of progress accelerating. What were some of the breakthroughs? That I can't really talk about. Yeah, I mean, I think it's important to recognize that it's not one individual breakthrough, really. It is the continuous application of lots of different things across the entire stack for many people. And it's mostly just a function of compute in many ways.
47:09There are obviously individual breakthroughs, but fundamentally progress has been pretty smooth. Like on the meter eval, if you look at progress over the last two years, you can plot it with a straight line, right? And so similar to Moore's law of the past and this kind of thing, even Moore's law is made up of lots of individual improvements. It's not any one critical breakthrough. It's more the accumulation of a huge amount of work in an environment where there's a sort of exogenous force of compute pushing progress forward. Okay, so maybe let's talk about progress at a more abstract level, but grounded in 2025.
47:49So a big part of the discussion seems to have been the evolution from a focus on pre-training to RL, which we touched upon a couple of times. Talk about the impact of RL and why is RL such a big part of the conversation today? For those listening, a good way to understand at a high level of pre-training in RL, pre-training is like skim reading every textbook in existence. And RL is like doing the work to problems and getting feedback on whether you were wrong or right. And there are actually a lot of things that you can only learn via RL. And a good example of this is the skill to say, I don't know, in response to a question.
48:29because in pre-training, remember, you're modeling the, you know, you're trying to predict what text is going to come next in the, you know, all of these, you know, textbooks, the entire internet in the world. And so the only reason you would say, I don't know, you know, as a pre-trained model is if you think the character that you're modeling in the text would say, I don't know. Like if it's a likely completion, right? Not whether you, in fact, don't know, but whether you think that the sort of player that you've pulled from this cast of characters that you could model would say I don't know.
49:07Whereas in reinforcement learning, you could, in theory, set up a battery of tests where there are things the model knows and things the model doesn't know. And you could reward it for correctly answering things it should know and penalize it for falsely answering when it doesn't know. And what it will then learn to do is it will learn to look up information inside itself and assess its own confidence in whether it knows that information. So saying I don't know or solving hallucinations intrinsically requires reinforcement learning in many ways. So that's one example. There's a whole bunch of things you can't otherwise know.
49:48So I think also an important change in this sort of like era of reasoning models and RL on language models is at the end of last year, RL on language models finally started to work. And I think OpenAI deserves a lot of credit for releasing the first like serious RL plus LLMs release with O1. And I think this really kicked off a pretty substantial change because it opened up a new axis of scaling, right? There was pre-training scaling, and now there's test time compute and RL scaling. I think this is something which all of the research labs were investigating already. One of the reasons that DeepSeq was able to follow so fast was that they'd actually already released papers in the direction of doing RL on language models before, for example.
50:37and so this was it was already an idea in the air but opening I deserves a credit for that crystallizing it releasing it and detailing the first public existence of those scaling laws and maybe to continue making this super educational how do test time compute and RL overlap yes one way of thinking about this is test time compute is doing a lot of reasoning and then RL is the feedback signal on whether or not that reasoning was right or wrong. And so test time compute is a way of answering questions that are hard for you to answer. Let's say I ask you a question that you just know off the cuff of your, like, you know, off the back of your hand, basically.
51:24It's not like from a field that you really know, whatever heuristic that you've already done. You've already baked that in, like your muscle memory, so to speak. but for something which requires you to really think and really learn. Like when you're first doing math, if I ask you a basic times table right now, you can say that off like this. But if you're a kid, you have to like do out the math and all this kind of stuff. You need to like do the reasoning chain to learn it. And then you get feedback on whether it's right or wrong. So test time compute lets you do harder problems than you can currently do off the cuff.
51:58And RL then allows you to sort of distill that back into the model. It's almost like a ladder. You can like constantly do slightly harder problems because you're learning strategies to do harder and harder and harder problems. Reinforcement learning is not a new concept. So we were talking about Richard Sutton, who's been doing work in the field for decades and others as well. And then there was, you know, AlphaGo, like the whole, that old line of very successful, impressive RL-based successes. So why is it that in 2025, there seems to be a breakthrough to apply those to LLMs? Yeah. In some ways, it's quite funny.
52:38A lot of the... Okay, yeah. How do I say it? Let's take the DeepSeek paper, for example. In the DeepSeek paper, they detail one, an approach that works, and two, a lot of approaches that don't work. Actually, some of the approaches that didn't work were the approaches that led to AlphaGo's success. one of the craziest things about RL on language models in the RL from verified rewards regime is it's almost the simplest possible thing. It's almost too simple to work. And this is, again, comes back to that question of taste where really, I think a lot of people thought this was just too simple to work.
53:19And so they tried more complex methods that ultimately ended up being harder to get to work. And there may still be juice in those methods. But it was actually really important to nail the simple thing first. And so I think people were like almost too ambitious with the RL strategies that they tried initially. I think there's also a minimum bar in LLM quality that is required. Like you need the model to be able to solve meaningfully difficult coding and math problems before you can get that feedback loop of, well, you solve these ones right and you solve these ones wrong. right um and i think also one of like maybe the unintuitive things is those reasoning chains of tokens uh people for a long time thought that you'd need to do something clever to give the model long-term coherency you have to remember that two years ago 8000 tokens was in long context for a language model you know two and a half years ago 8000 tokens was long and now models are using 8 ,000 or 30 ,000 tokens to reason about something, right?
54:32So there was this real phase shift in, oh, language models are smart enough, underlying priors, that they can solve sensibly difficult questions. They're actually reasonably coherent at longer context than we thought they would be coherent. And that this ability to reason in long chains of tokens can emerge naturally with the right feedback signal. And this is a little bit counterintuitive, I think. Most people wouldn't have expected off the bat that the ability to reason would emerge naturally. There was a lot of thought that you'd have to structure it. You'd have to provide strategies for it to do reasoning.
55:12You'd have to build all these things, you know, prompted and hinted and this kind of stuff. And actually it turns out, well, no, you give it math questions, tell it whether it got them right or wrong, and the model will learn. But it comes down to a bit of lesson in scale and search is just allow the model to search, have enough compute to run the experiments, and the model actually ends up figuring out a really effective and sensible strategy. And that's what's happening now, right? The big labs are basically giving a lot more compute to RL. Yeah. Minimum base model quality, minimum amount of compute for RL, sort of trust in the ability for long-term coherency, doing the simple thing that works.
55:50they still sound obvious but they're actually like a little bit counterintuitive sometimes. You mentioned the word AGI earlier so is your personal sentiment that the combination of ever more powerful LLMs plus RL gets us there with a side obvious question of what there actually means and what AGI means today Yeah, there's a few definitions that one could use. I think a useful one is better than most humans at most computer-facing tasks. Because I think that's a really important moment for the world, where we go, okay, intellectual labor is addressable via this set of algorithms, and that totally changes the world.
56:38I think there are other definitions that are stronger that you could use. One of those is... Stronger. That was pretty strong. Yeah, sorry. I mean harder to meet, maybe. Yes, yes, yeah. Because you could have this and it could still not learn as effectively as humans, right? We learn and generalize from very few examples. We have incredibly high what's called sample efficiency, whereas AI models need hundreds of thousands of times more experience, hundreds of thousands of lifetimes, basically, to learn the things that we learn.
57:14And they can, over those thousands of lifetimes, they do learn the skills that we do to an incredibly high degree of accuracy. I think one of the important change over the last year has been that RL has finally meant that we have sort of this algorithm that allows us to take a feedback loop and turn it into a model that is at least as good as the best humans at a given thing in a narrow domain. And you're seeing that with mathematics and you're seeing that with competition code, which are the two domains most amenable to this, where rapidly the models are becoming incredibly competent competition mathematicians and competition coders, right?
57:57There's nothing intrinsically different about competition code and math. It's just that they're really amenable to RL. And any other domain, but importantly, they demonstrate there's no intellectual ceiling on the models, right? They're capable of doing really tough reasoning, given the right feedback loop. so we think that that same approach generalizes to basically all other domains of human intellectual endeavor where given the right feedback loop these models will get good enough that they are at least as good as the best humans at a given thing and then once you have something that is at least as good as the best humans at a thing you can just run it you know a thousand in parallel or a hundred times faster you have something that's actually even just with that condition substantially smarter than any given human.
58:48And this is completely throwing aside whether or not it's possible to make something that is smarter than a human. It seems entirely plausible, right? Like the brain is ultimately a biological computer. It seems possible to make a better one. But the implications of this are pretty staggering, right? Which is that in the next two or three years, given the right feedback loops, given the right compute, given the right elbow grease, and this kind of stuff. We think that we as the AI industry are all on track to create something that is at least as capable as most humans on most computer-facing tasks, possibly as good as, you know, many of our best scientists at their, you know, fields.
59:30This is really wild. It'll be sharp and spiky. You know, this will, like, there'll be examples of things it can't do and this kind of stuff, but the world will change. What do you make of the, you know, the counter thesis of the, again, Rich Sutton or Yann LeChan that seem to be saying that like a different approach is needed or RL only. Where do you make of that debate? Yeah. I think that it's true that our models don't learn anywhere near as efficiently as humans do, right?
1:00:04They take, you know, a thousand lifetimes to learn. but this is I think fine because they can live those thousand lifetimes whether in simulations or doing a job at a thousand firms and so on and so forth I think that the maybe I would disentangle there's two arguments one is like architecturally that transformers are like insufficient I don't think that's true I think we haven't yet really found anything that transformers haven't been able to model provided sufficient data and sufficient compute I think RL as an objective is a pretty powerful one Rich Sutton is actually a big fan of RL as an objective he just thinks we're actually encoding too many priors in with pre-training and this kind of thing it's not an adequate representation of the world I think so far the evidence indicates that our current methods haven't yet found a problem domain that is not tractable with sufficient effort and yeah so things that would make me eat my words is like if there was some domain that we put a lot of effort into that just didn't move like the goal the goalpost like so the like benchmarks just didn't move as and we just couldn't make any progress for a year then I would be like okay yeah there's some fundamental limitation here but instead what I just constantly see is every time we make a benchmark that measures something we care about, progress is incredibly rapid along that.
1:01:39Yeah. I think this is like worth crying from the rooftops a little bit of like, guys, anything that we can measure seems to be improving really rapidly. Where does that get us in two or three years? I can't say for certain. Yeah. But I think it's worth building into like, you know, respect to worldviews that there's a pretty serious chance that we get. something that is AGI. So you think people don't realize, you know, it's always interesting, right? Because, you know, reading stuff like online in the last, you know, three, four months, it's like this theme of the, you know, we've reached a plateau, but basically saying the opposite, right?
1:02:21We are in an exponential curve and many people don't realize that it's the case. Exactly. And I mean, people have said that we're hitting a plateau every month for the last three years um and if you look at what we've come over the last three years it's incredible uh i think that one other thing that makes me think god we're not anywhere close to a plateau is i look at how these models are produced um and every part of it could be improved so much like it is a primitive pipeline held together by duct tape and the best efforts and elbow grease and late nights and like god it's actually i remember um uh i don't know if this is a good analogy or whatever but i remember i went sailing with a couple of friends a few months ago and the boat was so well designed it was just like clearly the product of you know like millennia or like you know centuries of like accumulated human design and effort and i was like wow like this is this is what it feels like to be in a uh sort of the the accumulation of a lot of human effort right it's Actually, it's pretty hard to beat today's best sailboat designs.
1:03:27But when I look at an LLM training pipeline, it is two and a half years of best effort, last minute, desperate effort. And there's just so much room to grow on every part of it. So first of all, Sonnet 4.5, which is described as the best coding model in the world, also seems to be performing across a lot of different other domains like economics, research, and finance. So it's already just to verbalize that. One of the things I was really excited by, actually, was that there was that GDP eval that OpenAI released. And I'm not sure, Sonnet 4.5 was only just released, so it's not on there. But 4.1 Opus was the leading model there.
1:04:14And I think that's a really interesting and really good eval because it demonstrates such a breadth of tasks across all parts of the economy. It's an eval that's across the various sectors of the economy, right? So manufacturing, and basically they took a bunch of experts to describe what success looks like. And now the models are going to be able to be measured not across just coding or some limited tasks, but across everything. Is that a fair way of describing it? I've wanted someone to do this for a long time. Take the Bureau of Labor Statistics. I think the most important input into policy would be take the Bureau of Labor Statistics, take all the jobs there, break them down to tasks and see whether the AI models are able to do that and measure progress over time, right?
1:04:53And this is obviously going to be in perfect measure. We'll probably reach like better than human on the GDP eval and it won't change anything economically because it'll be all the connective tissue and all the like, you know, the context and actually like the task won't be representative. But again, we'll then find better ways to measure these difficulties and we'll keep pushing benchmarks. I've wanted someone to do this for a long time. I'm really glad they did it. I'm really glad that our models were general and generally strong and sort of straight up top across all the areas. And I think policymakers should really look at this and extend it and really make an effort in investing and figuring out whether we are on track for what I've been claiming we're on track for.
1:05:36We can measure this. Yes. And we should be. Yes. So, yeah, just to close on that last theme. so awesome, all very exciting. What do we all do? How do we prepare for this world that seems to be around the corner? Yeah. I think the most actionable piece of advice is keep planning for a world where you as an individual have more leverage, right? Right now, I can use two coding agents to do twice the work that I could have done before. If coding agents progress in the way I've been saying in a year or two, you'll be able to manage a team, basically, that works 24-7 for you doing work. I think we should expect in the digital domain for individuals to get dramatically more leverage over the next couple of years.
1:06:29I think then many incredibly important problems, you're going to track, our world is so imperfect in so many ways. people still live in dramatic poverty health and medicine is unsolved housing is completely unsolved the world could be a million times better in so many different ways and what I hope is that people take initially models giving us leverage over the digital world and then hopefully models giving us leverage over the physical one through robotics to dramatically improve it Is that happening? Robotics is another thing that seems to be one of the key themes But on the other hand, to use actually the word hand, it seems like people are just still struggling to make hands move the way.
1:07:15So, like, the physics of it seems to be the limiting factor. Yeah, there's this thing called Moravec's paradox, right, which is that things which we find really easy, like manipulation, picking up objects, are really hard for AI. But maybe things which we find are hard, like reasoning through mathematical problems, are easy. I actually think Moravec's paradox is a little bit fake. and I think this is mostly a question of data availability and like, you know, RL signal and stuff. And I think one interesting, one reason to look at this, if you look at like robotic locomotion, so, you know, the ability of robots to walk around and balance and stuff and look at the videos of the unitary robots, this difference now versus two years ago is crazy.
1:07:56These things are incredibly agile. Like there's this video, I think someone kicking one over and it literally does a matrix kind of like get back up thing. It's crazy. This is because locomotion is a really easy RL signal. And right now you can pretty much, like locomotion is kind of solved, to be honest, with basic RL. Manipulation is a bit harder. But there's a few things that make me think that robotics is going to work. For starters, I've seen incredible progress from the robotics labs over this year. Really, they've gotten to the point where they can do pretty interesting, you know, basic physical tasks.
1:08:31Two is the existence of a large generator verifier gap, which is that one of the things that makes improving our models hard is we constantly need to find people who can beat the models at the things we want to improve them on. But with robotics, we're making really smart general models. So you can actually have these as teachers or judges for whether or not the robot is doing the right thing. If I say stack the red block on top of the blue block, we could then ask the language model, did it stack the blocks appropriately? If so, give it a reward. If not, don't. So you can use the generator verifier gap to give models feedback.
1:09:10And finally, for a long time in robotics, people thought they would have to solve long-term coherency in planning. And that's also something that language models have made easier. They can break things down into multiple steps. So all the robotics labs are focused really hard on making great motor policies, and they're making incredible progress. It's mostly just a data and feedback loop question. All right, it's been fascinating. I can think of another 40 questions that I would want to ask you right now, but you've been incredibly generous with your time. Thank you so much. This was terrific.
1:09:40Really appreciate it. It was a real pleasure. Thank you very much. Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
From the publisher
Sholto Douglas, a top AI researcher at Anthropic, discusses the breakthroughs behind Claude Sonnet 4.5—the world's leading coding model—and why we might be just 2-3 years from AI matching human-level performance on most computer-facing tasks.
You'll discover why RL on language models suddenly started working in 2024, how agents maintain coherency across 30-hour coding sessions through self-correction and memory systems, and why the "bitter lesson" of scale keeps proving clever priors wrong.
Sholto shares his path from top-50 world fencer to Google's Gemini team to Anthropic, explaining why great blog posts sometimes matter more than PhDs in AI research. He discusses the culture at big AI labs and why Anthropic is laser-focused on coding (it's the fastest path to both economic impact and AI-assisted AI research). Sholto also discusses how the training pipeline is still "held together by duct tape" with massive room to improve, and why every benchmark created shows continuous rapid progress with no plateau in sight.
Bold predictions: individuals will soon manage teams of AI agents working 24/7, robotics is about to experience coding-level breakthroughs, and policymakers should urgently track AI progress on real economic tasks. A clear-eyed look at where AI stands today and where it's headed in the next few years.
Anthropic
Website - https://www.anthropic.com
Twitter - https://x.com/AnthropicAI
Sholto Douglas
LinkedIn - https://www.linkedin.com/in/sholto
Twitter - https://x.com/_sholtodouglas
FIRSTMARK
Website - https://firstmark.com
Twitter - https://twitter.com/FirstMarkCap
Matt Turck (Managing Director)
LinkedIn - https://www.linkedin.com/in/turck/
Twitter - https://twitter.com/mattturck
(00:00) Intro
(01:09) The Rapid Pace of AI Releases at Anthropic
(02:49) Understanding Opus, Sonnet, and Haiku Model Tiers
(04:14) Shelto's Journey: From Australian Fencer to AI Researcher
(12:01) The Growing Pool of AI Talent
(16:16) Breaking Into AI Research Without Traditional Credentials
(18:29) What "Taste" Means in AI Research
(23:05) Moving to Google and Building Gemini's Inference Stack
(25:08) How Anthropic Differs from Other AI Labs
(31:46) Why Anthropic Is Laser-Focused on Coding
(36:40) Inside a 30-Hour Autonomous Coding Session
(38:41) Examples of What AI Can Build in 30 Hours
(43:13) The Breakthroughs That Enabled 30-Hour Runs
(46:28) What's Actually Driving the Performance Gains
(47:42) Pre-Training vs. Reinforcement Learning Explained
(52:11) Test-Time Compute and the New Scaling Paradigm
(55:55) Why RL on LLMs Finally Started Working
(59:38) Are We on Track to AGI?
(01:02:05) Why the "Plateau" Narrative Is Wrong
(01:03:41) Sonnet's Performance Across Economic Sectors
(01:05:47) Preparing for a World of 10–100x Individual Leverage
