1004: Recursive Self-Improvement

26 Jun 2026 · 10 min · 5 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Recursive self-improvement (RSI): an AI system improving AI research to build better successors in a compounding loop. The episode argues current progress is mostly AI-assisted coding, not full RSI, because humans still set goals and evaluate results.

Guest backgrounds

No guests; host Jon Krohn leads the episode. Mentions include Anthropic co-founder Jack Clark, I.J. Goode (1965), Andre Karpathy (OpenAI/Tesla), and Max Tegmark.

Key claims

Anthropic reports 80%+ of merged production code written by Claude (May 2026), up from low single digits the prior year; engineers ship ~8x more code/day via AI typing. Task autonomy is improving: agent task time doubled ~every 7 months, then ~every 4 months. RSI risks include compute/data bottlenecks and “recursive drift,” plus loss of human control.

Notable examples

DeepMind AlphaEvolve scheduling workloads to recover ~0.75% of compute; faster matrix multiplication speeding training ~1%. Karpathy’s tool ran ~700 training-script experiments overnight, cutting GPT-quality training time 2.02h to 1.80h (~11%). Claude Cowork spreadsheet generation from Gmail/Sheets in minutes.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding RSI: Current State of AI Coding

0:45 to 1:48

Discussion on AI-assisted coding and the difference between current practices and true recursive self-improvement.

“was written by Claude, the company's AI model.”

Historical Context and Acceleration Trends

1:48 to 3:02

Exploring the historical context of RSI and trends in AI task completion times.

“What's happening today is AI-assisted coding, where human engineers still set the goals, organize the work, and judge the results.”

Examples of AI Advancements in Task Management

3:02 to 4:02

Real-world examples illustrating AI's capabilities in improving algorithms and task efficiency.

“In practice, Frontier models went from reliably handling software tasks that take humans a few minutes to tackling work that would take a skilled engineer most of a working day.”

Future Scenarios for AI Development

5:52 to 8:00

Discussion on possible future scenarios for AI capabilities and the implications of self-improvement.

“Back to the Anthropic report that I was talking about at the outset of this episode.”

The Balance of Optimism and Caution in AI

8:00 to 9:56

Exploring the balance between AI productivity gains and the need for oversight and safety.

“The physicist Max Tegmark likens racing towards self-improving AI without adequate safeguards to flooring the accelerator on a highway with your eyes closed, fine for a while, right up until it very much isn't.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:This is episode number 1004 on recursive self-improvement.

0:08Jon Krohn:Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. Today's topic is recursive self-improvement or RSI for short. This is the idea that an AI system could get good enough at AI research to build its own more capable successor, which then builds an even better one and so on and and so on, and so on, in a loop that compounds with every turn. The phrase RSI exploded across the AI world this month after Anthropic published a report on its own internal engineering. The headline number is striking. As of May 2026, more than 80 % of the code merged into Anthropic's production codebase was written by Claude, the company's AI model.

0:51Jon Krohn:They also sponsor this podcast and this episode, but they don't have any editorial input. I decided to do this topic myself and all the research was done independently. So yeah, if you hear a Claude ad, it has nothing to do with the substance in this episode. Anyway, before Claude's coding agent launched in early 2025, instead of the kind of 80 % figure that Anthropic is reporting in May of this year. The figure last year was in the low single digits. The typical Anthropic engineer is now shipping roughly eight times as much code per day as they were two years ago, not by working harder, but by directing an AI that does the typing.

1:32Jon Krohn:In one case this past April, Anthropic said Claude shipped more than 800 fixes that cut a whole class of API errors a thousandfold. Work the engineer overseeing it estimated would have taken a human for years. So is this RSI? Is this recursive self-improvement? Not quite. No. And the distinction matters. What's happening today is AI-assisted coding, where human engineers still set the goals, organize the work, and judge the results. True RSI would be when the AI manages the whole endeavor, designing the experiments, writing the code, evaluating what worked, and deciding what to try next with little or no human in loop.

2:09Jon Krohn:The term itself isn't new. The British mathematician I.J. Goode described an intelligence explosion back in 1965, reasoning that a sufficiently capable machine could redesign itself and rapidly leave human intelligence behind. The concern was later formalized by AI safety researchers, and the rise of large language models has dragged RSI out of science fiction and into quarterly engineering reports. The evidence that we're inching toward RSI is accumulating. Researchers at the think Tank Meter, whom I talk about all the time on this show, have been measuring how long a task an AI agent can complete on its own, defined by how long the same task takes a human professional.

2:49Jon Krohn:They found that this time horizon had been doubling roughly every seven months for the past six years, but over the past year, it appears to have sped up to something closer to every four months. That is staggering. In practice, Frontier models went from reliably handling software tasks that take humans a few minutes to tackling work that would take a skilled engineer most of a working day. Anthropic's own breakdown tells a similar story. On the hardest, most open ended problems, where even the definition of success is fuzzy, its model's success rate jumped from under 20 % in late 2025 to 76 % by May of this year.

3:28Jon Krohn:That's only a few months. two concrete examples bring this wild acceleration to life in 2025 one of google deep mind systems alpha evolve started designing algorithms on its own it found a better way to schedule workloads across google's data centers that recovered nearly three quarters of a percent of the company's worldwide computing power and it discovered a faster method for matrix multiplication that sped up the training of google's flagship model by about one percent those are small percentages is a quarter of 1 % and 1%, but those are very small percentages on enormous numbers. So the absolute savings are huge.

4:04Jon Krohn:And crucially, the AI was improving the very machinery used to build AI. The second example I have for you is even more pointed. Andre Karpathy, a founding member of OpenAI's research team and the former head of AI at Tesla, released a small tool that lets an AI agent autonomously tune a model training script overnight. He pointed it at code he had already carefully optimized himself. Over about two days, the agent ran roughly 700 experiments, kept around 20 real improvements, and cut the time to train a GPT-quality model from 2.02 hours down to 1.80 hours. That's an 11 % speedup on already excellent code, and the agent caught efficiencies that Carpathie himself had missed.

4:43Jon Krohn:As Carpathie put it, none of the individual tricks were especially novel, but they stacked up and he didn't have to touch a thing. On this podcast, I'm always going on about how Claude Code is mind-blowing, but now Claude Cowork is making my jaw drop as well. For example, I recently wanted to quantify how healthy my sales pipeline is for my AI consulting business. I simply asked Claude to estimate my sales for the coming quarter, and it brought info from relevant Google Sheets and my Gmail to create a professional spreadsheet of clients with estimated revenue for each one. Whoa, this might have taken me a day.

5:16Instead, it was done flawlessly with Claude Cowork in minutes. Claude Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. Ah, and you'll appreciate that I can ask Cowork to show me data, such as my sales spreadsheet, and it provides an interactive chart right in the conversation. For problems worth solving, get started with Claude at Claude.ai slash superdata.

5:45That's Claude.ai slash superdata. and check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash superdata.

5:56Jon Krohn:Back to the Anthropic report that I was talking about at the outset of this episode. In that report, they laid out three scenarios for where this goes from here. In the first scenario, AI stays below the level of the best human engineers, and humans remain firmly in charge. In the second, which Anthropic and I would consider most likely, In that second scenario, AI-assisted engineering keeps accelerating, but humans still steer model research and development. In the third scenario, AI becomes capable of improving itself. Anthropic co-founder Jack Clark has put a 60 % probability on an AI system being able to fit into that final self-improving bucket, being able to create its own successor with no human involvement at all by the end of 2028, so a couple of years from now.

6:41Jon Krohn:Now, plenty of smart people think that timeline is too aggressive and their reasons are worth taking seriously. The first bottleneck is compute. Even with efficiency gains, each new generation of models needs more computing power to train, so progress is chained to the pace of data center construction, and every chip serving a paying customer is a chip not available for open-ended research. The second bottleneck is data. AI has improved fastest where success is cheap to verify automatically. Code either runs or it doesn't. A math proof is either valid or it isn't, and that lets models safely learn from data they generate themselves.

7:17Jon Krohn:It's far murkier to check whether a model has gotten better at creative writing or legal judgment, and there's a real risk of what researchers call recursive drift, where small errors in a model's own output compound as it trains on itself. There's also a sharper critique, which is that some of the framing is marketing. A market leader calling for the world to have the option to slow down frontier AI development is conveniently also a market leader asking its competitors to ease huff the gas. Several prominent researchers have pointed out that the gap between today's agenda coding and actual RSI is wider than the excitement marketing suggests.

7:55Jon Krohn:But Anthropik's leadership appears sincere in its concern, and they're not alone. The physicist Max Tegmark likens racing towards self-improving AI without adequate safeguards to flooring the accelerator on a highway with your eyes closed, fine for a while, right up until it very much isn't. The worry isn't a Hollywood robot uprising so much as a quieter loss of control. A world where models are trained by models to pursue goals set by models, with safety verified only by other models, and humans gradually edged out of the decisions that matter. So where does that leave us? I'd land, as regular listeners will have come to expect, somewhere in the optimistic middle.

8:34Jon Krohn:The productivity gains here are real and already in your hands, the same coding agents accelerating anthropics engineers are available to you and me right now, and they make building useful things dramatically cheaper and faster than they were even a year ago. At the same time, the closer we get to systems that improve themselves, the more it pays to keep our eyes open, to build in monitoring, human checkpoints, and oversight while these tools are still firmly under our direction. RSI is increasingly in the vernacular of our industry, and it's certainly a threshold to keep our eyes on. If you're concerned about runaway AI systems and would like to do something about it, I highly recommend checking out episode number 1007 of this podcast when it comes out in a couple of weeks.

9:15Jon Krohn:It'll feature Ben Todd explaining the greatest AI threats facing society and how you can get yourself into a career combating these risks. If you can't wait until July 7th when that episode comes out, you can check out Ben's previous appearance on this show from five years ago. That's episode number 497. All right. That's the end of today's episode. If you enjoyed it or know someone who might consider sharing this episode with them, leave a review of the show on your favorite podcasting platform or on YouTube, tag me in a LinkedIn post with your thoughts. And if you aren't already, be sure to subscribe to the show.

9:50Jon Krohn:Most importantly, however, we hope you'll just keep on listening until next time. Keep on rocking it out there. And I'm looking forward to enjoying another round of the super data science podcast with you very soon. you

From the publisher

Could an AI get good enough at AI research to build its own, more capable successor and kick off a compounding loop? That’s recursive self-improvement (RSI) and it surged into the conversation after Anthropic revealed that, as of May 2026, Claude wrote more than 80% of the code merged into its production codebase. In this Five-Minute Friday, Jon Krohn separates today’s AI-assisted coding from true RSI, walks through the accelerating evidence - METR’s shrinking task “time horizon,” Google DeepMind’s AlphaEvolve, Andrej Karpathy’s overnight training-tuner, weighs Jack Clark’s 60% bet that AI builds its own successor by 2028 against the compute, data and “marketing” skeptics. As ever, Jon lands in the optimistic middle.

Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1004⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
1004: Recursive Self-ImprovementSuper Data Science: ML & AI Podcast with Jon Krohn · 10 min
Listen in VO