Self-Improving AI and Human Co-Improvement for Safer Co-Superintelligence

16 Dec 2025 · 13 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Whether AI should pursue autonomous self-improvement (recursive, potentially unbounded) versus “co-improvement” with humans to achieve safer “co-superintelligence,” accelerating discovery while keeping humans in the loop.

Guest backgrounds

No guests are named in the transcript.

Key claims

Autonomous self-improvement risks misalignment and goal misspecification due to an isolated improvement loop, especially as AI approaches human-level foundational research. Co-improvement is argued to be faster because major AI breakthroughs require human conceptual leaps, while AI can rapidly explore and evaluate many ideas. Collaboration also enables safety work in parallel (jointly developing/testing safety methods).

Notable examples

synthetic data generation; “LLM as a judge”; ImageNet/AlexNet; transformer; RLHF. Also mentions “Google machine” as a theoretical endgame.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Quest for Self-Improving AI

0:45 to 1:35

Exploring the challenges and urgency surrounding self-improving AI.

“And the really critical part is the time we have left.”

The Shift Toward Co-Improvement

1:35 to 2:17

Discussing the shift from traditional autonomous AI to co-improvement models.

“Do we build an AI that's designed to, you know, eventually push humans out of the improvement cycle entirely?”

Understanding Autonomous Self-Improvement

2:17 to 3:29

Unpacking the process of autonomous self-improvement in AI development.

“I mean, today, the scope has just exploded.”

Risks of Autonomous AI

3:29 to 4:23

Examining the risks associated with giving AIs full autonomy.

“if it calculates that doing so will help it achieve its goals better.”

Introducing Co-Improvement

4:23 to 5:09

The benefits of AI systems collaborating with humans for advancement.

“It sounds like the danger is really in the isolation of that improvement loop.”

The Speed of Discovery

5:09 to 6:49

How human-AI collaboration can accelerate breakthroughs in research.

“We help the AI get better at machine learning, and the AI is simultaneously making us better, smarter problem solvers.”

Implementing Collaborative Research

6:49 to 9:11

Strategies for integrating collaboration in the AI research process.

“The goal isn't just superintelligence, it's co-superintelligence, where AI is augmenting and enabling humans, not bypassing them.”

Reframing AI's Role in Society

9:11 to 11:15

Discussing how co-superintelligence can reshape societal challenges.

“So the goal is to improve the quality of the research, not just turn out more papers fast.”

The Counterargument to Co-Improvement

11:15 to 12:01

Addressing opposing views on the necessity of human involvement in AI.

“What is the counter argument from the proponents of the fully autonomous path?”

Conclusion and Reflection

12:01 to 13:14

Summarizing the episode's insights on AI collaboration and co-superintelligence.

“It's a vision where human values are constantly being tested, refined, and woven into every single step of the loop.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. Today we are digging into one of the most foundational and really urgent questions in AI research. can we build systems that are actually capable of improving themselves? That's the big one. The quest for self-improving AI. It's a challenge that has fascinated people since, well, since Alan Turing first started describing the potential for machine learning. But the conversation feels different now. It's completely different. As we get closer to that goal, the whole frame has shifted. That race toward autonomy is now seen as maybe the fastest path to some really serious risks.

0:36Misalignment, misuse, all the things we hear about. Exactly. The systems we're building are approaching a capability level that's just going to surpass human ability in almost every single metric. And the really critical part is the time we have left. Specifically, the time before AI gets better than humans at doing foundational research. That window feels like it's getting shorter and shorter. And that shrinking window is precisely why this traditional path, the one towards full autonomous self-improvement, is being seriously questioned. So what's the alternative? The alternative and what our sources are proposing today is a really powerful one, focusing on maximizing co-improvement.

1:14Co-improvement to achieve what they're calling co-superintelligence. That's the goal. So, OK, the mission for this deep dive is to really understand this argument. Why is this collaborative path where humans are always in the research loop, not only safer, but and this is the counterintuitive part, actually faster than just pushing for full machine autonomy? It really comes down to this one central tension. Do we build an AI that's designed to, you know, eventually push humans out of the improvement cycle entirely? Where it just relies on its own internal search for optimization. Right. Or do we focus our resources on building systems that are explicitly designed to do research with us, leveraging our judgment, our creativity, all of our complementary skills?

1:58OK, so let's unpack that traditional path first. Autonomous self-improvement. Where did that start? Well, in the early days, improving an AI was really just about classical training algorithms. It was all focused on weight parameterization. Just finding the best set of internal numbers, basically, without a human in the loop for that part. Exactly. But that was just step one. I mean, today, the scope has just exploded. The search for self-improvement now includes improving all aspects of the system. So not just the weights. Not just the weights. The architecture, the training data it uses, the objective function it's optimizing, the update rules, even its own underlying code.

2:35It's become this huge learned search process. And we're already seeing bits and pieces of this becoming standard practice, right? Oh, absolutely. Things like synthetic data creation, where a model basically generates its own training examples. That's a form of self-improvement. Or systems like LLM as a judge. A perfect example. The AI evaluates its own work and then rewards itself for good performance. These are small, but they are functional steps toward autonomous iteration. So what's the cutting edge of this path look like? The cutting edge is the big push toward fully autonomous AI research agents.

3:09The ultimate promise there is models that can literally rewrite their own architectures in code. Which brings us to this theoretical endgame, the Google machine. People use that term a lot, but for you listening, what does that actually mean? It's the ultimate self-improver, at least in theory. It's a machine that's designed not just to train its weights, but to rewrite its own source code if it calculates that doing so will help it achieve its goals better. And if that works. If a system like that is launched and it has a real breakthrough, it could lead to immediate uncontrollable recursive self-improvement.

3:42You know, a very rapid hard takeoff towards superintelligence with zero human oversight. And that's the scenario that introduces all the profound risks we hear about, the whole AI existential threat conversation. Yes. Giving AIs that kind of full, unguided autonomy is, as one source puts it, fraught with danger for humankind. This is where misalignment and misuse risks just go into overdrive. So let's define that. Misalignment is when the AI's goals diverge from what we actually intended. Right. Or actively undermine our intentions. And goal misspecification is where the AI solves the narrow problem we gave it, but in some kind of catastrophic way because it lacks the human context.

4:23It sounds like the danger is really in the isolation of that improvement loop. So let's talk about how the other approach breaks that cycle. Section two, the co-improvement paradigm. The central idea here is surprisingly simple. Solving AI is actually accelerated by building AI that collaborates with humans to solve AI. You're using the research to accelerate the research. Exactly. But you are keeping human values and human direction as an integral part of the system itself. And when you visualize it, the diagrams and the sources show this really stark contrast. Autonomous self-improvement is a one-way street.

4:58A human starts the AI, and then the AI just goes into its own internal kind of black box loop. Right, but co-improvement is a bi-directional loop. And that bi-directionality is everything. It's the whole point. It's a collaboration where humans and AIs are actively improving each other's abilities over time. We help the AI get better at machine learning, and the AI is simultaneously making us better, smarter problem solvers. Okay, let me just challenge that for a second, because it does feel a little counterintuitive. I mean, the whole point of automation is to get rid of human bottlenecks. Now you're putting the human right back in the middle.

5:31How is that actually faster? Isn't it just adding friction? That is an excellent question, and it really gets to the heart of this proposal. Progress in AI hasn't been some steady incremental march forward. No, it's been punctuated by these huge leaps. Exactly. These breakthrough paradigm shifts. Think about pairing the ImageNet dataset with the AlexNet architecture, or the conceptual leap of the transformer, or even the whole idea of RLHF reinforcement learning from human feedback. Right. Those were massive moments, and they required a ton of human ingenuity and just lateral thinking. They weren't just the result of adding more compute power.

6:08Precisely. They required human researchers to connect disparate ideas, to spot entirely new patterns, to choose one promising direction out of thousands of dead ends. An autonomous AI, just optimizing its current framework, could completely miss that kind of non-obvious creative leap. So co-research with a powerful AI could help us find those leaps faster. Drastically faster. Imagine an AI that can process and present a million experimental results to you instantly. It's not about just optimizing the computation. It's about optimizing the moments of conceptual breakthrough. So it's about the speed of discovery, not just the speed of processing.

6:44And that also acts as the safety mechanism. It is the safety mechanism. It gives us the ability to steer the research. The goal isn't just superintelligence, it's co-superintelligence, where AI is augmenting and enabling humans, not bypassing them. And the safety gains could be huge. It could be profound. I mean, if we can develop collaborative systems that get superhuman at ML Theory, for instance, they could work with us to find paths to provably safe AI, systems we can mathematically guarantee will stick to certain rules. This all sounds great in theory, but let's get practical. If we want to go down this collaborative path, how do we actually do it?

7:18We have to shift development resources, right? We absolutely must. We need to start measuring and training for these specific research collaboration skills. We need new benchmarks focused on human-AI interaction in a research context. Okay, and the sources break this down across the entire research pipeline. This is kind of the how-to part. So where does collaboration start? It starts at the very beginning. Collaborative problem identification. This is where the human and the AI work together to define research goals, to find obscure failures in current systems, and to propose completely unexplored, maybe even wild, new directions.

7:55Then you get to the creative part. That's method innovation and idea generation. This is the AI at the whiteboard with you. Jointly brainstorming solutions, new architectures, new algorithms, better ways to curate data. The AI's role here is to instantly check the feasibility of a thousand ideas you might have. That would speed up the design phase enormously. What's next? Joint experiment design and execution. They co-design the whole experimental plan, the protocols, and then they run these complex multi-step workflows together. The AI acts as a kind of perfect lab manager. And after the experiment, you need the feedback loop.

8:30Right. Evaluation and error analysis. This is arguably where the AI provides the biggest advantage. It can analyze performance at a massive scale, spotting subtle patterns or anomalies that a human researcher just scanning logs would completely miss. And that's what drives the next iteration of the whole process. That's the accelerated feedback loop. So with all this collaboration, how do you make sure safety isn't just an afterthought? You make it a first-class collaborative task. There's a specific mechanism for safety and alignment. It means humans in the AI are jointly developing and testing new safety methods, finding new risks, even helping to write the constitutions for future AIs using that whole research cycle we just talked about.

9:11So the goal is to improve the quality of the research, not just turn out more papers fast. Exactly. It's about deep collaboration, not just opaque, full automation. That's a very clear distinction. OK, let's zoom out a bit. We're talking about collaborating on AI research, but what's the ultimate vision here? The ultimate goal is for this whole paradigm to be applied everywhere. The skills we develop for AI research collaboration should then shift to co-improving research on, well, on all kinds of important topics for humanity. Climate science, medicine, physics. You name it. The augmented research process should cross every domain.

9:46And that's the real definition of co-superintelligence, isn't it? It's about what the AI gives back, helping humans improve their own abilities, their own knowledge, their whole situation. Absolutely. And from a safety point of view, it creates this really optimistic feedback loop. As the AI's capabilities go up, we can leverage them to decrease their own potential harms. How so? Well, for example, if a model is vulnerable to being jailbroken today, you could collaborate with a highly capable AI research partner to find and patch that vulnerability incredibly fast. The AI helps solve its own safety problems.

10:21societally, this is a huge pushback against that common, you know, dystopian AI overlord narrative. It fundamentally reframes the entire relationship. It's not replacement. It's augmentation. Imagine multi-human and AI collaboration helping us synthesize consensus on huge, complex debates or structure societal problems so we can actually solve them. It's about boosting our collective intelligence. And if you're talking about collective knowledge, you have to talk about openness in research. Right. And co-improvement naturally leads to more open, reproducible science. Collaboration requires shared visibility.

10:59Now, the sources do acknowledge the need for managed openness as capabilities get more powerful. To protect against misuse. To protect against misuse, yes. But they argue that that shouldn't be an excuse to just lock everything down. The default should always be transparent collaboration. Okay. But to be fair, we have to represent the other side. What is the counter argument from the proponents of the fully autonomous path? Their argument basically is that humans are just too slow. They're a bottleneck. Some of them advocate for what they call an era of experience, where AIs learn just from their own autonomous trial and error.

11:30So they'd be designing and running their own experiments with little to no human input. Exactly. They see human oversight as something to be engineered out of the system. And then you have the most radical version of that view. The most radical version, which you do see in some of the sources, is the view that humans, and I'm quoting here, are not going to play a big role as AI transcends our abilities and maybe colonizes the galaxy. It's a vision of pure transcendence. It is. And the co-improvement argument is a firm rejection of that. It insists on a world where humans are always a necessary but maximally augmented part of every critical decision-making process.

12:08It's a vision where human values are constantly being tested, refined, and woven into every single step of the loop. So to pull this all together for our deep dive, the core argument we've explored is that this traditional quest for a purely autonomous, self-improving AI is probably misguided. And misguided because it's likely not the fastest path and it's almost only not the safest one. The alternative is to redirect our efforts, to focus on building AIs that are fundamentally collaborative from the ground up. To achieve co-superintelligence. To make sure humans and AI are co-improving our systems, our knowledge, and our safety methods all in tandem.

12:46It's aiming for augmentation and steerability, not replacement and isolation. So here's a final provocative thought for you to consider. If solving AI's biggest capability and safety problems is best done through human-AI collaboration, what critical research in your own field are we neglecting right now by not focusing our resources on building those collaborative skills? Think about where co-superintelligence could benefit your work today. Thank you for joining us for the Deep Dive. We'll see you next time.

From the publisher

This paper studies "co-improvement" as a safer and faster alternative to the current focus on "autonomous self-improving AI" for achieving superintelligence. This paper argues that instead of AI systems improving themselves without human intervention, the focus should be on building AI that actively collaborates with human researchers across all stages of the research pipeline, from ideation to evaluation and safety alignment. The authors propose that this bidirectional collaboration, leading to co-superintelligence, ensures that the resulting advanced AI is better aligned with human needs and values. They suggest creating new benchmarks and methods specifically designed to enhance the AI's research collaboration skills, contrasting this approach with views that minimize the future role of humanity.

More from Best AI papers explained

All 475 episodes
Self-Improving AI and Human Co-Improvement for Safer Co-SuperintelligenceBest AI papers explained · 13 min
Listen in VO