Welcome to the Era of Experience

14 Oct 2025 · 17 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues AI is shifting from learning from human data (LLMs trained on online text) toward an “era of experience,” where agents generate their own training data through long-term interaction, grounded actions, and objective feedback to enable new scientific discoveries beyond human knowledge.

Guest backgrounds

No guests are mentioned; it’s a host-led “A Deep Dive” discussion.

Key claims

Human-data progress is hitting a ceiling due to limited/expensive high-quality data. Experience-based agents learn via self-generated experiments, maintain long-term context, act in environments beyond text, use grounded rewards from measurable outcomes, and reason with internal world models rather than mimicking human thought. Risks include job displacement, safety/control challenges, and reduced interpretability; potential benefits include robustness and “correctable misalignment” via two-level goals (human “why,” learned “how”).

Notable examples

AlphaProof—trained on ~100,000 human proofs, then generated ~100 million additional proofs via reinforcement learning within formal math rules, with correctness enforced by the system’s objective constraints.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Shift in AI Paradigms

0:45 to 2:15

Exploration of the transition from simulation to human data in AI.

“Reinforcement learning agents getting superhuman in these closed tasks.”

Hitting the Ceiling

2:15 to 4:05

Discussion on the limitations of the current human data era in AI.

“We need AI to generate its own insights.”

Introduction to the Era of Experience

4:05 to 5:35

Understanding the era of experience and its core principles.

“That's the power of grounding it in the environment's rules.”

AlphaProof: A Bridge Case

5:35 to 8:05

Examining AlphaProof's achievement in mathematics and its implications.

“So they maintain context and memory over a long period.”

Shifts in AI Characteristics

8:05 to 10:51

Detailed look at four key characteristics defining the era of experience.

“For an education agent, maybe it's actual exam scores.”

Challenges and Consequences

10:51 to 12:55

Exploring potential risks and challenges of advanced AI systems.

“an accurate predictive understanding of how the world actually works based on its observations and actions.”

Safety Benefits of Experiential AI

12:55 to 14:00

Potential advantages of an adaptable AI approach in real-world scenarios.

“Interpretability becomes a massive challenge.”

Adapting AI Through Experience

14:00 to 15:46

Learn how AI systems can adjust their goals based on real-world feedback.

“but is being adapted based on experience to meet the why.”

The Challenges of Physical Reality in AI

15:46 to 16:36

Explore the limitations physical processes impose on AI advancements.

“And maybe we leave you, the listener, with one final thought from the sources.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00David Silver:Welcome to A Deep Dive. Today we're really digging into something big. A set of sources outlining, well, the next major shift in artificial intelligence.

0:10Richard S. Sutton:Yeah, it's pretty significant.

0:12David Silver:It's not just about, you know, slightly better chatbots. We're talking about a fundamental change. AI moving away from just relying on human data, human examples.

0:22Richard S. Sutton:And towards something new, something the sources are calling the era of experience.

0:26David Silver:Exactly. And our mission today is to unpack what that means, how we get there, and what it looks like.

0:32Richard S. Sutton:Right. And to really get it, you kind of have to look back first. We've had these phases, you see.

0:36David Silver:Like building blocks?

0:37Richard S. Sutton:Sort of. First was the era of simulation. Think AlphaGo mastering Go or chess.

0:44David Silver:Right. In those very controlled digital environments. Perfect information, clear rules.

0:48Richard S. Sutton:Exactly. Reinforcement learning agents getting superhuman in these closed tasks. That was impressive, but limited.

0:55David Silver:And then came the current era, the one that really brought AI into the mainstream.

0:59Richard S. Sutton:That's the era of human data. This is, you know, the GBTs, the large language models everyone's talking about.

1:04David Silver:Chat GBT and the like.

1:06Richard S. Sutton:Yes. They achieved this amazing generality, didn't they? By basically reading everything humans have ever put online or in books.

1:14David Silver:So they can write poetry, help with physics homework, even summarize complex legal stuff. Huge range.

1:21Richard S. Sutton:A sweeping range. That showed us what scale could do, learning from human knowledge.

1:25David Silver:Okay, but here's the tension the sources really highlight. This era of human data, it's apparently hitting a ceiling, a kind of brick wall.

1:34Richard S. Sutton:It seems like it. The problem is the really high-quality human data needed to push the best models significantly further, well, it's mostly been used up.

1:45David Silver:Consumed already.

1:46Richard S. Sutton:Largely, yes. or it's getting incredibly expensive or difficult to acquire more novel, high-quality data. So the progress we saw driven purely by learning from existing human texts and examples that supervise learning, it's slowing down. You can see it in the benchmarks.

2:01David Silver:So if we're aiming for something truly groundbreaking, like a new scientific principle or a medical breakthrough that humans haven't figured out yet.

2:07Richard S. Sutton:You're probably not going to find the key in data that just reflects what we already know. It's fundamentally limited by human understanding.

2:14David Silver:The next big leaps lie outside our current knowledge base.

2:17Richard S. Sutton:Precisely. We need AI to generate its own insights. And that's the core idea behind the era of experience.

2:25David Silver:Okay.

2:25Richard S. Sutton:Here, the agents learn primarily from their own interactions, their own experiments, generating new data as they go, data that improves over time.

2:33David Silver:And this self-generated data. The claim is it will eventually dwarf all the human data we've fed systems so far.

2:39Richard S. Sutton:That's the projection. Continuous learning from interaction creates an ever-expanding pool of experience.

2:45David Silver:Wow. Okay, that's a huge conceptual leap. From imitation to actual discovery. Before we break down the characteristics of this new era, is there a good example of this transition happening already? Like a bridge case?

2:58Richard S. Sutton:Yeah, the sources point to mathematics, specifically something called alpha-proof.

3:02David Silver:Alpha-proof. I remember reading about that. It achieved something quite remarkable in a human competition, didn't it?

3:07Richard S. Sutton:It did. It won a medal at the International Mathematical Olympiad, I mean, that's a domain of really high-level human reasoning.

3:15David Silver:Absolutely. So how did it manage that? How did it go beyond just learning from existing proofs?

3:21Richard S. Sutton:Well, this is where it gets interesting. It started with human data. It looked at about 100 ,000 formal proofs that humans had written over many years.

3:31David Silver:Okay, so that's the human data phase, learning the basics, the patterns.

3:34Richard S. Sutton:Right, but then it used reinforcement learning. It essentially started interacting with the rules of mathematics itself, trying to generate new proofs.

3:43David Silver:And how many did it generate?

3:44Richard S. Sutton:Get this. 100 million more proofs.

3:48David Silver:100 million. Autonomously.

3:50Richard S. Sutton:Through continuous interaction and self-correction within the formal system.

Read the full transcript

3:54David Silver:Okay, hold on. If it's generating that many on its own, how do you ensure quality? Isn't there a risk it just churns out loads of incorrect or useless proofs? How does it stay on track without constant human guidance?

4:08Richard S. Sutton:That's the power of grounding it in the environment's rules. Every step it proposes has to be mathematically valid according to the formal system it's interacting with. It's constantly being checked, not by a human opinion, but by the objective rules of math.

4:21David Silver:Ah, okay. The environment itself provides the correction.

4:23Richard S. Sutton:Exactly.

4:24David Silver:And that immense scale of exploration allows it to find pathways, novel strategies that humans just hadn't thought of or formalized.

4:33Richard S. Sutton:it's not just pattern matching. It's developing genuinely new problem-solving techniques.

4:38David Silver:So the argument is, if this works in math, it could work elsewhere.

4:42Richard S. Sutton:That's the contention. That once we harness the self-generation of knowledge in fields like physics, chemistry, biology, we could see truly unprecedented superhuman discoveries.

4:53David Silver:Right. So this brings us to the core of it. What makes an AI agent part of this era of experience? The sources lay out four key characteristics moving beyond just text.

5:03Richard S. Sutton:Yeah, four fundamental shifts from how we think about current AI. The first one tackles the limitation of time.

5:09David Silver:Time. How so?

5:10Richard S. Sutton:Think about current LLMs. You interact with them in short bursts, right? Snippets of conversation.

5:15David Silver:Yeah, you ask, it answers, maybe a few follow-ups, but then the context is often lost. It doesn't really remember you long-term.

5:21Richard S. Sutton:Exactly. It optimizes for the immediate reply. So the first dimension is streams versus snippets. Agents in the era of experience will operate over continuous streams of experience. We're talking long timescales, months, maybe even years.

5:36David Silver:Wow. So they maintain context and memory over a long period.

5:39Richard S. Sutton:Yes, which allows them to pursue complex long-term goals. Think about an AI trying to help someone genuinely improve their health over a year, or learn a new language fluently, or, you know, tackle a multi-stage scientific research project.

5:54David Silver:You can't do that in disconnected snippets. You need continuity.

5:57Richard S. Sutton:You need that long-term memory and the ability to adapt based on everything that's happened before.

6:02David Silver:Okay, so continuous time streams. What's the second big shift? If time is one constraint, we're removing.

6:07Richard S. Sutton:The next is how the AI interacts with the world, moving beyond just text.

6:12David Silver:Ah, right. Most LLMs live purely in the realm of language.

6:16Richard S. Sutton:Correct. The second dimension is grounded actions and observations. The era of human data really focused on human ways of interacting, reading, and typing.

6:24David Silver:Right.

6:24Richard S. Sutton:But intelligence, especially biological intelligence, uses sensors, motors. It acts in the world. So these new agents need to be grounded, meaning they interact with the real or digital world more directly, using motor control for robots maybe, or using sensors to perceive things beyond text, or even just operating a computer like a human does, mouse, keyboard, screen.

6:46David Silver:So an AI scientist could actually control lab equipment remotely or an AI assistant could navigate software for you.

6:54Richard S. Sutton:Precisely. It's about leaving the chat interface and engaging with the environment, physical or digital, autonomously.

7:00David Silver:OK, that's a big step, which leads logically to the third dimension, I imagine, if it's acting in the world.

7:07Richard S. Sutton:It needs feedback from the world. This is maybe the most crucial one. Grounded rewards.

7:11David Silver:Grounded rewards. Explain that. Current systems often rely on human feedback, right? We thumbs up or thumbs down and answer.

7:17Richard S. Sutton:Exactly. We provide the judgment, the reward signal. But the sources argue this puts an impenetrable ceiling on AI performance.

7:24David Silver:Why impenetrable?

7:25Richard S. Sutton:Because the AI can only ever get as good as the human judge's understanding. If the human doesn't recognize a brilliant unconventional strategy, they won't reward it. The AI can't surpass its teacher.

7:37David Silver:Ah, I see. If we want superhuman capabilities, the reward signal can't be limited by human intuition or knowledge.

7:44Richard S. Sutton:Correct. So the rewards need to be grounded in the environment itself. Signals that come directly from the consequences of the agent's actions in the real world. Objective signals.

7:54David Silver:Can you give some concrete examples? What does a grounded reward look like?

7:58Richard S. Sutton:Sure. Think about measurable things. For a health agent, it might be tracking resting heart rate or hours of deep sleep objective physiological data.

8:07David Silver:Okay.

8:07Richard S. Sutton:For an education agent, maybe it's actual exam scores. For climate goals, it could be measured CO2 levels. for material science, maybe the measured tensile strength or conductivity of a material the AI helped design.

8:19David Silver:So things that happen that can be measured independent of whether a human likes the outcome subjectively.

8:25Richard S. Sutton:Exactly. External objective events and signals.

8:27David Silver:But wait, if you just tell an AI maximize tensile strength, don't you risk unintended consequences. The classic paperclip maximizer problem where it optimizes that one thing to the exclusion of all loss, potentially dangerously. How do you guide these powerful long-term agents towards complex human values?

8:47Richard S. Sutton:That's a critical question. And the sources propose a more flexible approach. They call it a two-level goal system. The technical term is bi-level optimization.

8:56David Silver:Two levels. Okay, break that down.

8:57Richard S. Sutton:Think of it as separating the why from the how. The human provides the high-level goal of the why. Something like, help me improve my overall fitness.

9:05David Silver:Right, a broad objective.

9:06Richard S. Sutton:Then there's a lower-level system, maybe a neural network itself, that figures out the how. It learns to select and combine the grounded signals like heart rate, sleep duration, steps taken that best serve the high-level goal based on experience.

9:20David Silver:So the human sets the general direction, but the AI fine-tunes the specific measurable metrics it targets to achieve that direction.

9:27Richard S. Sutton:Yes. It provides a way to steer the AI towards complex goals without the human having to define every single reward parameter perfectly up front. The system adapts the how based on observing what actually works towards the why.

9:41David Silver:That sounds much more adaptable. Okay, that covers time, action, and rewards. What's the fourth dimension?

9:46Richard S. Sutton:This one is about the thinking process itself, non-human planning and reasoning.

9:51David Silver:Meaning not just copying how humans think.

9:54Richard S. Sutton:Exactly. A lot of current techniques, like chain of thought prompting, are essentially trying to get the AI to mimic human-like step-by-step reasoning.

10:03David Silver:Which seems useful, right? Makes it more interpretable.

10:05Richard S. Sutton:It can be, but there's a risk. If an AI learns only from human text and imitates human thought patterns, it might inherit our biases, our fallacies, even our outdated scientific models.

10:16David Silver:Ah, right. So if it reads enough old text, it might start reasoning based on, say, vitalism or thinking the sun revolves around the earth, because that's in the data.

10:25Richard S. Sutton:Potentially, yes, or more subtly, using heuristics that work for humans but aren't actually optimal or based on the true causal structure of the world. Think about things like animism in ancient cultures, or relying solely on classical physics when quantum effects are dominant.

10:42David Silver:So to get truly novel insights, the AI needs to reason based on reality, not just human precedent.

10:49Richard S. Sutton:That's the idea. It needs to build its own internal world model, an accurate predictive understanding of how the world actually works based on its observations and actions.

10:58David Silver:A model built from experience, not just text.

11:00Richard S. Sutton:Yes. And then it can plan its actions based on predicting the actual consequences in the environment, according to its model, rather than just following patterns of human thought. This could allow it to bypass limitations inherent in human cognition.

11:13David Silver:OK, these four dimensions, streams of experience, grounded actions, grounded rewards and non-human reasoning paint a picture of a very different kind of AI. If this comes to fruition, what are the big consequences the sources talk about? The potential upside.

11:27Richard S. Sutton:Well, the potential is pretty staggering. Imagine truly personalized agents adapting alongside you for years, helping with lifelong learning or managing chronic health conditions in a deeply integrated way.

11:40David Silver:That's the individual level. What about bigger picture?

11:42Richard S. Sutton:The most dramatic potential is probably in accelerating scientific discovery. Agents that can autonomously design experiments, perhaps control robotic labs, analyze the results, and iterate potentially discovering new drugs, new materials, new fundamental science at a pace far beyond human capacity.

12:00David Silver:A real step change in innovation.

12:02Richard S. Sutton:That's the hope.

12:03David Silver:But with that kind of power and autonomy comes risk. What are the major challenges or dangers highlighted in the sources?

12:10Richard S. Sutton:They identify a few key ones. First, job displacement. If AI can innovate and strategize long-term, the impact goes way beyond automating repetitive tasks. It starts touching creative and strategic roles.

12:22David Silver:Fields previously thought safe from automation.

12:25Richard S. Sutton:Potentially, yes. Second, safety risks might increase. You have autonomous agents pursuing complex, long-term goals. There are fewer natural points for human oversight or intervention compared to current systems.

12:36David Silver:More independent, harder to steer once they're running. That's the concern.

12:40Richard S. Sutton:And third, if these systems are using genuinely non-human ways of reasoning based on complex world models they built themselves.

12:48David Silver:They could become very difficult for us to understand. Black boxes, but even more so. Harder to interpret, broader to debug if something goes wrong.

12:57Richard S. Sutton:Exactly. Interpretability becomes a massive challenge.

13:00David Silver:That all sounds quite daunting. But you mentioned the sources also point out some potential safety benefits from this experiential approach, which seems counterintuitive.

13:10Richard S. Sutton:It does, but there are a couple of interesting points. The first is adaptability.

13:13David Silver:How is that a safety benefit?

13:15Richard S. Sutton:Well, think about it. A system that's constantly learning from its environment is inherently designed to deal with unexpected changes or failures. If a sensor breaks or the external world changes suddenly, like during a pandemic, for instance, a rigid pre-programmed system might just fail catastrophically. But an experiential agent can potentially observe the failure or the change, learn from it, and adapt its strategy to work around the problem.

13:40David Silver:So it's more robust to real-world messiness. It has a built-in error correction through learning.

13:48Richard S. Sutton:In a sense, yes. It's less brittle. The second benefit relates back to that two-level goal system we discussed.

13:54David Silver:Right, the why and the how.

13:56Richard S. Sutton:Yes. They call this correctable misalignment because that reward function, the how, isn't fixed, but is being adapted based on experience to meet the why.

14:05David Silver:You can potentially adjust the why if you see things going slightly wrong.

14:09Richard S. Sutton:Exactly. If the agent starts optimizing for, say, CO2 reduction in a way that has negative side effects who didn't anticipate a mild version of the paperclip problem, you don't necessarily have to scrap the whole system. You observe the negative outcome, you adjust the high-level goal or constraints, the why, and the system can then autonomously retune its low-level reward signals, the how, to correct its behavior.

14:33David Silver:So it allows for ongoing course correction based on real-world feedback and human concerns rather than needing perfect goal specification up front.

14:41Richard S. Sutton:That's the argument. It makes alignment potentially less fragile, more of an ongoing process.

14:46David Silver:Interesting. So this era of experience isn't just about creating more powerful AI, but maybe also a different, potentially more adaptable approach to managing it.

14:57Richard S. Sutton:It seems to be aiming for that. It's like a reconciliation. You take the broad knowledge capture of the large language model.

15:03David Silver:The era of human data.

15:04Richard S. Sutton:And you combine it with the powerful self-discovery and optimization we saw in earlier reinforcement learning systems like AlphaZero, which found new chess strategies.

15:13David Silver:The era of simulation's power, but applied broadly and grounded in reality.

15:18Richard S. Sutton:That's a good way to put it.

15:19David Silver:Okay, so let's try to wrap this deep dive up. The big takeaway seems to be that the future of AI isn't just about crunching more existing human data. It's shifting towards autonomous agents that interact with the world.

15:32Richard S. Sutton:Learn continuously over long periods.

15:35David Silver:And have their goals grounded in measurable real-world consequences, not just human preferences.

15:40Richard S. Sutton:AI stepping out of the library and into the laboratory, or maybe even the world at large.

15:44David Silver:A profound shift.

15:45Richard S. Sutton:Indeed. And maybe we leave you, the listener, with one final thought from the sources. Something to chew on. Even with super fast algorithms and AI that learns from experience, many of the most impactful advancements like testing a new drug, creating a new physical material, large engineering projects still rely on physical processes in the real world.

16:06David Silver:Right. You still have to run the clinical trial, build the prototype, synthesize the chemical. That takes actual time.

16:11Richard S. Sutton:Exactly. Drug trials take years. Building things takes time. This physical interaction, the need to actually do things in reality, might act as a kind of natural break.

16:22David Silver:A speed limit imposed by physics, regardless of how fast the AI can think.

16:27Richard S. Sutton:Perhaps. It raises the question, how much control, or at least breathing room, does the sheer pace of physical reality give us as we navigate this transition to more powerful experiential AI?

16:38David Silver:Something to definitely ponder. Thanks for joining us on The Deep Dive. We'll see you next time.

From the publisher

This "bitter lesson"-style position paper by David Silver and Richard S. Sutton introduces a shift in artificial intelligence from the "Era of Human Data" to the "Era of Experience." The authors argue that AI progress is slowing because the available pool of human-generated data is being exhausted, necessitating a new approach where agents learn predominantly from their own self-generated experience and interaction with the environment. This transition is characterized by agents that inhabit streams of experience rather than short interactions, utilize richly grounded actions and observations, optimize for grounded rewards from the environment, and engage in non-human planning and reasoning to achieve truly superhuman capabilities. The paper positions this new era as a critical next step that will reconcile the task-generality achieved by large language models with the self-discovery of knowledge exemplified by earlier reinforcement learning systems.


More from Best AI papers explained

All 475 episodes
Welcome to the Era of ExperienceBest AI papers explained · 17 min
Listen in VO