Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It Again

18 Aug 2026 · 54 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Why AI systems stop learning (weights frozen after pre/post-training), and how to enable continual deep learning plus continual world-model updates and abstraction/planning. Rich Sutton argues “all learning is continual,” and that the field wrongly treats it as special. He critiques LLM scaling limits from finite internet data and warns synthetic-data approaches still bottleneck on human domain expertise. He proposes continual backprop-style methods to avoid catastrophic forgetting, and a “big world hypothesis” that the real world is vastly more complex than any simulation, so agents must keep learning from experience.

Guests

Rich Sutton, reinforcement learning pioneer; invented reinforcement learning, wrote the seminal RL textbook, mentored key researchers (e.g., Dave Silver), and authored “The Bitter Lesson.” Khurram Javed, Sutton’s co-founder/former student from the University of Alberta; helped develop the “big world hypothesis” (wrote it up as a paper).

Key claims/examples

LLMs are positive/negative examples of “The Bitter Lesson”; AlphaGo/AlphaZero showed success from removing human priors; self-driving sim works only with human-in-the-loop sim fixes; continual learning requires algorithmic cures (step-size optimization, generate-and-test, continual backprop). Oak Lab aims to build a self-maintaining, self-consistent “mind” design that can learn and plan across scales.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Bitter Lesson and Reinforcement Learning

1:34 to 1:52

Explore the concept of continual learning and its relevance in AI today.

“lesson, the state of the world as we know it today, whether LLMs will get us there or not.”

Rich Sutton's Journey in AI

1:52 to 4:09

Discover Rich Sutton's personal experiences and motivations for pursuing AI research.

“What gave you the conviction to do that?”

Understanding Learning and the Mind

4:09 to 5:07

Delve into the fundamental aspects of learning as it relates to the mind and intelligence.

“They start questions saying how what I'm thinking is so different from everyone else.”

Impact on the Field of AI

5:07 to 5:40

Reflect on Rich's contributions to AI and their significance in the field.

“to feel they need to call it continual learning.”

The Bitter Lesson and Its Implications

5:40 to 7:35

Examine the essence of the Bitter Lesson and its implications for AI development.

“You've been able to educate a lot of students who push the frontier as well.”

Large Language Models and Learning

7:35 to 10:48

Discuss the relationship between large language models and the Bitter Lesson.

“You know, as the bitter lesson expresses, it's something that you can observe for a long time, for many decades.”

Synthetic Data Generation in AI

10:48 to 13:14

Learn about the challenges of synthetic data generation and its limitations.

“too much on human knowledge, and it eventually holds us back.”

The Limitations of Synthetic Data

14:01 to 16:40

Explore the complexities and limitations of synthetic data in AI.

“It doesn't matter if you generate synthetic data that captures 50 different universes.”

The Role of Human Engineers in Simulation

16:41 to 17:59

Discuss the importance of human engineers in creating and refining simulations.

“So I think the important question to ask here is how many engineers were involved in building that simulation?”

The Bitter Lesson: Prior Knowledge vs. Learning

18:00 to 21:09

Examine the conflict between leveraging prior knowledge and the need for continual learning.

“Okay, I want to move to another part of the bitter lesson, removing human knowledge.”
Show all 26 chapters

Imagining Future Learning Machines

21:10 to 22:02

Speculate on how future machines could learn from experience and interact with their environments.

“And that's what all that matters in the long run.”

The Shortcomings of Current AI Models

22:03 to 23:28

Critique the limitations of current AI models in terms of learning and adaptation.

“Well, it could look like a robot, but it also could live entirely on the Internet.”

Learning from Human Experience

23:29 to 25:59

Analyze how human learning and adaptability can inspire AI development.

“The only point that the big disagreement is we don't let them learn after that.”

Critique of Supervised Learning

26:00 to 28:00

Debate the relevance of supervised learning in comparison to natural learning processes.

“And that is the capability I think that's extremely useful we would want in our systems.”

The Nature of Learning and Schooling

28:00 to 29:00

Discussion on how animals and humans learn differently, with a focus on the role of school and experience in learning.

“And school is a very special thing that even we didn't have up until, I don't know, a few hundred years ago.”

Understanding Intelligence: Abstraction and Learning

29:00 to 30:50

Exploration of the differences between human intelligence and animal capabilities, emphasizing abstraction in learning.

“to force me to go to school and deal with all the structure.”

The Role of Experience in Human Learning

30:50 to 34:30

How humans and machines can learn from experience, and the necessity for imaginative thinking in learning processes.

“Does your world model, I guess, span sensory motor learning all the way up?”

Challenges in Machine Learning: Abstractions and Knowledge

34:30 to 37:50

Discussion on the gaps in current machine learning systems regarding abstraction and continual deep learning.

“and you can expose this at the edge of human knowledge, but you can also study this problem at the sensory motor stream level.”

The Alberta Plan for AI Research

37:50 to 41:20

Insights into Rich Sutton's 12-point Alberta Plan for advancing AI research, focusing on continual deep learning.

“And more importantly, where would that come from?”

Advancements in Continual Learning Algorithms

41:20 to 42:07

Exploration of new algorithms for continual learning in AI, addressing issues like catastrophic forgetting.

“We published in Nature a couple of years ago, and it is exactly like backprop, but you also plant new seeds of units that are newly initialized with random weights.”

Exploring Learning Algorithms in AI

42:07 to 43:13

Learn about the application of new algorithms for continual learning in AI models.

“And that's what we hope to do in the next couple of years.”

The Ambitious Goals of Their Company

43:13 to 45:26

Understand the ambitious vision for AI that combines broad and specialized knowledge.

“There was a point in 2016 to 2018 where a lot of people were exploring these ideas quite a bit.”

Feasibility of Future AI Developments

45:26 to 47:16

Discuss the potential challenges and breakthroughs in future AI technology.

“It almost seems that it's such an ambitious vision.”

Challenges in AI Paradigm Shifts

47:16 to 49:16

Examine the difficulties that come with shifting paradigms in AI research.

“You have lots of people at research labs that have access to way more than that.”

Concept of Multiple Learning Systems

49:16 to 52:04

Learn about the significance of having multiple systems for learning in AI.

“If everything goes right, we implement the architecture, we can have genuine continual learning, and we can form abstractions so that we can do planning and reasoning.”

Building a Cohesive Team for AI Development

52:04 to 53:12

Discover the approach to building a small, aligned team for AI advancements.

“What kind of people are you looking for?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Khurram Javed:People think I have a radical point of view sometimes. They start questions saying how what I'm thinking is so different from everyone else. But I don't see it that way at all. I see it as like I'm thinking the ordinary way. It's just everyone else that's thinking a bit weird. And I mean that like, you know, it's just the recent times people are thinking weird. Before there was all this AI craziness, you talk about, you wouldn't have to say continual learning because it wouldn't make any sense to talk about learning that wasn't continual. All learning is continual. We always act and we learn. That's just the normal way of thinking.

0:35Khurram Javed:I'm not weird. The field is weird. The field, they need to call it continual learning. It's just learning.

0:59Rich Sutton:We are honored to have the great Rich Sutton with us here today. Rich, you invented reinforcement learning. You wrote the seminal textbook. You had the key students in the field, folks like Dave Silver. you wrote the essay, The Bitter Lesson, that I believe is the Bible of the field, and you have just been one of the greats in propelling the field forward. So thank you for taking the time to join us today. Rich is joined by Khuram Javed, his co-founder and former student from the University of Alberta. The two of you have set off to found Oak Lab. I'm very excited to talk to you about that today.

1:33Rich Sutton:So for today's session, we're going to start talking about The Bitter lesson, the state of the world as we know it today, whether LLMs will get us there or not. And then we're going to transition to start talking about your research agenda and your plan for Oak. Rich, maybe take us back. I was going to start with a better lesson, but I actually want to start earlier than that. Decades ago, you decided to dedicate your career to reinforcement learning, to deep reinforcement learning in particular, and you established the University of Alberta as a bastion of that back when I think the field was very much in its infancy.

2:08Rich Sutton:What gave you the conviction to do that?

2:09Khurram Javed:What else are you going to do? We're trying to figure out the mind and learning is a central part of the mind and having a goal is a central part of the mind, central part of intelligence. Yeah. So I was just doubling down on what I was always thinking.

2:26Rich Sutton:Did people think you were crazy at the time? um it was a winter it was an ai winter what year was this it was in 2003 okay and it's kind of

2:40Khurram Javed:crazy actually the truth because i was like really sick i was dying i was actually dying of cancer in 2003 and uh but i i wasn't quite dead you know i've been trying for a number of years And I wasn't dead. I was in another remission. And so I said, well, I'm not dying. I haven't succeeded in dying. So I might as well, you know, it's going on long enough. I might as well just try to get another job. And so I went to Alberta and started teaching there. And then in the end, I didn't die. It's kind of amazing. It's like that. I'm joking about it now, but it was quite serious. And it's an even more poignant question.

3:22Khurram Javed:Why did I continue to work on this research stuff when I only had a few months? I would always keep reminded what I think it's Benjamin Franklin is supposed to have said, that if you ever wonder why someone is doing something, it's almost always one of two things. It's either habit or vanity. So I think that was probably true. Maybe it was a habit to just kept doing what I always was doing, or maybe it was vanity. I don't know. I think it was more like habit because I was dying. Wow.

3:57Rich Sutton:Divine intervention.

3:58Khurram Javed:It's always been easy for me to be very determined. And I'm going to go even longer on this answer.

4:06Rich Sutton:Please do.

4:08Khurram Javed:People think I have a radical point of view sometimes. They start questions saying how what I'm thinking is so different from everyone else. But I don't see it that way at all. I see it as like I'm thinking the ordinary way. It's just everyone else that's thinking a bit weird. And I mean that like, you know, it's just the recent times people are thinking weird. If you look back, what people thought about the mind for, you know, even just a decade, you'll find the kind of thoughts that, you know, learning is important. You've got to have a goal. And, you know, perception is important. We are low level beings.

4:48Khurram Javed:We are generating actions and receiving data at a fast speed, and yet we have to think at higher levels. And, you know, go back a few, before there was all this AI craziness, you talk about, you wouldn't have to say continual learning because it wouldn't make any sense to talk about learning that wasn't continual. All learning is continual. You know, it's not a special phase. We always act and we learn. That's just the normal way of thinking. I'm not weird. The field is weird. to feel they need to call it continual learning. It's just learning.

5:22Rich Sutton:I'm not weird. Everybody else is. That's a good motto to live by. We're going to have to send out a next post about that. We're very happy that you lived on. The field is happy that you lived on. And thank you for pushing the frontier of AI. I'm really happy. Thank you for pushing the frontier of AI. I am sure really happy. And thank you for all of that. You've been able to educate a lot of students who push the frontier as well. How'd you pick them? How'd you over the last 20, 30 years?

5:51Khurram Javed:Oh, you are giving me opportunities to be humble. I like to be humble and point out how all these great decisions just happen. And that's the way I feel about students. I don't feel that I choose them very well. Sometimes I'm lucky, sometimes I'm unlucky. I don't feel I'm particularly good at picking my students. I'm looking at Kerm now. I think sometimes you end up with really great ones. David picked David Silver picked me yeah how is it that I got you Karim yeah that was also so I

6:25Rich Sutton:finished my master's not with you uh and I was planning to join industry and then we were collaborating on a project which also just started organically like there was something I worked on that which was in a meeting then they mentioned that I worked on it so I got pulled into it we started collaborating it went really well like I felt so happy with that collaboration which also felt really good about it. And then six months down the road, we had made some progress and it just made sense to convert that into a thesis proposal. So no point did I apply, at no point did I ask, should you be my PhD advisor?

7:01Rich Sutton:We worked together, then we decided this would be a pretty good thesis. And then after that, I applied for the PhD. Life works in unexpected ways. Take us to 2019. You wrote The Bitter Lesson, which has become the modern tome. 2019 was a funny time to be writing that piece because ImageNet was 2009, AlphaGo was 2015. What caused you in 2019 to reflect and to write that? Because it was before the current kind of scaling paradigm around large language models had taken off, but it was after deep learning had really proven itself.

7:36Khurram Javed:Well, it was a long time coming. You know, as the bitter lesson expresses, it's something that you can observe for a long time, for many decades. And it's definitely at least as much due to the round of symbolic AI, which I lived through. It's all about not getting distracted by trying to put in your human knowledge and just paying attention to what the problem needs and how you can scale with computation. I know I wrote versions of it at least a year before. And I gave talks. I gave a talk a year before. And it wasn't a particular response to the moment. It was a particular response to my long experience.

8:22Khurram Javed:Different people trying to think in different ways about how you can make smart systems.

8:29Rich Sutton:What is the essence of the bitter lesson? Yes, it's not a bitter lesson. It may be the phrase that I hear used the most in my meetings these days. Is this bitter lesson pilled? Is it not bitter lesson pilled? Yeah. I would imagine, given the popularity of the phrase, it's probably been tortured and misused in different ways that you didn't originally intend it. So what do you think, what is the essence of it, and where do you think people go wrong in their attempt to understand it?

8:54Khurram Javed:Yeah, you're making me think about X now. I recently made a post where I tried to do the bitter lesson in 26 words. It goes something like, don't be distracted by human knowledge as AI traditionally has been many times. Instead, focus on learning methods that will scale with computation, like search and like learning. So it's really all about focusing on algorithms and improvements. It's not saying you don't need fancy algorithms. You need fancy algorithms, but you want fancy algorithms that will scale with computation. Rather than scaling with data. Rather than scaling with human input. Yeah.

9:45And then the question, if I can anticipate, yeah, what about large language models?

9:53Khurram Javed:Yeah. Are they consistent or inconsistent with your essay? And I thought about this, and I think there's another expose about it, but the conclusion is that it's both a positive example and a negative example of the bitter lesson. First, large language models enabled enormous scaling with computation, and you could just drink in the internet and scale so much. So it was a way of getting a much more capable system just by methods at scale. Then after that, as you go on further, it eventually gets limited by that information. information. The internet is finite, and it's hard to get more examples.

10:39Khurram Javed:And the world is big, and the world is massively bigger than everything we stored on the internet. And so in the end, it seems like it could be, I guess that would be a positive example of when, you know, we relied too much on human knowledge, and it eventually holds us back.

11:00Rich Sutton:Can I just push on this a little Well, then.

11:02Khurram Javed:Yeah.

11:02Rich Sutton:It seems like a lot of what the foundation model labs are working on right now is synthetic data generation in order to kind of get us beyond the fossil fuel that is the existing human internet. Is synthetic data generation kind of as part of this LLM scaling paradigm, is that a general method that leverages computation?

11:24Khurram Javed:No. Why? That's just a big mistake.

11:27Rich Sutton:Why?

11:28Khurram Javed:Well, it's such a big, maybe it's the next big lesson. It's been floating around Alberta for five or ten years. And we call it the big world perspective or the big world hypothesis. Kheram, who eventually wrote it up as a paper, there's a little paper called the big world hypothesis.

11:50Rich Sutton:The big world is that the world is infinitely big. There are infinitely many things to learn. and you can have people generating these synthetic data sets, but there would always be more things to learn. And because of that, if you could just learn from experience, if you could remove the humans from the loop, then you would have systems that can do everything because the world is big. There are many tasks that we want them to do, and they would be able to do anything by learning from their own experience. Going back to the synthetic data question too, who decides what's a good synthetic data and what's the bad synthetic data?

12:26Rich Sutton:because I can write a program that can output a lot of synthetic data, which would hurt programs. Right now, I would say humans decide. And that's the bottleneck where, okay, you can have humans deciding how to generate these data sets, but you need human experts who know what's a good data set and what's a bad data set for that approach to scale. So it is bottlenecked by humans. Doesn't my loss curve decide, like, how much better did I get with this data set versus that data set? Right. But if all the engineers, OpenAI Anthropic, or all the big neo labs, that engineers went on vacation, who would generate the synthetic data?

12:58Rich Sutton:That's the question. It doesn't come from agent's experience. It's not something that the agent is generating itself. Some human has to decide what is the right synthetic data to generate. And that requires human expertise. So for example, if you want a system to do something very challenging from a physics point of view, maybe you want a drone that flies with echolocation, like a bat, for example, what's the right synthetic data for that? I think you would need to hire domain experts to go figure out what is the right data and generate it. And then maybe you would be able to learn from that. But the domain expert has to exist first.

13:34Rich Sutton:So we are bottlenecked by human expertise at that point. But you can have infinite synthetic worlds. The existing world is finite. But let's go back to the echolocation thing, right? That's what I want. I want a drone that can localize itself and move with echolocation. That's my goal. The robot, that's a robot, is generating its own experience. So it could totally learn from its own experience, but it wouldn't be able to. It doesn't matter how much synthetic data you generate. It doesn't matter if you generate synthetic data that captures 50 different universes. It would not allow you to do that task without humans figuring it out first.

14:12Khurram Javed:First, just say the synthetic data is wrong. I mean, it won't be correct. It'll be a synthetic world. It won't be the real world. And it will matter. The world is incredibly complex. If you write a little program, because there's going to be a little program that will generate the synthetic data, it'll be a small world. So, for example, what's important to me is what's going on in your mind right now. Okay? And you're saying, why don't I get some synthetic data to tell me what's going on in other people's minds? No, there's no way we can have synthetic data for other people's minds, and other people's minds matter to us.

14:56Khurram Javed:You know, like, I talked to you guys about investing today, so I care what's going on in your minds, but how can I get synthetic data on such a thing? Really, you can't even get synthetic data on anything. You can't get synthetic data on how the drone is going to interact with its environment, in the physical world, in the friction and wear in the motors of this robot. The world is infinitely complex, and any simulation of it is microscopic. The big world hypothesis, let's say what it is, is that the world is massively more complex than your mind, than any agent. And this is obvious because the world contains many other agents.

15:40Khurram Javed:So, because the world is massively complex, there's no way you can do anything that might claim to be optimal or perfect. You're going to be imperfect and you have to have approximations and those approximations will be severe. And so, because of that, that is ultimately the reason why we have to continue learning. If you want to think of it as a reason. We have to continue learning because we'll encounter some particular part of this immense world. And we'll have to learn an approximation that's tuned to the part of the world we're in, not to all the other parts that we're not in.

16:20Rich Sutton:I'm going to push on this one more time. And sorry, I'm being argumentative for the sake of being argumentative. I'm trying to understand. My understanding is that the newest cohort of self-driving car companies, many of them were primarily trained in sim. and then they have to do some sort of post-training, I guess, to make sure they work in the real world, but that it's been a very effective pipeline. Yeah. So I think the important question to ask here is how many engineers were involved in building that simulation? And are we ready to say that the only problem worth solving are those where we can hire a large team of engineers to first make a simulation?

16:58Rich Sutton:And I'm sure they had to do multiple iterations where they made the simulation, they learned it, they realized there was a sim-to-real gap that was not acceptable. Then they fixed it. So there is this human in the loop fixing the simulation. Like they are getting feedback from the real world, humans, and then they are fixing the simulation. Why can't we just remove the human and let the agent do it itself?

Read the full transcript

17:18Khurram Javed:And then when it actually drives, again, something unexpected will happen.

17:22Rich Sutton:And that's when you actually want to learn from experience. So your point is there's just so much more data that's going to come from experience than there possibly can be from humans curating, creating data. Yeah, and I think there is obviously value in learning from simulation, and there is a way of doing it. The agent can learn a model from its own experience, and when the agent learns it, it's much better because if the model is incorrect, it can fix it by continuing learning. If the humans are making a simulator, then the model only gets updated when the humans figure out that something's wrong.

17:54Rich Sutton:So, yes, planning is important. the agents should learn from simulators, but simulators they make themselves. Okay, I want to move to another part of the bitter lesson, removing human knowledge. From your essay, quote, seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain. But the only thing that matters in the long run is the leveraging of computation. And if I, you know, your former student Dave with AlphaGo and AlphaZero, for me, that was an example of a triumph of removing human priors. Did that result surprise you?

18:29Rich Sutton:I guess why or why not?

18:31Khurram Javed:Of course, it made me very happy. It made me feel vindicated. It could have gone either way. It wasn't that because prior knowledge can help. There's nothing wrong with prior knowledge. And I say this right at the very beginning of the bitter lesson. I say, there's no reason why there has to be a conflict between prior knowledge and then learning knowledge. You can put some prior in there and then start learning. There's no reason in principle why these have to be opposed. In fact, they're all about knowledge. Life is gaining knowledge and having knowledge. How somehow nature and nurture became enemies?

19:12But really, prior learning is what you already

19:15Khurram Javed:and then you learn more. And they should be friends. But as I say at the beginning of the bitter lesson, in practice, they have been enemies. In practice, people who had an affection for existing human knowledge ended up wanting that to win. And so they wanted to minimize or dismiss learning. And so now I'm sure your sense of me is that I'm someone who loves learning and wants to dismiss prior knowledge. But I'm really someone who's interested in the mind. The mind is you have prior knowledge and then you get more. And then once you've gotten more, then that becomes your prior knowledge as you get more and more and more.

19:59Khurram Javed:And these two things work together. I end up appearing to be someone who's interested in learning primarily because all the rest of the world is talking about all you need is enough knowledge. You don't need to learn. You know, large language models are, oh, we're going to put all this knowledge into the system, in the large language model will not learn when it runs. You know, it's talking to people, it's interacting. It is absolutely, the weights never change. So, you know, I am not the weird one. It's you guys that are the weird one that think that that's possible, that you could possibly, you know, they claim they can make a PhD level experience and expertise out of something that doesn't learn at all anymore.

20:44You know, so, you know, I'm not the weird one.

20:47Khurram Javed:So your recommendation is let the algorithms run for a much, much longer period of time before feeding it. Continually learn. Before you feed it data or prior data. Or drip the prior data along the way, prior knowledge. So both are important. But in the long run, you've got to gain and structure the gaining of new knowledge. And that's what all that matters in the long run. And as you are doing this, yeah, there'll be some that you got previously. Like how it works if we look into the future when we have intelligent robots, will we have them all learn from scratch? Or will we like copy them and ask them to keep learning from wherever they are?

21:39Khurram Javed:I mean, they'll be digital and it'll be easy to copy them. And so instead of having like this huge thing where we're spending zillions of dollars to retrain them from the Internet, we'll just copy the agent and keep learning from there. And so in some sense, the prior knowledge should be dismissed because you're just going to copy it from the previous robot.

22:02Rich Sutton:So why don't you just describe for us what you think a machine or a computer that learns from experience looks like?

22:08Khurram Javed:Well, it could look like a robot, but it also could live entirely on the Internet. You could like, for example, routing of packets through the Internet and do that in a way that's sensitive to experience and becomes better over time. Or you can interact, be the user interface that's interacting with people, like on your phone or on your computer, and it becomes better over time. You know, like an intelligent assistant, you know, has to become better over time, has to know what you want.

22:36Rich Sutton:Would your contention be that the current paradigm of, you know, the popular LLM-based assistants, would your contention be that these are not experiential learners or continual learners? And if so, what is the fundamental gap? Are you serious?

22:55Khurram Javed:I mean, obviously.

22:56Rich Sutton:They learn memories about me. They're, you know, they're doing some in-context learning.

23:03Khurram Javed:Their weights never change.

23:05Rich Sutton:And by the way, is a small number of the weights changing sufficient, or do you need all the weights to be changing?

23:10Khurram Javed:Well, so think of all the structuring and generation of new concepts that went into creating the large language models. All that is the weight learning. And you want to be able to continue doing that. You don't want that to happen just once.

23:28Rich Sutton:Another way of saying it is we do too much pre-training and post-training before we launch the models. They don't learn after that. The only point that the big disagreement is we don't let them learn after that. Yeah, we don't let them learn after that. We can do as much pre-training as we want. That's okay. Post-training is fine. But then when I'm using the model, it stops learning. You can give it more context. You can change the state of the model by giving it more context. So it has already learned that if the state is different, if the state says something new, then it will use that to make the next prediction.

23:59Rich Sutton:But the model is unlearning. Cursor is tab autocomplete model. it does get updated based on those models those ways change those are two examples like Cursors, Tab and I think the Composer they were also updating those are two examples of continual learning but it can be much better so the way they do it as far as I understand is a lot of people are using Tab they collect all this data so coming from millions of users or thousands of users and then they do one update of the policy from this batch data so this could work but imagine I want to teach this model something specific. I don't want to fight with 100 ,000 other people about what they want to teach their models.

24:39Rich Sutton:I want to teach my model something very specific and I want to do it to my version of the model. I don't care about the shared knowledge that the model has coming from other people. And so it's a very inefficient way of doing it. It seems like the way that this is currently done is that there's fundamental skills maybe that are learned in the weights that are common to everybody. And then there's personalization that happens in the form of context, right? Right, yeah. Is that not the right mental model for how learning should work? Like should all the contexts live in the weights themselves? So context can be in the state too.

25:14Rich Sutton:Could be both, but you still need to be able to update the weights. So if I give you an example, some really good use studies are with human disabilities. When humans go through something that changes their mind or some sensors, you can see them adapt. So for example, we have proprioception, We have internal sensors that tell us how the body is positioned, and we use this for walking. There are cases where people lose this ability completely, and then they can't walk at all because that is literally the foundation of their walking policies. It is ingrained in the brain. But then over the course of two, three years, they can learn to walk again by looking at their feet.

25:50Rich Sutton:So visual feedback through that. So brain is insanely plastic in the sense that it can learn a lot of things, something that has been true for 20 years, when it stops being true, it can go and update that and get rid of that. And that is the capability I think that's extremely useful we would want in our systems. What is there for us to learn from how human babies or animals learn? And how much inspiration do you take from that?

26:16Khurram Javed:Well, we take a lot of inspiration. We don't take it as a requirement that the AI has to behave like the natural system, like babies or people or animals. But it's a source of inspiration. Inspiration, but not constraint from animal learning.

26:35Rich Sutton:Be consistent with the bitter lesson. Yeah. Where do you think we should most seek to draw inspiration from the way that biological learning works that is not present in today's systems?

26:47Khurram Javed:I feel like I'm just giving opinions now, but they're just obvious opinions. So I think it's apparent that no animal learns by supervised learning. Because we don't get examples of how our muscles should twitch. And that's our output. But all of school is supervised learning. I know. Absolutely not. But even if it was, school is like a tiny fraction of what we learn. Like we learn to see and we learn to walk and we learn how the world works. But even in school, no one tells us how we should twitch our muscles.

27:30Rich Sutton:The knowledge skills I acquire were from supervised learning in school. I don't want to say that learning from others, transmission from others is not important.

27:41Khurram Javed:It's like extremely important and language is extremely important. But what are we missing? There is no supervised learning. There's no targets that are given to us. You know, you hear the right answer is, you know, what's the capital of France? And we know the answer is Paris. okay but no one tells me how I should pronounce Paris you say the answer is Paris and I listen to you and I hear your words and you know I will make some other muscle motions to produce the answer Paris it's not literally supervised learning anyway yeah so I think it's really true I mean well anyway the first thing is the school is it is irrelevant like you know squirrels don't go to school and learn that.

28:37Khurram Javed:They might. Animals don't learn that way. And school is a very special thing that even we didn't have up until, I don't know, a few hundred years ago. But isn't it something we've evolved to? It's not part of the essence of intelligence. And it's a distraction to think of that as your primary example of learning is this thing which we didn't do as animals.

28:59Rich Sutton:I wish you had been around to tell my parents that. to force me to go to school and deal with all the structure. The thing is, squirrels are wonderful at jumping off trees, but squirrels can't prove math theorems. And if I want to learn how to prove a math theorem, I go to school.

29:14Khurram Javed:Yeah. They also don't have DVDs and iPods. You know, a lot of things. They can do things that we can't do. But math theorems, yeah, and they don't play chess. You know, it's sort of like Moravex Paradox. They're these advanced things that we think of as really intelligent, but they're sort of easy for computers to do as opposed to all these regular things that are hard, like moving and seeing with attention and everything. I think supervised learning is a good thing. you know just mentions i like to think look for obvious things no one tells us how to twitch our muscles by giving us examples because they couldn't possibly because we have had to twitch our muscles we've had to figure that out yeah and their answer would be wrong right so if i moved my

30:10Rich Sutton:mouth and my tongue and my vocal cords exactly the same way that rich does to pronounce paris i'm sure a very different sound would come out so in some sense rich or no one knows the right way of producing a sound with my body. Only I know that. Yeah. It seems to me that many of the most, I guess, the most raw, like sensory motor capabilities, especially related to movement in the physical world, I agree with you that that seems something that's inherently learned from experience. It seems to me, though, that there are higher levels of abstraction that bring us closer to, you know, what makes humans great.

30:46And much of that doesn't live in this low level

30:49Rich Sutton:of sensory motor learning. Does your world model, I guess, span sensory motor learning all the way up?

30:56Khurram Javed:Yeah, that's the ambition. Absolutely. And squirrels, by the way, can do some enormously abstract things.

31:03Rich Sutton:What's the coolest thing a squirrel can do?

31:05Khurram Javed:Well, it can always get into your bird feeder, no matter what obstacles you put in the way. You know, it can find new ways to jump and climb and do lots of things.

31:16Rich Sutton:You calculate trajectories pretty well. Animals are pretty good at understanding the physical world without the mental calculations that we think we are doing when we think about launching ourselves into space. Breaking a fall, they can do it in real time the right way to prevent injuries. Okay, fair enough.

31:34Khurram Javed:I think it's just a question of degree between, and I like to think that animals, other animals, are very close to humans. I think it's hubristic to try to emphasize what we do differently, you know, how we're different from animals. It's better to see the commonalities. And I think we are just in a question of degree. It's degree and of course society and culture give us big advantages. Language gives us big advantages. Can I just push on some of this? Yeah, good.

32:03Rich Sutton:Because I want to back up, Sonia. So I believe animals and children learn from experience and do incredible things learning from experience. And when my son was two or three or four, I'm like, wow, this is really interesting that my son can learn these things without nobody really teaching him how to do these things. But at the same time, what Sonia is saying is like what makes human uniquely human, to be able to go to outer space, build a rocket. Those are not things that are learned 100 % from experience because before you launch the rocket, you actually have to abstract thinking through it in a way that is not learned from quote-unquote experience because you don't know if it's going to work or not.

32:51Rich Sutton:You have to imagine it. How do we teach a machine to imagine things that were not available before? That's probably the thing that we're trying to push on because we're not quite understanding that.

33:04Khurram Javed:We're absolutely going to agree with you there. You have to be able to plan. you have to be able to imagine.

33:09Rich Sutton:Would you say that humans a thousand years ago, before they had done most of the things that we were talking about, were they as intelligent? If, for example, someone from that era was exposed to this new culture, would they be able to get the same skills and start doing useful things? Even over the last 10 ,000 years, I don't think the human brain has evolved that much. Fundamentally the same machine. Fundamentally the same machine. But we've built up 10 ,000 years of knowledge. Yes. And I get to learn 10 ,000 years of knowledge by going to school through supervised learning. And I get all that much, much faster than trying to learn through experience.

33:47Rich Sutton:So I think you're totally right. So we would want our systems to learn from experience. And part of their experience would be getting exposed to our culture and then learning about our culture. They should learn from that. That's all good. But let's talk about when someone goes and does a paradigm shifting thing. So everyone gives the example of Einstein, but I think there are many examples. Learning is that too, like looking at learning thing versus programming things. So when these paradigm shifts happen, I would say it's a human who has accumulated all this knowledge. And then from their experience, they're building new abstractions, they're planning with them, and then they're discovering new knowledge.

34:24Rich Sutton:And that skill of coming up with new abstractions and then learning more and more and more and more planning with them, that skill is totally missing in our current systems. and you can expose this at the edge of human knowledge, but you can also study this problem at the sensory motor stream level. So we're not arguing with the principle.

34:46Khurram Javed:We need to form abstractions so we can reason at a high level. You guys are coming close to doing that thing that I said we should never do, which is argue is prior knowledge important or gaining knowledge important. You know, that's what you guys just said. You said you're still going to have to learn things and you're saying, oh, I can get things from my culture and from prior knowledge. But these should not fight for each other.

35:10Rich Sutton:On the exact thing around paradigm shifts, how do we create a machine that understands when to shift the paradigm? I think through its experience, right? So it would have to, through its own experience, it can't rely on human knowledge because we're assuming the humans see one paradigm and we want a different way of looking at things. And so through its experience, it has to find something that is better. Maybe it generalizes better and makes better predictions. Maybe it's better in some other ways, but it has to be through its own experience.

35:40Khurram Javed:The big challenge that we don't see in our field, the ability we don't see in our field yet, is the ability to learn a model and then plan with a model. We can do the math things and we can do AlphaGo because the games, we know the model. We know how the moves work. And in math, we know what the operators are. We know lean will take us from one state of knowledge to the state of the proof to the next state. But if we have to learn the models, there are no, I'm going to say it, it's probably maybe a weird example, but a counter example. But I can say there's no instances of learning the model and then planning with the model in our field.

36:22Rich Sutton:at least not with self-discovered abstractions. So there are people who say, I'm just going to learn a model of what happens in the next second or next millisecond. But that's not how our models work. Our models are more abstract. Our models are quite different. So one of the things I like about what you're doing here is you're not just sitting around pontificating or lamenting the state of the world as it is. You're very action-oriented. That's why you've started a company. So let's start talking about that a bit. In 2022, Rich, you laid out a very specific 12-point plan, the Alberta Plan for AI Research.

37:02Rich Sutton:Maybe tell us about that.

37:03Khurram Javed:So the Alberta Plan came about because we did have general ideas, but we also needed to convert them into smaller chunks. So the 12 steps are the attempt to crystallize particular chunks. there's a very important early step, step two, which is continual deep learning. And we think that one is like almost the most important because it unlocks everything else. If you could do continual deep learning, you could then continually update your model of the world. And then if you knew how to do the abstraction rights and like the second half of the steps are all about how to get the abstractions right.

37:48Khurram Javed:and not only by abstractions right what I mean but I don't mean get the right abstractions because no one can say what the right abstractions are that depends on the world that you're in your agent would have to learn the correct abstractions for whatever world it's in and so you know maybe those are the two key things you have to find the right abstractions and then you have to do able to continual deep learning

38:10Rich Sutton:I think that a lot of the people in the field realize that we need models we need to plan with them but the abstractions tell us what the model should be conditioned on. So what should the model predict? What are you going to do? And then something is going to happen. And more importantly, where would that come from? So I really like the example of elite athletes. If you ask elite athletes about how they do certain things, they would have weird niche terminologies for doing very specific things. They were like, you know, I do this thing and they would have a name for it if they communicate. Sometimes they don't even have a name for it if they're just doing it alone.

38:45Rich Sutton:So how did they come up with those abstractions? That's, in some sense, a crucial thing that's missing that the later half of Alberta plan answers. Can we talk about the continual deep learning part? Is it an algorithmic gap that exists today or is it just a practical deployment infrastructure data privacy gap? Because if I wanted to do, call it naive updating of weights based on user interaction, I can do that today, right? And so what, in your opinion, is the biggest thing that we're missing to kind of get to continual deep learning? Yeah, so it's absolutely an algorithmic gap. You can do the naive thing, but then you'll see all sorts of problems.

39:24Rich Sutton:For example, if you say I'm going to take one sample and then I'm going to update my whole model with that one sample, you will run into this problem that now all of the previous knowledge in the model, it's impacted negatively. And the way currently we get around this is exactly what cursor does. they don't use one example. They use a large badge coming from a lot of users. So in use cases where you can have that, you can do continued learning. But most use cases, you don't have that. Most use cases, you have a single stream of data. And then if you apply it to the naive thing, it just completely destroys your prior knowledge in a very destructive way.

40:01Khurram Javed:Catastrophic forgetting. Yeah. That is so, yeah. But it's totally curable. You have to have the right algorithm. What's the cure? Yeah, exactly. What is the cure? You know, first you need to do what we call step size optimization. And it means every weight in your network has to have a separate step size. So some will move fast, some will move slow. And you will have to meta-learn these step sizes for each weight. Most of your network will have weights that have tiny step sizes. So then when you train on a new example, they don't get destroyed. It happens just to the right places. And then secondly, you have to use some form of generate and test, which is in feature space.

40:49So you come up with new features or new units and without following gradients, because gradients is a very slow process.

40:55Khurram Javed:You only move in a direction if you know it's the helpful one. And that's always going to be very slow and doesn't give you a path to grow more and more complex and to have sustained learning. You need to have something that just proposes a bunch of new units and then goes from there. I guess so there is a specific thing I can say to make it at least concrete, which is to say we have this algorithm called continual backprop. We published in Nature a couple of years ago, and it is exactly like backprop, but you also plant new seeds of units that are newly initialized with random weights. Backprop only has random weights at the beginning of time.

41:40Khurram Javed:And then as you go on, all that randomness, all that variety from the randomness gets used up. And with continual backprop, you keep injecting a bit of randomness, a bit of generate and test, a bit of generate. And then the operation of backprop is the tester. So you need that. And if you put those together really well, I think you'll have a new generation of massively superior continual deep learning. And that's what we hope to do in the next couple of years.

42:10Rich Sutton:Wonderful. Do you think that these algorithms can be applied to the current state of affairs with people scaling LLMs and trying to get them to do continual learning without catastrophic forgetting? Yeah, absolutely. I think it's, so I don't think that you could take an existing model and say, I'm going to just start updating it with these algorithms because these algorithms meta learn how to learn. So really you have to say, I'm going to learn from scratch. So let's say I learn a new foundation model, but I'm going to learn with these new algorithms. These new algorithms, in addition to learning the knowledge, they're also going to learn how to learn future things.

42:48Rich Sutton:So they're learning two things at the same time. And then I think you would be able to learn new things without catastrophic. Is that the most radical thing that you're trying to do in your company from the current state of affairs to try to do these two things at the same time? most radical thing. I think that's... This goes back to I'm not crazy, everyone else is crazy. That's perhaps not totally radical. There was a point in 2016 to 2018 where a lot of people were exploring these ideas quite a bit. They were doing it in a much more limited setting. So they would say, we have a distribution of problems.

43:27Rich Sutton:And then in this specific case, we'll do it. Whereas we want to do it from a single stream of experience. So our method should be more generally applicable. So I think many people have explored this, but no one has explored this in the general setting where the resulting algorithm would be applicable everywhere.

43:43Khurram Javed:So what would be the most radical thing that your company, your new company is trying to do that other people are not doing? What's the most ambitious thing? Remember, I don't think I'm weird, so I don't want to say it's radical. What's the most ambitious thing? thing, I think, is to try to have the full spectrum of knowledge, both about the tiny things and about the big things. You know, like thinking about how you take an airplane from one city to another. That's a very big, you know, it's more, it's like your space flight example, but it's just kind of more common sense to think about because we all, many of us take airplanes and all of us use abstractions on all kinds of our life.

44:26Khurram Javed:And even the squirrels use abstractions. So to have that spectrum of knowledge from the small to the big and to treat it in a uniform way and to be able to have it self-maintaining. You know, the big question is always, you have your knowledge-based system and what keeps the knowledge in it correct? Well, what keeps the knowledge correct in a large language model is, well, people did a lot of post-training and they made it sure it was correct, and then they freeze it after that. So that's what keeps it correct. But really, our minds, we are always changing things, and yet something keeps it organized and coherent and settling back into a good place rather than drifting off into crazy land.

45:15Khurram Javed:And that is, I think, our biggest ambition, to have a mind that is self-consistent and can keep training itself and make it coherent.

45:26Rich Sutton:I love that. Can I ask? It almost seems that it's such an ambitious vision. And the idea that all these things can be unified into a single mind is so ambitious.

45:37Khurram Javed:It's within reach. I think it's within reach. Here it's 2026. Yeah. And our computers are so fast. You know, is it so ambitious that it's out of reach? Or do we have already inklings of how all the steps can be done? I think we have a vision and inklings. I don't think it's inappropriate.

46:02Rich Sutton:Your vision involves a trillion parameter model with 20 watts. That seems pretty ambitious. That is ambitious. In some sense, with current technology, I would say it's also impossible. Just storing a trillion parameters in memory would probably use more than 20 watts of energy with current memory technologies. But we are really thinking of, okay, things are getting better. Computation is getting cheaper. It is getting more energy efficient. So where would be in 5 to 10 years? And I think five to ten years with the right algorithms, and we can totally be in a world where this would be possible.

46:43Khurram Javed:So five to ten years is two orders of magnitude of Moore's Law. It's a standard improvement if we double every 18 months. Ten years would be two orders of magnitude. And so for Kerm's statement to be plausible, then today you should be able to do it for 20 watts, 200 magnitude, 2 ,000 watts. If you can do it with 2 ,000 watts today, then in 10 years you'll be able to do it for 20 watts.

47:14Rich Sutton:I think you can do it for 2 ,000 watts. You have lots of people at research labs that have access to way more than that. Yeah, I think we can be more efficient than that even now with the right algorithm. If we can be more efficient than that, then why aren't we? It's not like people just want to spend all their money, spend all our money.

47:37Rich Sutton:Sometimes it seems like they want to.

47:39Khurram Javed:Yeah. Doesn't it? That's how they show they're real men, by using lots of energy.

47:45Rich Sutton:at least when i look at different research groups i don't even see anyone believing in it's possible and i think if you don't believe in it you're just not going to work on the technical problems and work through them is it that it's not possible or it's that there's so much waste in the system like one which one is it is is there someone who knows how to do it efficiently yeah and then there's 10 times the number of people in the same lab doing all these other things and so nine out of 10 people are wasting? In some sense, the way I think about it is that we are stuck in a local minnow. So if we want to move towards this new kind of algorithms, it is almost impossible that things will not get worse before they get better.

48:27Rich Sutton:So when we start exploring these new directions, you're not going to get state-of-the-art performance from day one, because it is a different paradigm. But that path leads to similar performance at a higher energy. And these big labs, they are so locked into a product that it is not possible for them to pursue a path where things get worse first. Because their current paradigm allows them to keep scaling. And this new paradigm, they have to take a bet. And they have to figure out some technical things that are difficult that we have thought about it for many years. We know people who have thought about these things for many years.

49:05Rich Sutton:And when I talk to them, it makes sense that it's doable. But you need to think about those challenges for a long period of time. So if everything goes right with Oak, what happens with the company? What kind of company are you building?

49:20Khurram Javed:If everything goes right, we implement the architecture, we can have genuine continual learning, and we can form abstractions so that we can do planning and reasoning. and we have sort of like true intelligence. And then, you know, it's hard to imagine just exactly what will happen by then. But I think... Humans will become irrelevant? I don't think that's true at all. We don't either. I think the world becomes exciting and even more exciting and interesting for humans. But in particular, I think there are the... You have to wonder about the large language models. They might be at risk. when this eventually happens, you know, I'm sure they'll get a good run.

50:06Khurram Javed:They've already had a good run. You know, they've been very successful. And let me say, just for clarity, that large language models are an amazing scientific breakthrough, a breakthrough in the skillful use of language by neural networks. It's totally unanticipated. You know, it was always a holdout for symbolic methods in language, and they have totally changed how that's thought about now. Yeah. So it's a big breakthrough. So it's frustrating to me that we have to, you know, just celebrate that we've made this great progress in the subset of the problem of AI and enjoy that. Instead, it has to pretend to be all of AI.

50:46Khurram Javed:All of intelligence is not fluid, capable use of language. There's so much more. It's an important part. You know, it's like 20 % or a quarter of intelligence. There's more. We're not done.

51:01Rich Sutton:Yeah. If everything goes right, are you imagining that there's a single mind that can do everything from learn how to swing from tree branches to make a spaceship to all these various things we've talked about today? Is it a single mind? And is it a single set of weights that can do all these things? Or is it...

51:22Khurram Javed:It's a single design. Okay. There'll be many minds.

51:26Rich Sutton:Okay, so it's a single design that reacts to different environments. And different versions of that mind would learn different things because they have different experience. This sort of goes back to the big world hypothesis that there are infinitely many things to learn. So one system cannot learn infinitely many things. And I think Rich already mentioned this, but if you have two of these systems, So if you have two of the largest systems in the world, then it is trivial that they cannot model each other because they're equally complex. So a single system would never be able to get to a point where it can learn everything.

51:59Rich Sutton:It would always be multiple systems that are learning from their own experience. You guys are hiring. What kind of people are you looking for? We are hiring. The initial team, most of it, we already have in our mind. So these are people who have thought about these ideas for a long time. and we are going to take a slightly different approach because this is a different paradigm. It doesn't make sense to become large very quickly because in some sense, everyone we hire has to come to see what we see and not everyone sees that. So we're going to start small, slowly grow to maybe a handful or two or three and then go from there.

52:40Khurram Javed:We want to be super aligned. We want to be super aligned. So that we can be very productive working together and scaling the progress.

52:50Rich Sutton:Absolutely. Very, very cool. Wonderful. I love this conversation. Thank you for taking the time to share what you're up to. You are an unusually deep thinker about where reinforcement learning and algorithmic design will go. And it was a true pleasure to get to explore it together with you today. So thank you.

53:10Khurram Javed:Thank you very much. Thank you. It's our pleasure.

53:20Thank you.

From the publisher

Rich Sutton, who helped pioneer reinforcement learning and wrote the seminal AI essay The Bitter Lesson, has now cofounded Oak Lab with his former student Khurram Javed. Their goal: to build agents that continuously learn from their own experience rather than from us. Rich doesn't think he holds a radical view: "I'm not weird. The field is weird." He says all learning is continual, and the field is the one that needed a new name for it. Rich and Khurram argue synthetic data is "a big mistake." Their "big world hypothesis" is that the world is massively more complex than any agent or simulator, so approximations have to be updated continuously rather than frozen at deployment. Rich calls LLMs an unanticipated scientific breakthrough, but says they represent roughly a quarter of intelligence. He says catastrophic forgetting is "totally curable" with the ideas behind their continual backprop algorithm. Khurram explains why the frontier labs can't follow: they sit in a local minimum where a new paradigm gets worse before it gets better. Their target, five to ten years out, is a trillion-parameter mind that keeps learning, stays coherent, and runs on 20 watts.

Hosted by Sonya Huang and Alfred Lin, Sequoia Capital

More from Training Data

All 110 episodes
Rich Sutton and Khurram Javed: Why AI Models Stop Learning, and How to Start It AgainTraining Data · 54 min
Listen in VO