#170 Richard Sutton on Pursuing AGI Through Reinforcement Learning

22 Feb 2024 · 56 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. - Episode #170: Richard Sutton on Pursuing AGI Through Reinforcement Learning

Podcast Overview Host: Craig S. Smith Guest: Richard Sutton, Professor at the University of Alberta and research scientist at Keen Technologies Episode Focus: The pursuit of Artificial General Intelligence (AGI) through reinforcement learning and insights from the Alberta Plan for AI development.

---

Key Topics Discussed

  1. Introduction to Richard Sutton
  2. Renowned scientist in artificial intelligence with a focus on reinforcement learning.
  3. Discusses his research and contributions to the field, particularly in policy gradient methods and temporal difference learning.
  1. The Alberta Plan
  2. Objective: A five-year research agenda aimed at building embodied agents capable of learning and interacting with their environment.
  3. Methodology: Emphasis on reinforcement learning as a means to develop AI that can set and achieve goals through environmental interactions.
  1. Core Concepts in AI Development
  2. Computational Power:
  3. The evolution of AI is heavily influenced by the exponential increase in computational power, as per Moore's Law.
  4. Sutton emphasizes the implication of this growth on various scientific fields, not just AI.
  • Reinforcement Learning vs. Supervised Learning:
  • Sutton discusses the distinctions and relevance of both learning methods in the context of AI development.
  • The conversation touches on how supervised learning has historically overshadowed reinforcement learning, despite the latter being crucial for building autonomous agents.
  1. The Horde Architecture
  2. Sutton introduces the Horde Architecture, which consists of multiple agents (demons) working on different subtasks simultaneously, all contributing to the overarching goal of the system.
  1. AI and Human Intelligence
  2. Intelligence Augmentation (IA): Sutton posits that the ultimate goal should be to enhance human intelligence using AI, rather than replacing it.
  3. Sutton critiques current AI models (e.g., large language models) for lacking true understanding and goal-directed behavior.

---

Insights from the Conversation

Importance of Goals in AI

  • Sutton asserts that any intelligent system must have clear goals and the ability to understand and prioritize them.
  • This understanding distinguishes effective AI systems from models that just react to stimuli without true comprehension.

Ethical Considerations and Societal Impact

  • The discussion also touches on the ethical implications of AI development, dismissing the notion of a "doomsday" scenario often discussed in media.
  • Sutton argues that AI is a tool that can be used for good or ill, depending on human intentions.

Future Prospects for AGI

  • Sutton and Craig discuss the ambitious timeline of achieving AGI by 2030, with Sutton suggesting a 25% chance of success.
  • The need for continued advancements in reinforcement learning and embodied AI is emphasized as critical to realizing intelligent systems that can interact with the world effectively.

---

Conclusion The episode concludes with a reflection on the rapid advancements in AI and the vital role of understanding human intelligence to develop systems that complement rather than replace it. Sutton's insights provide a hopeful perspective on the future of AGI, stressing the importance of responsible development and the potential for AI to enhance human capabilities.

---

Key Takeaways

  • Computational Power: Essential for AI evolution, impacting all sciences.
  • Reinforcement Learning: A critical avenue for developing intelligent systems that learn from interaction.
  • Horde Architecture: Concept for decentralized problem-solving in AI.
  • AI & Ethics: AI tools can be beneficial or harmful based on their application.
  • Future of AGI: Aiming for AGI by 2030 is ambitious yet achievable with the right focus and research.

---

For more insights and a transcript of this episode, visit [Eye on A.I.](https://eye-on-ai.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00There isn't a science around that isn't profoundly influenced by the availability of massive computing power and just greater regular computing power. It's the story of our age. It's not just the story of AI. The idea is to leverage computation, to make useful things and understand the mind. All these things need a lot of computation. It's the fact that computation is becoming more plentiful and cheaper exponentially for on the order of 100 years and can be expected to continue going that way. It looks like doubling every two years, now every 18 months, and that keeps happening 18 months after 18 months after 18 months, and it means you double and you double and things get qualitatively different every decade.

0:41And that's happened for a long time, for many decades, and will happen more so in the future. So we have that to look forward to. I think it's what we really should mean when we say the singularity. The singularity is that we have this exploding, it's a slow explosion of computer power and that that is fundamentally changing things. Hi, I'm Craig Smith and this is Eye on AI. In this episode, I speak with Richard Sutton, the father of reinforcement learning and professor at the University of Alberta. We discussed his cooperation with John Carmack on Keen, a startup that vows to reach artificial general intelligence by 2030.

1:19Richard also talked about the Alberta Plan, his ambitious five-year research agenda focused on building embodied agents with the capability to learn and plan through interactions with their environment. Sutton provides insights into the current state of progress, new algorithmic developments, and trade-offs between simulated and physical environments in training, and the ultimate goal of creating AGI. I hope you find the conversation as amazing as I did. So why don't you start by introducing yourself. I assume people know who you are. I've had you on the podcast before. But for those new listeners, tell us who you are, where you are.

2:07And then we'll talk about the Alberta Plan, which I find pretty exciting. Thank you, Greg. Well, I'm Richard Sutton. I'm a scientist. I've been studying artificial intelligence for like 45 years, a long time. And I'm up in north at the University of Alberta in Canada. And I'm a professor in the computer science, computing science department. And also I'm a researcher at Keen Technologies. And I got lots of lots of titles and sub roles. But basically, I'm just trying to figure out how the mind works. And I've tried to do it in a very broad interdisciplinary way, reading all the different thinkers on the subject.

2:52And just from the point of view of psychology and how the brain might work as well as from the point of view of computing science. Yeah, I've read a number of the recent papers and I can see this thread developing. And I don't know whether it's just that you're writing more and so the thoughts are more developed in print or whether they're developing in your mind. But from 2019, when you wrote The Bitter Lesson, you talked about the idea that it's really increasing computation and it's driving a lot of things, a lot of progress. That kind of coincided with OpenAI's scaling of the transformer model.

3:44I talked to Ilya Sutskova and I asked him whether your essay had triggered their interest in scaling. And he said, no, it was coincidental. incidental but uh I kept first can we talk about that about how scale scaling and and the availability of computational resources and Moore's law has driven a lot of what's happened in artificial intelligence research uh almost more than novel algorithms well I think the first thing to be aware of is it's been driving things that are not just artificial intelligence. It's been driving all the sciences and all the engineering developments in the world. There isn't a science around that isn't profoundly influenced by the availability of massive computing power and just greater regular computing power.

4:45It's the story of our age. It's not just the story of AI. It's not particularly the story of AI. AI has always known that it needs computation. The idea is to leverage computation to make useful things and understand the mind.

5:06Yeah, now it's true that those of us who are interested in connectionist systems or distributed networks, nowadays just called neural networks, not particularly good terms, I always shudder a little bit when I use it. But But those of us that have been doing that have, those of us that have been doing learning, I think that learning is important for intelligence. These, all these things need a lot of computation. And so they're, they are limited by the computation available at the time. Okay, so let's, let's be, what is this thing? What is the, it's a Moore's law. Let's call it Moore's law. It's the fact that computation is becoming more plentiful and cheaper exponentially for on the order of 100 years and can be expected to continue going that way.

5:54So exponentially, it's like doubling every two years, now every 18 months. And that keeps happening 18 months after 18 months after 18 months. And it means you double and you double and things get qualitatively different every decade. and that's happened for a long time, for many decades, and will happen more so in the future. So we have that to look forward to. That will continue having a tremendous influence on everything that's done. On the other hand, it's just normal. It's just what you would expect. And those of you who worked on AI for a long time have just expect and planned for. And now it's coming, but it's an exponential.

6:44So exponentials are self-similar. So that means they look the same at every point in time. Every year, you're doubling in a year and a half. So it's an explosion, because every exponential is an explosion. It's sort of, I think it's what we really should mean when we say the singularity. The singularity is that we have this exploding, it's a slow explosion of computer power and that that is fundamentally changing things. Yeah, and I had a really interesting conversation almost a year ago with Aidan Gomez, who was on the team that designed the transformer algorithm at Google. And he now has a startup called Cohere.

7:27He's Canadian. And he said an interesting thing that he believes it could have been almost any algorithm. It didn't have to be the transformer, that the community got behind the transformer, poured resources into it, continued to scale it, and it was scalable. I mean, that was important, that it's a scalable architecture. But it didn't have to be the transformer. And that made me think of you, because um so transformers they the way he described it at its core it's a stock of multi-layer perceptrons with attention you scale it feed it data and it does learns to understand language or at least seems to understand language but it's got all these obvious limitations i've been talking a lot over the last couple of years to Yann LeCun, about world models.

8:32And that, to me, sounded like a much more exciting direction for general intelligence because not all intelligence is contained in language or at least most or even less so in human text. And then I see you guys come along with the Alberta plan. And that sounded even more exciting to me. So how do you, so the Alberta plan, you're building, the idea is to build an agent, ultimately an embodied agent, that has a world model or can create a world model through interactions with its environment. chairman how is that different from uh lucum's approach uh at a very basic level very basic level a good is that they're a very similar idea it's uh you look at the parts of his architecture and the parts of the architecture put forth in the alberta plan they line up one one for one um yeah we're trying to do the same thing um we're going about it slightly different and we could talk about that but i think to just focus on the differences might even be to distract from the big message the big message is that you have to have a goal and you have to have a model of the world and um and then everything is driven by using that model to take action and to plan action at various levels of abstraction in order to to achieve the goal okay so to me this is what intelligence is understand the world use your understanding to get to achieve to achieve your your goals I'd like to formulate the goals as as a reward and I'm super comfortable with that other people sort of grudgingly accept rewards even though it seems kind of low level that it's a natural approach, I think.

10:50I think it's something that almost makes more sense to people who aren't steep in deep learning and supervised learning. And one thing I found interesting in the roadmap that you've laid out for the Alberta plan, you start with supervised learning. And why is that? Is it just because it's easy? Yeah, I guess we do in a sense because we want to focus on, well, continual learning, learning continually, which is sort of an obvious thing. You almost, what learning means, it has something that goes on at all times. But the first steps, getting continual learning with nonlinear networks is still challenging, even for supervised learning.

11:33And so it's natural to start at the simplest possible case, which involves the fewest other factors. And that's a supervised learning case. Yeah. Yeah, it's funny. Let me just say a few words about that because there's sort of been a fight through a struggle throughout the decades between supervised learning and reinforcement learning. You know, there's only so much oxygen for learning methods and all the attention that's paid to supervised learning somewhat detracts from reinforcement learning. So there's a bit of a friendly competition. And supervised learning has always won the competition because supervised learning is so much more easy to put into practice and for people to use.

12:19And it's sort of less ambitious, but it's really important. And really, those of us who do reinforcement learning or try to make whole agent architectures, we are consumers of supervised learning algorithms. we will use them as components of our overall architecture. So we need them and we can work on them and we need to structure them for our purposes. But I saw one of your talks, you make a distinction between AI tools and AI agents and supervised learning falls into the tool category. Can you sort of start and talk about the evolution of the Alberta plan and then present to listeners what it is in its simplest form.

13:07And that'll give me a structure on which to hang questions. The Alberta Plan is an attempt to understand intelligence as primarily a learning phenomenon, assist something that comes to understand its environment and and then drives the environment to achieve goals so the first step in the alberta plan is the structure between the agent the environment and their interaction form the interaction there's not exchanging states you're exchanging observations like sensor sensors visual touch auditory um it's all abstract to those particulars but it's got to be genuine observations and not state because state we don't we don't really have access to directly uh so that you know the principles number one principle i'm trying to remember them as i speak but number one principle is this this agent environment interaction is sacrosanct uh and number two is that learning uh or everything is is we could say continual I think we call it we say temporally uniform uh temporally symmetric in in the Alberta plan which means that there are no special phases where you like training and test there's just life goes on and on you get rewards uh or you don't get or you or you don't get the reward you want and you get your observations and there there is no teacher other than um rewards, pains, and pleasures.

14:48And maybe I'm not getting the four principles right, but another important point is that you are going to be forming a model, and so you're going to plan. It's both trial and error learning directly from experience and learning a model and then planning with the model. Both of these are an important part of intelligence. Okay, so that's the background. And then we outline their 12 steps. And the 12 steps really start with, let's have learning that is temporal uniform. Let's have meta learning. And meta learning, maybe I should stop on that for a moment. Meta learning means learning to learn, not just learning one function.

15:31But once you are continually learning, you're learning this and you're learning that, you get many, many experiences learning. And you can get better at learning. You can use those repeated experience with repeatedly learning to make future learning episodes more efficient. So as part of that, you learn representations, you learn features, you learn step sizes. Okay, so continual learning and then all the algorithms. And once we add meta learning and continual learning, we have to and supervised learning. Then we extend that to reinforcement learning, which involves its own set of issues, more interesting temporal relationships.

16:17And I think like the first six steps are crafting the basic algorithms of reinforcement, working through them again to be continual and meta. And then we start to bring in the challenging issues like learning off policy and learning models of the world and then planning. And just to jump to the end, the last step is about AI, IA. Opposite AI is IA, intelligence augmentation, where we combine computers, AIs, with our own minds to make our own minds stronger. okay now one of the key steps in there was off policy learning and learning a model of the world um off policy learning means you want to be able to learn about things that you're not doing or you're not because uh or not doing all the way to completion so even like to recognize an object you you look at the object and you say how would you you have to define that in some objective way and the best way to to just do that is as a sub problem so yeah maybe maybe i'll i'll just sort of stop there the i the most interesting uh strategy uh distinctive strategy the alberta plan is the pose is that the mind works by posing sub problems for itself and then working on them and it's it's not it's sure it's got a main problem which is to get reward, but it also has many thousands of sub-problems it's also working on simultaneously.

18:05And since it's not behaving, it cannot behave for all thousand problems at once. It has to pick one problem, like perhaps the main problem, and behave according to that. So all the other things have to be able to learn from data that's not exactly on what they would do. And this is called off-policy learning. And it's a key to learning to achieve auxiliary sub-problems. And also it's a key to efficiently learning a model of the world. Yeah, you have something called the horde architecture. Is that where that comes in when you break a problem down into multiple subtasks that you learn? There was one paper where we worked on that idea.

18:50We developed that idea. The horde is the horde of subproblems. Each demon in the horde, which could be almost viewed like a single neuron in a neural network, is working towards a different task, trying to predict a different thing, or maybe trying to attain a different thing. it's the view of the mind as decentralized there is one goal and everything is ultimately driven towards one goal but still it's a useful structure to have different parts driving towards other goals. How did you get together with John Cormack? Was that primarily because you need the funding and it gives you a vehicle to raise capital?

19:37No, seriously. I mean, you know, Dan Lacoon's got meta behind him. Well, it's just not really comparable. John's company is great, but it's still like a$20 million company, which is plenty of money for what we want to do now. John and I got together because we had similar ideas about what was needed and also what was not needed to get to AI or AGI. Yeah, so I read a newspaper article, an interview that John did down in Texas, and I just could see that he was thinking about the way, thinking about things the way I was, even though our backgrounds were quite different. He thought of intelligence, you had to, there's a few principles that needed to be worked out rather than so it isn't a huge program to write it's a few principles we have to figure those out um not that many maybe a maybe 10 000 lines instead of 10 million lines of code so it's easy to get it's relatively it's still it's still it's still hard to get basic research funding in the world it's easy to get funding towards applications of ai large language models particularly.

21:00Anyway, I've really enjoyed working at Keen and being able to focus on the ideas. And it's a calm company.

21:15There's a lot of thinking involved, a lot of contemplation. There is also experiments and we're trying to get the engineering side of it is really important. but for me it's been really great just to be able to you know regroup my thoughts and think about them very carefully and push them forward but keen is is implementing the alberta plan is that right i mean that's that's uh the project well the alberta plan is a research plan it's like a five-year research plan and so research is something you don't implement research is something you conduct and and it doesn't always end up the way you want but um yeah yeah i wouldn't say implement is the right word not yet but but but the the work you're doing at keen is is informed by the alberta yeah i'm absolutely i'm working on the alberta plan uh uh and and the end goal at keen is to create the embodied intelligence described by the Alberta plan.

22:20You don't sound very confident. Well, a plan is just a plan. And I think there's a good chance that it will work out as planned. But a five-year plan, you make another one after four or three years. Yeah, so I wouldn't presume to know how it's going to work out. But at the same time, we have to make our bets. We have to think hard about it, just knowing we may well be right. You know, your work is primarily in reinforcement learning. You wrote the book on reinforcement learning, temporal difference learning, and lambda, and all of that. is this I mean this is this seems a much more ambitious project is this was it the success of the transformer scaling that said well you know let's do that with RL why are these guys you know everyone's celebrating what they're doing but there's much more to be done no So what you see in the Alberta plan is perhaps bigger than the book, but this has always been the plan.

23:44We've always, in AI, tried to understand all of the mind and reproduce it in computers. And so that is a big, enormous ambition. That's what it's always been. So the large language models are a bit disappointing in some sense. I mean, it's really good that people are getting excited and people are wanting to learn about it. But it's not, I don't envision that it's the direction that will be most productive to pursue. Now, you know, who knows? What I do know is it's not the most direction that's useful for me to pursue. I'm much more interested in actions and goals and how an agent can tell what's true and what's not true.

24:35All of those things are missing from large language models.

24:41So, no, they're not really – what they are doing that's important is they're showing what you can do with computation and networks and learning. that you can get enormously complex things and you can incorporate a lot of data. It just shows the power for those who needed to be shown that. And it could be an interface between humans and whatever you end up creating, the agents you end up creating. You still need a language interface to communicate. Yeah, but I doubt that what we're doing with large language models today will contribute to that. Oh, is that right? Yeah. In other words, the models that you want to build, the agents you want to build, would learn language as part of the learning process.

25:42Yeah, so it's like we say, language is last, not language first. Large language models are language first. We just say language is last. just as jan lakoon says we need to do you know rat level intelligence and then cat level intelligence and we have to get those figured out before we should try to make human level intelligence uh so where are you on the plan i mean you you figured out reinforcement learning you can build agents uh you there are various architectures for creating representations from various kinds of sensory input and at that representation level, then you can plan efficiently.

26:32So where in all of that are you in your research? Well, it's a little hard to explain non-technically, but you can say some things. certainly you can say that the various steps are not done entirely sequentially. You're always looking for areas of opportunity where you can make an increment of progress. And those could be, you know, on step 10 or they could be on step three. But I could also try to be very rough and say that we're at about step four now. um we are still doing things where we're changing the basic underlying uh fundamental reinforcement learning algorithms we are not done with that we need more efficient algorithms and uh i'm excited about some of the changes uh new ideas we're developing recently about how that can be done can you talk about those new ideas at all okay well one of the big things is efficient off policy learning and the use of importance sampling.

27:40Important sampling is where you see how likely you're to do things under your target and your behavior policies and you adjust the returns based on those ratios of those two. And for a long time I thought that was the only way to adjust the returns. But now the forward correction of the returns I think can be done by changing your expectations. So like if you're expecting a good thing to happen, you're expecting a good action to be taken and then a different action was taken, a more exploratory action. So this is a deviation from your target policy, which would be more greedy. And one way to take into account the deviation from the target policy is to just say, oh, okay, now i've done something not best so i'm just going to adjust my level now you can expect a little a little less and there's a way there's a systematic way of doing that that's gives us a new way to handle the off policyness of of our returns and so this gives a whole new family of algorithms so that's exciting now for exciting maybe mostly for me I think maybe the most accessible direction of excitement, of novelty, is in continual.

29:07So I'm going to say a bunch of things. And to me, they're all going to have the same solution. Continual learning, meta learning, representation learning, learning to learn, learning how to generalize, how to construct a state representation, feature finding. that whole thing is coming, and it will be a kind of, it's just a new kind of way, a new kind of way of doing learning in deep networks. And I call it dynamic learning nets. See, dynamic learning nets have learning at three levels, whereas usually our neural networks only learn at one level. They learn at the level of the weights. And in addition, we also want to learn at the level of step sizes.

29:52So all of every place you have a weight in your network, you're also going to have a step size. So a step size is sometimes called a learning rate. It's much better to call it a step size because a learning rate will be influenced by many other things. So if we imagine a whole network, all these weights, next to each weight is a step size that is adjusted by an adaptive process that's adapted in a meta-learning way, a meta-gradient way towards making the system learn better rather than just perform better at an instantaneous moment in time learning rates or step sizes don't affect the function they don't affect the function implemented at a particular point in time they don't affect what the network does they affect how what the network learns and so if you can tune the step sizes you also get learning to learn and learning to generalize well and things like that the last three the last element that we wanted to have be adaptive weights step sizes the third one is the connection pattern so who's connected to who and so this will be done by an accretive process like let's say you start with a linear unit and it learns say a value function or a policy and it does the best it can with the features available and and then it needs to induce the creation of new features because you need to learn a nonlinear function of your original signals and so you need to create new features that have become available to that linear unit and in this way you grow in a sort of organic way a system that can learn nonlinear functions so this is just a different way of ending up with a deep network that was all learned including all the features dynamic learning nets.

31:42Where is the data, the input data coming from? Well, the input data and reinforcement just comes from life, from doing things, seeing things. Right. There is no label data set. Yeah, maybe I should have said this from the very beginning. The whole idea of, I call it experiential AI, is that no one makes you data. You grow up as a baby and you play with things and you see things and you do things. And that's the data. And the trick of reinforcement learning is how do you turn that kind of data into something you can learn from and grow a mind from? So the beauty and the limitation of supervised learning is they say, well, let's not worry about that for now.

32:22Let's assume that somehow we have a data set with labeled things. And let's work on this sub-problem. That's a great idea. Work on a sub-problem, figure it out, and then move on to the next thing. But really, we have to move on to the next thing. We have to worry about how the data set, quote data set, is automatically created from the training information. There is never a data set. Data set is such a misleading term. It suggests that it's easy to have this thing and store this thing and curate this thing. Like, really, life is full of you do things, things happen, and then they're gone. You know, everything is fleeting.

33:00You don't have a record of it, and it would be enormously complex and not all that valuable to have a record of it. The feeling is totally different in reinforcement learning and supervised learning. And particularly the way I would address it. You know, many people do reinforcement learning by creating a buffer or a record of all the experiences that have been retained, that have been occurred at least for some period of time. And I think that's an appealing, but it's not where the action, the answer is. The answer is embracing the fleeting nature of data and making most of it when it happens and then letting it go.

33:47Well, that's why you want to make an embodied system so that you have all the five senses or more? So you need, as you say, an embodied system, an interactive system that influences its input stream, its sensory stream, and that you get that interaction for a long period of time. You can do this in simulation or you can do it in robotics. There's still, I still know what's the best way or if the best way is to do both or maybe first one and then the other. John is interested in having learning from video and he likes his his his his view of the experience is you have massive numbers of video streams like you're viewing you know 500 channels of television and then you can switch switch to look at one look at another one other people in Keen my close colleague Joseph Modile, he's interested in robotics and he thinks the best way to get an appropriate data stream is to actually build robotic hardware.

35:01You know it's important that the world be large and complex because the worlds we want to address are large and complex. And so you want things like video and you want large data streams. Now you can use simulations to generate even video streams, simulated video, but inevitably those simulated worlds are really quite simple. They have an underlying simplicity. They have objects perhaps in three-dimensional straight structure, maybe they're rigid objects, and the vision is a very particular geometric form. They are generated and they are made up worlds and they're generated so they're they're really the worlds are are less complex than the agent their goal would be to have spend most computer power working on the mind and just a little bit to create the simulated data and and that's that's the reverse the way it really is right every person is maybe has a has a complex brain but their world is much more complex not just because the world consists of all these physics and matter, but it also consists of other minds, other brains and other minds out there.

36:16And what goes on in their minds matters. And so the world is inherently vastly more complex than the agent. And we reverse that when we work on simulated worlds, which is always concerning. Anyway, those are some of the issues and the trade-offs between working with simulations or with physical worlds. nonetheless you need to develop the architecture and the algorithms before you worry about the data data stream i would think yeah but you want to develop the right algorithms and if you're working with the world that's not representative of your target world in an important way it can be misleading but you're right and that's what we that's what we strive to do you know i know if you know, but I think in my own work, it's almost always I want to focus on some issues.

37:06So I make a really simple instance of that issue, like, you know, a five-state world. And I study the hell out of it, but I don't like try to take advantage of its smallness. You know, I study algorithms that are in some sense even simpler than the simple world. And I stress those algorithms and see what their abilities are. So we always, you know, it's always part of research is we We simplify the world, understand it fully. Just like a physicist might make a simplified world with a ball rolling down a ramp. And it's a really simple world. And you try to eliminate the friction. And you eliminate other weird effects and just see things in their simplest form.

Read the full transcript

37:47Yeah. Have you paid much attention to Alex Kendall's work at Wave AI? Do you know that company? It's an autonomous driving company. they have a world model called gaia one uh and it's it's it's similar to what uh yanlokun's doing it it you know encodes uh representations from from video from live video and then plans uh based on those representations and it can control a car from the representation space, it's actually pretty remarkable. So let's talk about the world model and what kind of world model would be appropriate for autonomous driving.

38:45so let me say some things that are mistakes very natural seeming but mistakes in my opinion the mistake would be to make like a physics model of the world or to try to make something that could simulate the world and produce the video frames you don't you don't you don't want the video frames of the future that's not the way you think um instead you think oh i could i could go to the market and maybe there would be strawberries okay you're not creating a visual uh a video you're saying you're like jumping to the market and then your strawberries could be you know different sizes and positions and and and still uh there's not a video there's an idea that will happen if you go to the market um so uh people have realized this like jan lakoon used to talk about um generating video of the future and then he realized it would be blurry and and now he realizes um that you you need to produce outcomes of your model that are not like not at all like video streams.

40:00They're not like observations at all. They're like constructed states that are the outcome of the action. Okay, so this is very different from a partial differential equation model of the world. And so it's very different from what self-driving car companies start with. Self-driving car companies start with physics and geometry and things that are calibrated by human understanding, engineers' understanding of the world and driving. But I suspect that's going to be – I mean, what do I know? I'm not into self-driving. I don't do self-driving cars, but I know that like Tesla is and Elon Musk is. And so their goal is to make some – they started like everyone else with engineering models.

40:54But I think my understanding now is that they're building sort of more conceptual models that are based on artificial neural networks. OK. And so rather than starting with geometry and understood things, they're just getting massive amounts of data and training it to make a model. We need a model that is at the level of high level consequences, not at the level of low level things like pixels and video. So one way you do that is by having state features that are at a more advanced level. You say, oh, this is a car rather than this is a video frame.

41:33And so and then basically it's as simple as you need abstraction in both state and time. Abstraction in state is like saying there will be strawberries when I get to the market. And abstraction in time is saying, oh, I can go to the market and then in 20 minutes I will be there. probably and other things will be the same or related in natural ways so we want to be able to think about I could go to the market you also want to think oh I could pick up the coke can I could move a finger and that will have certain consequences these these all these things that we know you think are vastly different scales going to the market is like 20 minutes you know taking it taking a new job you know might be a year uh deciding to study a topic also might be a period of time we think and we analyze the consequences like you wanted to meet with me today and you know we arranged it we set it up it was your planning uh took you know place over weeks and some cases months and and yet and we assembled the the event of this interview by by planning all that and exchanging high-level messages.

42:49All that, it's silly to think that that's done at the level of imagining videos that we might see with our eyes or audio signals that we might hear. So we need models that are abstract in time and state. And as a reinforcement learning person, there's a particular set of technologies that I naturally turn towards to do that. the prediction is based on multi-step prediction by temporal difference learning the planning is done by dynamic programming essentially value iteration but where the steps the are not low-level actions but they're called options they're high-level ways of behaving with that terminate so there are things like going to the market and they'll terminate when and you're at the market.

43:42So at a certain conceptual level, it's clear where we want to go, to me, with abstract models in time and state, built with options and features.

43:59We did write one paper recently published in the AI journal on the notion of planning using uh sub problems on the stop progression stop means subtask option model and planning put all those things together and you can do the full progression from from the data stream to abstract planning and that's that's what we're trying to put together yeah yeah and i i sort of misspoke talking about gaia one about that model i mean it they they the input is video uh it creates a representation and it plans and and and takes action in the representation and plans actions in the representation space you can then decode that into video to see what it's doing, but you're not planning in the video space.

45:01So what's your ambition with this? You'll figure out to refine the algorithms, the reinforcement learning algorithms. They need to be scalable. Once you have that, then you move on and start scaling them with compute and following your roadmap? Or am I simplifying it too much? You know, we want to understand how the mind works, and then we're going to make a mind, or some minds, or some mind, amount of mind. And it will be useful in all ways, in all sorts of ways, economically useful. It will also be useful to us to extend the capabilities of our own minds. if we can understand how our minds work, we can augment them so that they can work better.

46:02Yeah, we're gonna... the key step is understanding, and then there would be millions of uses. I don't think it's going to be as simple as making workers sort of like slaves for us to direct. I don't think it'll be as simple as that. That maybe gives a lower bound on potential utility. Some of our story for at Keen is we say that, well, suppose you could make a virtual worker. This would be enormously useful. Much of the work that we all do from day to day doesn't require a physical presence. It doesn't require a robot. Much of which we do is just shuffling information around. We can do most things through a video interface.

46:55So why can't we make workers that are extremely useful by playing the roles that people play in many cases? That's sort of a lower bound in what can be done. I think much more can be done, and there'll be much more interesting things to be done.

47:13And then there's a question of what should be done. um yeah those are those are rich philosophical questions and practical questions for the economy yeah uh the the uh i've seen your uh well and one thing on reinforcement learning and sort of supervised learning sort of took over for a while now it's transformer-based generative uh ai but uh during the supervised learning phase uh the argument was that uh higher knowledge is all supervised learning and and it's still supervised it's still supervised in general ai large language models they the the training information is the next token the next word and that's taken as as the correct action the analogy you gave me was uh you know because the analogy that that's always given is that you know a child sees an elephant the mother says that's an elephant and the child very quickly can generalize and and recognize other elements elephants maybe it makes a mistake and the mother corrects it and says no that's a cow and and that was always given as an example of supervised learning but maybe it's reinforcement learning maybe it's the child's reward from the mother praising him for remembering the label the point is a child has has well-developed concepts classes concepts um before and then and then when it's you know when its mother says that is an elephant uh there's already an extensive understanding on the child's and a part of you know what the space is what the objects are and and this this the thing that is being labeled.

49:17The label is the least interesting part of that. And the child has already learned all the other most interesting parts of what it means to have animals and moving things and objects in its world. The label is the least interesting part of it. You were talking about agents that could be virtual workers. Already, reinforcement learning people are building agents and using large language models and knowledge bases to you know carry out tasks knowledge-based tasks so what you're talking about is is more than linguistic tasks or knowledge-based tasks. You're talking about physical planning and physical tasks.

50:13Is that right? The key thing is having goals. And if you have, for example, an assistant help you plan your day, organize your day, or do tasks for you, I'm thinking it's very important that the system is able to have goals and is able to understand your goals, I think it's probably the most important part of an assistant is to understand the purposes involved. And large language models don't understand, don't really understand the purposes involved. They will appear to a little bit, but the corner cases always come up. And once you spend a bit of time. You're always in a corner case. And so an AI system is a system that after a bit does silly things that don't respect the goals that you have or that have been given to it.

51:11That's not going to be a useful assistant. So I mean, I don't want to be critical of large language models. They're very, very useful, but it shouldn't be viewed as a criticism to say that they're also at the same time have rather important limitations. It's not a competition in that sense. Are you concerned at all? Do you ascribe to the threat debate? No, I think the, I don't, I don't, I think the doomers are, they're not just wrong. I think, I think they're blindingly biased. The bias is blinding them to what's going on. Basically, AI is a broadly applicable technology. It's not like, it's not like nuclear weapons.

51:55It's not like, like a bio weapons it can be used for all kinds of things and it's not it's not uh it's uh the way we deal with such things is we we uh we try to use them well and there are there will be people that use them uh for bad things and then you know this is just normal there's normal technology is it can be used by good people or bad people the the doomers the doomers are just saying, oh, somehow there's going to be, it's going to be, that it's bad in the same way that nuclear weapons are bad. That's just, they're just blinded by that metaphor, by the thinking that the AI will be out to kill them.

52:41That's just, it's just silly. And I don't think, well, the doomers don't actually give coherent reasons for what they believe. And so it's hard to argue with them uh so maybe it's fair just to hold that they're they're biased and blind i don't accept i don't accept an argument unless it's a proper argument uh so so where you say you're maybe at stage four in the research uh carmack says 2030 uh that's you know it's far enough out there that maybe people won't remember in 2030 that he said 2030 uh it's always uh you know i've 2030 has been out there for a long time and it's it's it's uh you can't it doesn't recede it's always been 2030 for the uh computer power reaching human scale um quantities yeah but anyway 2030 is is a reasonable reasonable target for us understanding everything that we need in order to make a real mind yeah i'm good with that yeah that's you you have to be ambitious i've always said that 2030 is a 25 chance of of of achieving a real intelligence a real human level intelligence, 25 % chance.

54:13So probably not, but it's a big enough chunk of probability that an ambitious person should work towards it and try to make it true. And it does depend upon what we do and not just the unfolding of the universe. So we should try to do that. The big thing that's happening right now is the public is coming to grips with what it means for us to understand the mind and have the ability to create minded things. And so that is a big transformation, it's a big change in our worldview. And so we absolutely need all kinds of people to help us become easy and have an understanding of what's happening as we achieve human level designed intelligence.

55:08That's it for this week's episode. I want to thank Richard for his time. If you want to read a transcript of today's conversation, you can find one on our website, eye on AI. That's E-Y-E hyphen O-N dot A-I. In the meantime, remember, the singularity may be getting closer, but AI is already changing your world. So pay attention.

From the publisher

Join host Craig Smith on episode #170 of Eye on AI, for a riveting conversation with Richard Sutton, currently serving as a professor of computing science at the University of Alberta and a research scientist at Keen Technologies.

Sutton is considered one of the founders of modern computational reinforcement learning, having several significant contributions to the field, including temporal difference learning and policy gradient methods.

In this episode, we go through the Alberta Plan for AI development, the transformative potential of reinforcement learning, and the future of AI in augmenting human intelligence.

Richard Sutton shares insights on the importance of computational power, the impact of large language models, and the vision for AI that interacts with the world through goals and learning from its environment. 

We also explore the challenges and opportunities in making AI more embodied and goal-oriented, and how this approach could revolutionize our interaction with technology. 

A must-listen for anyone interested in the cutting-edge advancements in AI and its societal implications.

 

Don't forget to rate us on Apple Podcast and Spotify if you enjoyed this episode!

 

This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.

Download NetSuite's popular KPI Checklist, designed to give you consistently excellent performance - absolutely free at https://netsuite.com/EYEONAI



Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI



(00:00) Preview and Introduction

(02:15) AI's Evolution: Insights from Richard Sutton

(07:08) Breaking Down AI: From Algorithms to AGI

(10:50) The Alberta Experiment: A New Approach to AI Learning

(18:27) The Horde Architecture Explained

(21:23) Power Collaboration: Carmack, Keen, and the Future of AI

(25:04) Expanding AI's Learning Capabilities

(31:34) Is AI the Future of Technology?

(35:29) The Next Step in AI: Experiential Learning and Embodiment

(40:00) AI's Building Blocks: Algorithms for a Smarter Tomorrow

(45:59) The Strategy of AI: Planning and Representation

(49:27) Learning Methods Face-Off: Reinforcement vs. Supervised

(52:53) The 2030 Vision: Aiming for True AI Intelligence?

 

 

More from Eye On A.I.

All 266 episodes
#170 Richard Sutton on Pursuing AGI Through Reinforcement LearningEye On A.I. · 56 min
Listen in VO