#241 Patrick M. Pilarski: The Alberta Plan’s Roadmap to AI and AGI

7 Mar 2025 · 1 h 2 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode Notes

Episode #241

Patrick M. Pilarski: The Alberta Plan’s Roadmap to AI and AGI

Episode Description In this episode, Craig S. Smith interviews Patrick M. Pilarski, a Canada CIFAR AI Chair and professor at the University of Alberta. The discussion centers around The Alberta Plan, which outlines a strategy for achieving Artificial General Intelligence (AGI) through reinforcement learning and experience-based AI.

Key Themes and Concepts

  • The Alberta Plan: A roadmap to develop AGI by emphasizing continual learning and the integration of real-time experience into AI systems. This contrasts with traditional machine learning approaches that rely heavily on pre-trained models and massive datasets.
  • Reinforcement Learning: The episode discusses the superiority of reinforcement learning over pre-trained models for building intelligent systems that learn and adapt from their own experiences, akin to how children learn through trial and error.
  • Bionic Medicine: Pilarski shares insights from his work on AI-powered prosthetics that enhance human-machine interaction, revealing how AI can augment rather than replace human capabilities.

What You'll Learn

  • The advantages of reinforcement learning for AGI development compared to traditional pre-trained models.
  • Four core principles of The Alberta Plan:
  • Experience-Based Learning: Agents learn from their ongoing experiences in real-time.
  • Continuity in Learning and Acting: Learning processes are continuous, avoiding the separation of training and application.
  • Computational Efficiency: Every computation should be valuable to optimize the learning process.
  • Multi-Agent Interactions: The importance of considering interactions with other agents in the environment.
  • The significance of continual learning in AI to prevent catastrophic forgetting, ensuring that new knowledge integrates with old without loss.
  • Real-world applications of reinforcement learning in various fields, including plasma control and robotics.

Episode Breakdown

  • (00:00) Introduction to The Alberta Plan
  • (02:22) Guest Introduction: Patrick Pilarski
  • (05:49) Core Principles of The Alberta Plan
  • (07:46) Experience-Based Learning in AI
  • (08:40) Comparison: Reinforcement Learning vs. Pre-Trained Models
  • (12:45) AI's Relationship with the Environment
  • (16:23) The Role of Reward in AI Decision-Making
  • (18:26) The Importance of Continual Learning
  • (21:57) AI Applications: Fusion, Data Centers, Robotics
  • (27:56) AI Learning Like Humans: Predictive Models
  • (31:24) Learning Without Massive Pre-Trained Models
  • (35:19) Control Theory vs. Reinforcement Learning
  • (40:16) Future of Continual Learning in AI
  • (44:33) Reinforcement Learning in Prosthetics
  • (50:47) The End Goal of The Alberta Plan

Conclusion The episode emphasizes that the future of AI lies not just in accumulating more data but in creating systems that can think, adapt, and learn from their experiences. The Alberta Plan proposes a framework for achieving this through reinforcement learning, continually evolving AI's capabilities over time.

Call to Action For those interested in the evolution of AI and the quest for AGI, this episode offers critical insights into the ongoing development of intelligent systems and the role of continual learning in this dynamic field.

---

Additional Resources

  • Visit [NetSuite by Oracle](https://netsuite.com/EYEONAI) for further information on cloud financial systems and AI integration.
  • Explore the Alberta Plan and its implications for the future of AI research and applications.

---

This succinct summary captures the essence and critical insights from the podcast episode, providing a structured overview for quick reference and deeper understanding.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The Alberta Plan is an approach to pursuing really strong artificial intelligence, which has deep roots. So it's something that I think a lot of us at the University of Alberta and at the Alberta Machine Intelligence Institute, and it's a precursor to the AICML, have pursued over the years. And it is indeed sort of both a structure for how we might think about building intelligent agents that can learn from their own experience, and also a set of commitments that sort of tie us to an approach to doing research, which we believe is going to be able to support generating these kind of new future technologies.

0:35What does the future hold for business? Ask nine experts and get 10 answers. Bull market, bear market, rates rising or falling, inflation going up or down. Can somebody please invent a crystal ball? Until then, over 40 ,000 enterprises have future-proofed their business with NetSuite by Oracle, The number one cloud ERP, bringing accounting, financial management, inventory, HR into one fluid platform. With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions. With real-time insights and forecasting, you're peering into the future with actionable data.

1:27If I were a larger organization, this is the product I'd use. Whether your company is earning millions or even hundreds of millions, NetSuite helps you respond to immediate challenges and seize your biggest opportunities. Speaking of opportunities, download the CFO's Guide to AI and Machine Learning at netsuite.com slash IonAI. That's netsuite, N-E-T-S-U-I-T-E dot com slash IonAI, E-Y-E-O-N-A-I all run together to get the CFO's Guide to AI and Machine Learning. The guide is free to you at netsuite.com slash IonAI. NetSuite.com slash IonAI. I'm Patrick Polarski. Thanks for having me. I'm a Canada CFAR AI chair and a professor actually in the Department of Medicine at the University of Alberta and also a fellow and on the board of directors of AMI, the Alberta Machine Intelligence Institute here in Edmonton, Alberta, Canada.

2:37So my line of research, I mean, I was trained as an electrical computer engineer. I worked on all different forms of machine learning over the years, everything from evolutionary computing, which I was like, to swarm intelligence, to things like supervised machine learning, computer vision, and reinforcement learning. So that's the sort of trajectory I followed. I'm currently working mostly in bionic medicine. Artificial intelligence applied to how humans and machines interact, and in particular, our research lab is working on a very cool, I think very ambitious research project where we're looking at literally bolting robots into the bones of the human body.

3:17We work on artificial intelligence and transformative surgeries and ways of measuring people's interactions with prostheses. So we're looking at say people with amputations or other forms of limb difference, how we might begin to build a new sort of science and art of prosthetic restoration, Neuroprosthesis for people who use artificial arms and hands in particular. And so yeah, we're looking at bone anchoring surgeries, how you can rewire the human nervous system to engage better with a decision making machine like an artificial limb, artificial hand, and then building a new generation of artificial intelligence technologies that really sort of glue that person and machine together in an informational sense to act as a broker to translate what the human wants from what the artificial limb is actually able to deliver.

4:05So yeah, we're working on bionic medicine. It's very, very cool. What I think how this ties into our conversation today is that it's in fact, it is an example of machine learning technologies, reinforced learning technologies that are deployed in very low compute environments. Like we're not imagining an artificial limb that continually talks to like a supercomputer in space every time it wants to try to open and close a hand. We want devices that people can wear on their body or attach their body that when they drive under a bridge or in a tunnel on the freeway, that it's not going to stop communicating with the Internet and stop working for them.

4:41So a lot of what we think about is how you have very powerful technologies that can actually be embedded within the devices that people wear on their body. So I think that ties into what we're going to chat about today. Yes, the Alberta plan. That was what I was interested in. I spoke to Richard, Rich Sutton about it the last time I had him on the podcast, but we didn't go deeply into the Alberta plan, and I've seen it referenced more recently in reinforcement learning. because of the use in training large models is getting more general attention than it has for a while. So I wanted to hear for listeners as well as myself exactly what is the Alberta Plan and where are you in executing the Alberta Plan and how that relates to what's happening in reinforcement learning research.

5:47Yeah, the Alberta plan is an approach to pursuing really strong artificial general intelligence, which has deep roots. So it's something that I think a lot of us at the University of Alberta and at the Alberta Machine Intelligence Institute, and it's a precursor to the AICML, have pursued over the years. And it is indeed sort of both a structure for how we might think about building intelligent agents that can learn from their own experience, and also a set of commitments that sort of tie us to an approach doing research, which we believe is going to be able to support generating these kind of new future technologies.

6:26So the Alberta Plan itself is something that we have, again, talked about for years. Recently we actually wrote it up, that's the, I think the manuscript that you're referring to is the Alberta plan document that Michael Bollingrich Sutton and I put out a couple years back. And that is where we tried to really sort of codify the commitments of the Alberta plan and also what we think is a very reasonable 12-step plan to actually executing on it to build really powerful thinking machines that can learn from their own interactions with their world. And I do think there are tie-ins from this directly to what we're seeing and how RL is being applied to everything from tuning large language models to controlling industrial processes.

7:09So I'll touch on that in a little bit. But maybe I'll start with saying that like, the Alberta plan is a plan for us to pursue intelligence, the study of intelligence, like this is something that we're it's a very big tent kind of thing. This is not like a Yeah, Mike Rich and I are sitting here in Alberta working the Alberta plan. This is something that that does include many of our colleagues all around the world. People we have not met are pursuing research in line with the Alberta plan. Our students and our teams here at the University of Alberta and Amy have been going out in the world and pursuing pieces of it.

7:42But I'll say that there's maybe four main commitments that we have in the Alberta plan that help us think about how we might build reinforcement learning systems, things that actually solve reinforcement learning problems from their own experience. And one is that, in fact, it is very experience-based, that we believe in agents that are living within the flow of time and another commitment is that they aren't really like a temporal uniformity there's no like special training and testing period we're not like taking a billion dollars to distill down all of reddit plus the rest of the internet and turn it into a model the idea in the albert plan is that we believe that that all of the data that flows into the system that stream of experience the agent living within its stream of experience that experience is always being incrementally added to what it already knows and that we're not trying to divide up time into epochs of batches of learning versus batches of batches of acting, that acting and learning are dovetailed together in an inseparable way.

8:40That's part of the third commitment as well, which is that we really like in the Alberta plan, we talk a lot about respecting computation, respecting every one of those flops we're using, every single act of computation should be valuable. And if we can be as computationally efficient as possible and how we think about the learning process, then we can just do more learning. Like imagine, imagine like a bunny rabbit hopping through the briar bushes. Like if you make its operation more efficient, it can predict more things about its environment. It can start to build up expectations about its world in a way that we couldn't otherwise.

9:13I guess then the last commitment, then I'll let you jump back in with more questions before we go into sort of like what's in the Alberta plan. But I think the other really exciting thing about the Alberta plan is that, and this was sort of us like working back and Mike Bowling and I working back and forth with Rich Sutton about like, well, what do we think about multi-agent interactions? But I think the last thing that we want to observe and we want to sort of commit to studying in the Alberta plan is the idea that, well, there's other agents in the world, that an agent, a learning machine, whether or not it has to like identify, codify, build a explicit representation for another agent that lives in the world with it, that there are other agents in the world.

9:53They create complexity, they create patterns that are learnable, and they create the richness in the environment and also potential to amplify the capacity of a decision-making agent. So this is the bit of the Alberta plan that I'm perhaps most passionate about, is the longstanding idea of intelligence amplification. Like going right back to Ashby in the 1950s and the post-cyberdeticists and thinking about the ability to directly paraphrase Aspie, like that intelligence, much like physical ability might be amplified by tools that, that we can use one intelligence to multiply the abilities of another.

10:33And that, that in fact, artificial intelligence and computation is one of those powerful amplifiers we might have. So that last commitment, the Alberta plan that I hold pretty near and dear is the, is very much that we might start to see the ability of intelligence to multiply the capacity of other intelligences and that. So we, again, want to make sure that we're explicitly thinking about the contributions of other intelligence systems in the world of an agent, any kind of agent, machine or human. Yeah. And, and this, as I said, I've spoken to Rich about the Alberta plan and some of his work.

11:11And you differentiate, I mean, the goal is, is artificial general intelligence, right ultimately uh and right now the the sort of uh trajectory that gets the most attention are the pre-trained transformer uh models and and adding compute and adding data uh and you know measuring against benchmarks and with the idea that that will eventually get us to artificial general intelligence. But there's also people like Yan Le Koon, and I'm going to be talking again to Fei-Fei Li in a couple of months about world models, the idea that you need direct experience, that you're not going to reach general intelligence through recorded, particularly textual data that embodies human knowledge.

12:22How does, well, first of all, is your approach purely reinforcement learning? Is it a combination of reinforcement and symbolic learning, or are you using transformers? and how does it relate to what Jan and Fei-Fei are talking about? I think the idea that an agent needs to take into account the interactions between itself, whatever the dotted line, we always debate this in the community, like where is that dotted line that separates the agent and the environment? And it's a very flexible line. I think it's a very deliberate choice that, I mean, if I was an industrial process designer, I would be very deliberate choosing where that interface is.

13:09If I was a thinking machine that was trying to live its life in its flow of experience, that dotted line might change over time even. So I think that what's interesting about the separation of the agent and the environment is that it does sort of necessitate a relationship between them. Like we have to assume that the agent's learning is going to be with respect to what in a number we would call the aperture between the agent and the world and that's that is that that's true fundamental piece of the equation so to separate out a like let's just just again let's distill down the internet i'm being i'm trivializing it this is an amazing miraculous act that we're able to create machines that are able to like hold in their digital minds the sum total of the internet and all these private data islands that are attached to it like this is this is miraculous So I don't want to trivialize that.

14:02But the relationship between an agent and his environment is special. And if we do truly want to build thinking machines that can live within their flow of experience in whatever their world is, whether it's the world of protocol buffers that are flowing between data centers, or the human world of desks and tables and chairs and doors, we have to really deeply consider that interaction between the agent and the environment. That's part of its learning. um and there's a another piece of that which is that the the body of the agent whatever it is whether it's again some kind of network router or a protocol buffer or a piece of metal that is uh that is moving itself through the world that there is physical intelligence within that system and so rodney brooks has had some amazing uh uh long discussions about this and has written quite eloquently about it, that like the physical intelligence, the way that the body itself, not just translates the intelligence of whatever the mind is, but in fact, contains some of the aspects of intelligence itself.

15:06This is something that necessitates our consideration of the environment. So the environment, agent, environment. Yeah, I think we do need to think about that. We can't assume, and that's not to say that we can't have an agent that watches a ton of YouTube videos and then is able to actually go out there and interact with the world itself. But if we want agents that not just can interact with the world, but can continue to adapt to worlds and to new worlds that they encounter, it's very important that we think about the environment. And so the representations we choose, the data streams that we integrate into agent learning, it's absolutely critical.

15:40So this is, again, the perspective that we take. The Alberta plan perspective is very much that, that what is being learned is about that aperture between the agent and the environment. And what's cool about that is that what's learned can take many forms. And we often think about value functions. We think about expectations about the future that are about a signal of interest or many different signals of interest. Those can then be used to sculpt behavior. Those forecasts an agent makes, the way that the agent starts to build up expectations about its future experience It might be about that cool signal that we call reward.

16:16We love rewards. Nice, precious, very good signal. Agent environment loop, there's a reward signal going there too. But there's all these other things to learn. And so, I mean, when we think about the models that an agent use, and I'm going to define, I'm going to say what I mean by model, because we all talk about, oh, there's a machine learning model. Oh, this is a model. Oh, I have a plastic model of a car. When I think about model, I really do think about something that an agent has formed, an animal or machine agent has formed that is able to in some way carry the likeness of something in this case mostly its environment has interactions with its environment and so that that's a very again big tent kind of approach to model but i i do think about this sort of internal construct which allows an agent to be able to represent to manipulate to improve its computation and its ability to leverage its experience about the world to choose and make actions.

17:12So if we have that model, the nice thing about that is that could we put in a large pre-trained model, like something that we'd call a foundation model, you could absolutely leverage something like that. And I don't think anything that at least my research would pursue, and I think I can speak for Mike and maybe Rich on this too, is that I think like that we can imagine permissive models that take into account all sorts of expectations about the world or all sorts of ways of representing the world and the way that streams experience might unfold in that world. And we might imagine, so that's saying we could admit things that have been learned that were not learned directly by an agent.

17:48I think that's okay. We can also admit into our definition of models and how an agent might use it to improve its action and its computation. We might imagine models that are learned directly from its raw experience. Now, I think speaking for myself and I think my colleagues as well, is that that area, learning from its raw experience and unlocking the true potential that is, is underexplored, comparatively speaking. And that is something that has great power. And so I think that's something that we would want to continue to reinforce, is that reinforcement learning is able to build up this understanding of an agent's interactions with its world.

18:26Like the reinforcement learning agent is going to build up forecasts about its future from its raw sensory motor experience. That's really critical. but it's compatible with other kinds of models that might be learned. And that, again, could be the YouTube watching agent that's like, I really, I watched a bunch of robots and they were like, I watched all the video games, like speed runs I could see. And now I'm going to use that to like inform my motor control. That is totally okay. I like that. I think that just makes sense. And that's being, quite frankly, it's being responsible with our computation.

18:55If we're spending all this time and money training these gigantic systems, which are able to, again, and figure out how a robot might move from a YouTube video. Well, we should be using those. We shouldn't be redoing that computation every single time that we want to try to learn, have a robot learn to walk or to have an agent control itself in a grid world. Like we should be leveraging our past compute in as flexible the way as possible. That's just responsible use of resources. Yeah. So is that related at all to like DeepSeq's use use of direct reinforcement learning to train. It sounds like it might be.

19:35You were saying that reinforcement learning, the base model could be a foundation model. It could be something else. I mean, Jan has something called JEPA. Is that right? Anyway, I'll cut that out. But yeah, so what's, how is that related to this reinforcement learning training that suddenly everyone's talking about? Yeah. So I'll start with my personal perspective of the kind of models that I like. I really like considering agents that have predictive models of their world. And this means that like the kind of model they might form is about the raw bits of experience flowing across all of those wires into their digital mind.

20:25So for years, I have taken what normal engineers would say were perfectly reasonable signals, like the joint angle of a finger or the current going into a servo motor or the pressure that something. I've taken those perfectly reasonable scalar floating point numbers and just turned them into bits and stuff those bits into the mind of the agent so that it can start to predict the future of how those bits are going to change. So that's like, as my disclaimer, my starting perspective is very much of a blinking lights on wires. the agent is predicting its raw experience without that experience being tagged or labeled, or in fact, even separated into bins that we might call a floating point number or an 8-bit integer.

21:05Okay. So that's my default perspective. I love predictive models. I think agents that leverage their predictive models have amazing potential to interact with diverse environments in ways that agents that are structurally biased with our own designer, like human way of segmenting the world might be limited. So all of that said, like the idea that you might start to use reinforcement learning to take whatever kind of base understanding, the ability to act or make decisions or to produce tokens or to produce motor trajectories, to be able to take reinforcement learning and place it within that, to be able to sculpt and optimize or to adapt that behavior.

21:44This feels like just a absolutely the right win. I think we probably see this happening in biology all the time. I think that, like you mentioned, DeepSeek, I think that that is another great example. It goes back to the we actually had a we were discussing this at a reading group yesterday, we were brought in the the paper from David Silver, Satinder Singh, doing a pre cup and Rich Sutton called reward is enough. I don't, I don't know if you're familiar with that. I'm sure you're familiar with rewards enough. But like, to paraphrase poorly paraphrase, of what our colleague said in that paper was that if you have a single signal of reward, this precious signal, then if an agent is truly trying to maximize the extended amount of that signal that it collects over time, and the task requires it to learn anything from chain of thought reasoning or social interaction or vision, if the task to be solved requires that the agent build this advanced capacity, simple or advanced capacity of what we consider intelligence, then that signal of reward is enough to drive the agent to create whatever it is.

22:54Again, for the deep seek paper, this is like the idea that it might start to create chains of reasoning, that it might be able to unfold in like its little think brackets, it'll actually unfold the steps it took to get to a solution. And then the reward signal itself from all those myriad tasks is in fact driving the creation of those chains of reasoning outside of other sort of like labeled examples that are being stuffed in. Of course, compatible with sculpting from all this other kind of input information. So I think that like the reward is enough, like this is now stuck in my head because I talked about it yesterday, but the reward is enough idea I think is very powerful because it does let us think about how we might close the gap between all of those thinking machines that we have created and thinking could be very simple ways or very powerful ways.

23:44And the like complete cohort of goals that we might have for the technology that surrounds us in our lives. Like this is a, if there is this missing piece, like, oh, Hey, I really would love chain of thought reasoning, or, you know what? Social intelligence is going to be really important or, oh, wow. I got to have a better idea of how bio signals flow from the human body to like a robotic device attached to that body that thinking about reinforcement learning and a very much tabula rasa, perhaps reinforcement learning as the thing that can pull those two, the outcome and that prior foundation of technology together.

24:15I think that's a really powerful thing that we should continue to consider. Yeah, that just brings to mind very early on when I was trying to understand this stuff and supervised learning was reigning supreme. that somebody was saying, you know, and everyone uses the example of a mother teaching her child what a cow is. You know, that's a cow, and the child learns. That's the label, and from then on when it sees this thing, but there's the argument that it's actually reinforcement learning, that the child is getting praise from the mother when it recognizes the cow. And, yeah, so reinforcement learning can be seen as kind of the root algorithm operating in learning.

25:20But how do you implement that? I mean, can you build a model, an intelligent model from scratch purely with reinforcement learning, or do you need a base model, as you were saying? It's interesting, actually, I think, just to follow up on the idea about the cow. I'll go back to a conversation. I think this was like, this must have been 2009. So Rich Sutton and I were having a conversation back in 2009 about chairs and how you think about chairs and how a machine might think about chairs. And I promise this is more interesting than you saying like chairs are – that like the normal way we might think about a chair or hope a machine thinks about a chair.

26:03Like look at the vast swath of computer vision literature, which has considered objects like chairs and tables and doors. You might think about the chair as like how wide is it? How tall is the chair? Where's the seat? Is the seat hard or soft? how far is the chair from the door so that's a very normal objective worldview like this is sort of thinking from the perspective of the frame of reference of the world what is how do i define the chair its extrema its boundaries its properties there's another view this is the view that i think rich was advocating during that conversation there's another view on that chair that we might take which is well what's the probability if i were to try to sit down right now i would actually sit or would I fall?

26:46Like there's, if I were to walk three steps, would I bump into something with my shins? There are these very subjective, this is the agent's frame of reference. Now that the frame of reference inside the agent, like with respect to the data stream accessible to the agent, what do its predictions say about his interactions with the world around it? And so I think a lot of early sensory motor learning, you look at motor babbling in babies, like they start to learn to move their head around and then their eyeballs and they start to be able to figure out that, oh, wow, like, if I move my head this way or that way, I there's a noticeable change in my sensory experience, I predict that if I'm able to, if I do this, I am going to be able to see the ball that's outside my vision, everything from like, the object permanence to actually, Perry's UDA has some really nice work on this with robots doing the same thing that babies do.

27:32It's a there's a fantastic paper from quite a few years back. But, but it's really neat, because this says that the system is now beginning to make these like elegant scientific tests, the agent as a scientist saying, well, every time I move my head this way, I start to see this other thing. Or I realized that if I, if I move all of these bits of myself, something appears in front of my face, it's my own hand. Oh, that's so cool. So like, there's this rudimentary fundamental learning process, which unfolds, you see, it happens with like baby animals and humans and otherwise, that is rooted in these predictions.

28:07And there's like, again a whole elegant cohort of neuroscience literature as well which says oh like the fact that i feel like i'm moving my arm is in fact not the sensors reporting back that i'm moving my arm it is my predictions about my arms movement and then if something gets in the way like my hand or something then there's correction that happens with predictions but the feeling of agency even over my own movements might be due to the predictions in my mind prediction precedes control, says Daniel Wolport and colleagues, in motor learning. And so from that perspective, you can imagine chaining together, again, going back to reward is enough, chaining together all of that process of learning to say that, well, yeah, you should be able to learn everything you possibly would need to know about the world just from signals of reward and predictions of your raw sensory motor behavior.

28:57This is the prediction perspective, the subjective agent frame prediction. now like is that the fastest way to build a robot butler or a self-driving car i well probably not like let's be honest if i was if i was going to go out and start building a self-driving car right now i mean i want to pitch this to someone i was like hey i would happily help your company build a self-driving car give me some cars and let me start smashing them into stuff and they actually said maybe which i thought was a really positive sign but if you wanted to take that sort of constructivist so it's back to constructivism if you want to take the constructivist view the subjective agent frame, constructivist view of learning a very hard task like self-driving or like full problem solving at the math Olympiad or something like that.

29:41Well, I mean, there might be easier or harder ways. What I think is neat, though, is if we look at sort of the trajectory of like AlphaGo, everyone loves talking about AlphaGo. I also love talking about AlphaGo. Go for it right here. And like if we start to think about a version of AlphaGo that learned from watching human examples and then had these sort of gaps in its mind's eye, had certain delusions, and then go to like an alpha zero, which learned through self-play that it was able to drop some of those sort of data biases that let it find solutions it never would have found otherwise. And I said earlier, like, oh, wow, I love having a machine see the world through its raw bits without me even putting like floating point numbers on top of it.

Read the full transcript

30:24Part of that is saying like the agent learning from as raw a sensory motor signal stream as possible allows it to sort of undo some of the assumptions I might be making as a designer, which means it might be able to find just a much better solution, maybe even faster than the other way around. Now, I think, you know, building up human language, it takes a while for a human to learn human language through trial and error, through motor babbling. It's going to take a while to learn language. Would we rather just like, again, juice the internet, make some amazing lemonade and use that as our foundation model?

30:54Like, like, yum, large language model, pre-trained, it's great. I think, well, of course we would. This makes a lot of sense. Everything everyone is doing, I think, does make a lot of sense. But there's this other path, and that's why, as again, going back to the Alberta plan, why I think we want to make sure this stayed in the public mind's eye is that there are alternate and compatible views, which allow us to maybe fill in some of the gaps, or maybe to start from a very different foundation point and get to something that is like, maybe our current curve goes up and then it sort of starts to taper off.

31:24Maybe there's another approach where you do pursue this raw sensory motor learning about the stream experience, but then that actually has like a higher final peak. So let's keep, again, we don't want to have a monoculture. So it's keeping that diversity in our mind stream. And then also being really like not being dogmatic or like siloed about it and saying, well, where is the opportunity to do something like we saw in DeepSeek where they're like gonna throw just like raw reinforcement learning to help the model get to the next step of its chain of reasoning. I think it's really important that we aren't, again, too siloed or too cliquey in our communities.

31:59Be like, no, I'm going to sit over here and do raw sensory motor learning and ignore everyone else. I have the luxury. Yeah. Are there projects to do just that? Start with raw sensory motor learning. I presume that the learning then is encoded in the weights of a neural network, right? I mean, when there is a learning, but is there a project to start with, you know, whether it be YouTube videos or a robot arm or something and see how much learning can accumulate and whether it scales? Yeah, I think there's there's like a happily a pretty spectacular library of examples globally. Now I think we're starting to see, and like, I think where we, where we see them most is in specialized targeted domains.

33:00So like the, the work from deep mind on stabilizing plasma in a tokamak reactor, decision-making system acting thousands of times a second on a very diverse stream of data to literally control a lobe of plasma hotter than the heart of the sun. That's really neat. That's reinforcement. It's learn from reward. And then we see systems that can do elite coding tasks that were also driven directly by reward. And so as we start to see all of these different domains, like, well, one of those is very sensory motor. I would argue that a stabilizing plasma in a tokamak fusion reactor is, in fact, very much a sensory motor experience kind of thing.

33:38Same thing with cooling a data center or even compressing bits in YouTube using MuZero. things like that are much more sensory motor they're very much like again the the nerves firing down the spinal cord and letting my arms do all sorts of cool things and then something like well i mean doing really well at like elite coding tasks and like programming competitions well reward driving that feels more like the languagey thing like that that feels a lot more cerebral it feels like you know that like the deep thinking not just like the reactive i'm a tiger running through the woods kind of thinking but what we're seeing is that reinforcement learning applied in its sort of raw sense is being able to generate insight in both of those domains.

34:19So we'd be here probably till, I don't know, next week if I were to start the list off. I usually, if I give talks, I actually have this best of hits collection at the beginning of my talk. I'm like, hey, what's been happening in reinforcement learning? Well, have you met Gran Turismo Sophie? Okay, cool. You have? That's neat. Our colleagues at Sony AI did a really cool job of building a racing system which could do high performance superhuman driving while also learn the implicit rules of driving in racing, not cutting people off. You don't sideswipe people. There's rules of racing that I was able to begin to build up.

34:53And that was also, again, through process of reward-based learning. So going back to reward is enough in this case, you need some social interaction skills. Well, you could, if the task requires it, you can get that through reward. Do you want to be able to solve a programming task? Well, yeah. I mean, you, if you need to be able to think about data structures and the way you chain together methods and, and, uh, different computational artifacts, well, yeah, you could drive that with reward. There's a whole bunch of them thinking, well, what's the, like, is it just bits on a wire? Like Patrick proposes.

35:24Okay. Probably not. Like, this is a very bold, like if we were to go all the way to sort of like Joseph O'Dale is another person who really enjoys sort of the, the bits on the wires, Rich does too. But like thinking about an agent's experience just in terms of raw ones and zeros without any kind of labeling on them there's a there's a like there's a quick shortcut you can take which maybe you shouldn't take again if you believe my argument about like design biases but there's a quick shortcut say well let's just actually use floating point numbers and integers like patrick said we shouldn't and then there's another which like well let's just use like you know we'll use more more advanced data structures as the representational elements the foundational elements the machine is starting to use the actions it might take.

36:05And then you might say, well, wouldn't it be faster if we just use some of those that were already learned through processes of supervised learning? I mean, I like to think about supervised learning, and again, a lot of our colleagues would say that it's like learning through labeled examples. So as opposed to supervised, well, that was like learning through labeled examples. Do we learn through trial and error? Do we learn from the structure of the data? There's cases where each of those might be a great way to start to build the fundamental units of experience that an agent is operating upon.

36:33And constructing those, you could say, yes, there's like we back in 2011, a large number has put out a paper on generalized value functions. And one of the reasons we were excited about that, one of the reasons that we thought about it's like generalizing the idea of learning a value function about reward, like expectations of how much reward will I accumulate in the future, assuming I care about the future in this way, to do that for any signal of interest. The reason we've been passionate about this and continue to pursue this for well over decade now is because we viewed those predictions, those forecasts, those expectations of future signals to be those foundational elements of building up other more advanced behaviors for an agent.

37:12So that again goes the choice of well what is the fundamental units of an experience stream. Yeah but the examples you're giving are sort of narrow domains. How does that generalize into uh or how how do you foresee that generalizing into general intelligence i like when i even think about the like attention is all you need transformers what are we getting from some of these systems like i i view them as like the multimodal versions of themselves not just like a stream that's predicting language tokens but to think about like well what about a vision transformer What about these systems that are truly multimodal?

37:54When you extend that one step further and just think about what the operations are happening in these systems, they're creating associations across space and time. Like many different pieces of information without any of this like sort of a bit like staged convolutional foveation on a specific part of the data. Like these, these systems are able to consider the flow of time, spans of time, which are getting increasingly vast and then creating associations across that like large span of time, but also the different, the different channels of information could be anything from like the stream of tokens in a sentence alongside the audio data coming in alongside the time series of the vision pixels to the random bits on a wire that I was proposing to you earlier that the associations across space and time are very, very powerful.

38:40I mean, we like Dale Sherman's and colleagues have the, have an awesome paper where they're like, yeah, if you like give transformers redire operations, they're turning complete. Grossly paraphrasing Dale and colleagues on that paper, but it's really cool. Cause this is like, oh, well, yeah, this is, and I'll sort of paraphrase Dale as well here and say that this is another way of thinking about doing computation. It's a fundamentally different one. And so what's happening with all this, we have, we have cross attention. We have the ability to chain together facets of information in the beautiful gemstone of the past experience of an agent to create predictions about the future.

39:16Those predictions are acted upon. They're either going to move a joint or more likely they're going to be able to spit out a new word token or to spit out a new pixel in a generative video. The same thing happens in the base process of learning a generalized value function. you're thinking about based on a representation of some kind, which we hope is able to form, again, associations across both space and time, that a system is able to generate a prediction about the future. And these predictions for generalized value function, for instance, are temporally extended. They might be thinking about the expected sum of a signal conditioned on a policy that assumes some kind of preference over the time window that it's being integrated.

39:55but either way this is you could definitely do token prediction with that if you wanted to you can also predict the future current on a robot arm and start it moving early to be able to like catch a ball mid-flight so the associations across different modalities or channels of information and across time and using that to create a prediction like these are different ways of achieving the same thing you could imagine using them flexibly together you could imagine and committing hard to one of them and working on that. So I do, I feel like that's the potential here is to see it as, oh, wow, this system is able to put together the fact that I had coffee last week, plus the fact that I'm going to buy milk and then get like, oh, you're probably gonna make coffee today.

40:37Like that's a very sort of high level human thing or like it's grading my cashmere or whatever. Like the systems are going to be able to put together these pieces of information to make predictions about the future. And then how you think about the form of those is maybe where I see people are starting to disagree, but sort of zoomed out a little bit to bring together the cross-attention idea with some of those token prediction in terms of raw sense remote experience. There's a lot of unity there that I think that we're just not celebrating, quite frankly. And I really like it because it does, like some of the tricks that we're starting to see, and the tricks is not to trivialize the elegance of the computation for like taking a vast history and giving the access, like a large language model, having access to a vast history of tokens.

41:25Like me sitting there with like Google Gemini and typing and having this lovely conversation about like structural neuroscience. Like there's a lot of tokens that are going in. Like we chatted for an hour. And like the tokens, like getting these long histories and then still thinking about the relationships across these vast spans, I think is something that is exciting and new. Like history is hard. And often we try to make it, Like a lot of the work we do for the Albertan line, we're thinking about how to make learning incremental by making sure that every new sample that comes in is sort of bolted out the old so we're not keeping around histories.

41:58But there's new ways that we can start to think about this maybe close to Markov level unfolding of experience based on these technologies that can integrate long, vast histories. Context. Context is critical. I think we're getting closer to context. This is the thing that warms my heart. Reward is all you need. I mean, the idea is that reinforcement learning can build up this world knowledge or this representation of reality through which it then can make an agent can make predictions and act in the world. But, as you said, there are quicker ways to get to a baseline and maybe reinforcement learning.

42:45You start with that baseline and go from there. I mean, I've been talking to Pedro Domingo. I don't know if you know him. He had that book, The Master Algorithm. and he's saying that, you know, it's not going to be one of these schools of thought. It's going to be a blend. And that's why I was interested in DeepSeq's reinforcement learning training. And I met recently Mark Raybert from Boston Dynamics or originally from Boston Dynamics. And he does not believe that end-to-end neural networks can, that robots can operate on an end-to-end neural network control. That you need a traditional control theory.

43:44you need very exact equations for the precision that's required. And Boston Dynamics, for all they do, that's not machine learning. So how do you feel about that? I mean, you're working on prosthetic arms or legs. How do you use control theory and then build up knowledge through reinforcement learning starting with control theory? Or is it raw reinforcement learning and, you know, you can eventually get to the precision required? I mean, yeah, if you could talk about that. We've tried a number of different approaches over the years. And maybe I'll say one that continues to shine for me. And it works really well in a clinical setting.

44:40Works well in an environment where you're trying to bring together technology, people who build technology with clinical practitioners who are experts in their clinical fields. And that's that I love starting with an approach that the system will begin at the gold standard, whatever it is. And this is especially for human facing technologies. That's where I mostly work. Human facing technologies. Like I would very much like a system that starts at the gold standard and then gets better with experience. It does learn through interactions with its raw stream experience. It is able to integrate those other facets of experience into its learning and improve or optimize its behavior from the gold standard.

45:16But that then fails back to the gold standard. If and when it makes a mistake, then it still defaults back to the like whatever that gold standard operation would be. So this actually says that if we are, of course, like there are cases, especially human facing cases, where do we want to appeal to control theory? We sure do. Does control theory currently allow us to think hard about soft robotics, deformable systems? Imagine a liquid metal robot that's able to change its form completely and has very strange dynamics like jamming systems, anything else that you might see in like any of these sort of reconfigurable robotic systems.

45:52Well, there's branches of control theory which are attacking those problems, but these are hard problems. So there's case where, yeah, we can have this piece is definitely going to be nailed by control theory. Can we improve upon it to make it more resilient, more agile, more able to adapt to the world around it through learning about its stream of experience? I think the answer is absolutely. Like some of our colleagues here are, RL Core Technologies is a company from Edmonton and some of our colleagues from the University of Alberta are leading that company and they do like reinforcement learning technologies for one of their examples is for water treatment.

46:24You can take a water treatment system and start to think about how RL can be used to improve the ability of a community to treat the water and make sure their drinking water is safe and reliable. There's going to be some parts of that. And I think you want to interview like Adam White or Martha White about this. I think that would be a great guest to have on the show to talk about how they're pursuing this. But again, to paraphrase what they've been working on, I'm like, yeah, you're going to have some parts of this like water treatment plant, which are very much like solid gold standard control engineering, but then you're going to be using this learning through experience, learning through trial and error process to improve the pieces that can be approved upon that gold standard, such that it can adapt to new circumstances to be more adaptable or to be able to operate for long periods of time without constant human intervention.

47:09So I think these are really the space explorations of the great space where you're like, yeah, do I, do I want some kind of control system probably governing my like plasma thrusters? I sure do. But do I want, the spaceship to like steer itself like maybe learn to avoid that yeah yeah that would be great in new environments but there's like certain low-level processes like our biological processes in our own bodies like the the protein machinery or our cells is going about doing its business um and yeah there's some adaptation going on there but a lot of what we think about is intelligence is is the the the thing that's moving around all those little protein machines right Yeah.

47:48So the reinforcement learning is learning, right? It's trial and error, trial and error. And it's that learning then, and I'm sorry, I'll cut this out if I'm revealing too much ignorance or if I'm off base. But, you know, I understand neural networks and backpropagation and how the network learns. But with large language models, once a model is pre-trained, its learning is kind of fixed. I mean, you can fine tune it, but if you start, I mean, there's this problem of catastrophic forgetting. If you run it long enough, then it's going to start forgetting or the weights are going to start changing in the network that encoded previous knowledge.

48:52How do you deal with that with reinforcement learning? Is it, what happens to the learning? Does it, is it encoded in the weights of a network? And then can you just keep expanding that network so that you never lose any of that learning? This is an excellent question. I think I'll take this question to the domain of continual learning. So this is really exciting. Like I think a lot of us in the community are passionate about continual learning. And that I mean, a system that is able to continue to learn, adapt and change. It's very constructivist in nature. Like you would hope that a system is able to continue to integrate new information into what it knows, not overwrite the old stuff, unless the old stuff is supposed to be overwritten to be able to flexibly compose skills and behaviors it's learned in the past in a new way and new situations and to continue to do so forever, literally forever.

49:46So this would be continual learning. And there's some good review papers. Kimmy Kirtapala and Joanna Precup have this awesome review paper on continual learning, which is definitely worth going through. But so continual learning, the way the field, like there's a new conference on continual learning colas, which is really cool. And there you see in that community, a variety of approaches and that you called them both out. So I like this. Like one is, well, okay, I have a giant, let's imagine I have a giant large language model. Okay, cool. It's got all of these factoids baked into it in whatever way, whether there's a grandmother neuron or whether it's a collection of activations that allow us to have the concept of a grandmother.

50:23You'd want maybe that system to be able to be incrementally improved without having to go back and spend how many billions of dollars, or I guess now millions of dollars if you take the deep seek approach, how many billions or millions of dollars to retrain the whole thing from scratch on the data. Well, maybe you don't want to do that. Maybe you just want to take new information and have it be able to be integrated into that system. Like my brain, I don't want to retrain my brain every single time I learn something new by reading a paper or like talking to you today. I don't want to have to retrain my whole brain.

50:53So there's a whole community of people that are in fact looking at that problem. And there's been some really good advances there, which is again, taking a piece of information and being able to use that to upgrade an already in place system. And again, like if we go down to the brass tacks of it, like there's, I don't think we have any analog to like glial cells in the neuroscience of our digital system so for the most part we're looking at weights and you are saying oh like can i modify either the weights or some other holding memory variables that are impacting the operation of the network if there's other shenanigans going on then can you go and take this new piece of information actually modify those one those ones and zeros and the the other numbers that are comprised that comprise those ones and zeros um so that's one approach is to say yeah i would love my continual learning to be a thing i would I'd love my large language model to continue to learn.

51:39I would love to be able to take whatever news article is popped online, which hopefully isn't written also by a large language model, because then, as we saw from that recent paper, you start to get these collapse iterations happening. So, okay, that's one branch of how we might think about continual learning. And the other is that, well, we actually don't want just be sort of modifying like a large language model. So we're actually thinking about the process again, those ways as being constantly updated from the raw experience in a very incremental way. And that's saying, I'm not taking new factoid and just jamming it in.

52:11It's saying that every single moment of interaction is already changing the system. And both of those communities, like taking both of those views, you're modifying weights, you're modifying numbers that are stored in some kind of digital mind of the machine. But you're hoping that they don't overwrite things that shouldn't be overwritten. They respect context. Like if I'm upstairs making tea or I'm drinking tea down here, Or maybe that I want to be acting differently. I want to be thinking about my predictions about likelihood of spilling hot water myself differently. So I think those are really important.

52:44And you do want to see that continual learning can happen for long periods of time. There's been like some of the there was a various paper that by some of our university of our colleagues that actually showed ways you might think about stable continual learning. This came out came out recently. And so that paper as well is starting to say, we have new ways to think about allowing a system to continually keep changing its weights without having catastrophic, like forgetting or catastrophic failure after like hundreds of millions of epochs of learning or time steps of learning. Like you want systems that, that should stay stable for very long periods of time.

53:18If I send out a continually learning system on a spacecraft, it's going to be Alpha Centauri. I sure hope that like we don't find out, oh, shoot, our learning approach after 100 million time steps, everything just flatlines. And it's like it predicts zero for every single outcome. Like, boom, there goes your spaceship. So I think we like as we imagine long spans of experience, machines that learn and act over long spans that we do definitely want to consider both of those styles. One, which is very strategically, surgically adding new information into an existing system in a way that retains old content without getting rid of the essence of the new content.

53:53Or systems that can continually update everything in the network in a flexible way that is more of an incremental hands-off way. I think we want to consider both of those. And luckily, the community is doing both of those. Yeah, and just intuitively, and maybe it's the first, can you just keep expanding the network, adding layers or adding nodes per layer or something as it learns? Well, this is a very fun topic of conversation. You should have Rich back on the show to talk about his current thoughts on this. I won't want to take any of his thunder and share it. But the two examples I was just pitching are both sort of assuming that the network architecture itself stays the same, that the com nets connected to the fully connected network, all that stuff.

54:46You have all of the pieces of the network that are more or less the same shape and size throughout the course of learning. And then you're thinking about how do you modify those? There's a very reasonable, natural approach that you might consider, which is, well, how do we have the architecture of the system? So this is a separate piece of continual learning. Like one is, well, we've got a set sort of skeleton. And now we're just thinking about how beefy the muscles are in different parts of the skeleton. And the other is, well, let's just add more bones onto the skeleton and grow muscles around them.

55:13The weights being the muscles and the skeleton being the architecture connections. So I feel like that's a really very interesting. And it's not like there's deep lines of thinking on this. This is not just like, oh, shoot, we just figured this out this year. We should think about it. But respecting all of those old lineages of thought, like how we might grow a network from scratch is very constructivist. It's like you start with a single weight and a single input and output and a single weight. And well, clearly you're not solving an XOR problem with that. Like, sure, sure not. Okay, well now we have like, oh, we got like a bunch of inputs and you're still not solving an XOR problem.

55:46Oh, but if you just add this and then it crosses, oh, cool, you can solve, you can solve a different kind of problem now. You're like your perceptron goes into a multi-layer perceptron. So thinking about how a system can use its own learning experience to grow and change its architecture in addition to its weights in a continual sense is a very interesting, I think, a very exciting open problem. And luckily, people are thinking about it. Again, go talk to Rich sometime and see what he's been up to. In your work applying reinforcement learning to prosthetics, can you just talk a little bit about that?

56:17And maybe we'll do a separate episode at some point. but uh uh how how how you're approaching that yeah i think uh just as sort of a reprise the like we're really thinking about the interaction between a human and artificial limb and how artificial intelligence and machine learning technologies can allow someone to use an artificial limb to do all the things they want to do in their daily life that they're not limited by the technology that they want to express themselves they want to engage in their communities they can do that so how do we use ai machine learning to sort of power that interaction and the view that we take and the view that i've sort of based my entire research line on is in fact a very controversial one i'd say which is that to take the to consider the relationship between a person and their artificial limb as a multi-agent relationship now this is like this is actually very contemporary because previously the sensors the actuators the computation available sort of prohibited us thinking about the machine part of that relationship as an agent.

57:21Like it wasn't able to build models of the person in any appreciable way. It wasn't able to consider the environmental factors. And so I think now what we're seeing is that with these surges in computational capacity and the ability for systems to be more like agents, we can now actually think about what would it mean to have a person and their bionic limb jointly acting in the world, working together to try to affect change in their shared environment in both space and time. And I think that is a really neat and different perspective. And so, yeah, we've been considering that quite actively. That's the unique sort of line into it that we've taken.

57:56Okay. Well, I don't want to take up too much of your time. As you can see, I'm operating kind of at the outer edge of my knowledge. But that's what I like about these conversations. I'm sort of expanding my network. I mean, to that point, that's what the brain does, right? It grows new neurons. You're absolutely right. And we have left behind the thinking of the brain as this crystalline entity that slowly decays over time. We know there's new neurons created. We now know that it is, in fact, a system that is continually refreshing itself. That's a big change in the neuroscience community over the last 50 years.

58:37And I think that is, again, a nice way for us to consider how we might think about our digital systems as well. The Alberta Plan is an approach, is a point of view that sees reinforcement learning as the key to building intelligence. And there are various research projects under the Alberta Plan that are either domain-specific or more theoretical. what's the end game of the Alberta plan if there is one if you look at the like so we have there's 12 steps to the Alberta plan in addition to those sort of those commitments we have I think it is like if you look at them all it's it thinks about prediction about being able to have control behaviors about being able to planning with learned models and if you put those all together you get a system and this is the end game it's a very simple one it's right and I think in the first paragraph of our manuscript which is that you get a system that can truly learn from its raw sensory motor experience and continue to do so over whatever its lifespan might be in any environment, any environment, not even a human environment.

59:48And I think that's a really beautiful thing to aspire to and something that will help us more fundamentally understand the phenomena of intelligence itself. What does the future hold for business? Ask nine experts and get 10 answers. Bull market, bear market, rates rising or falling, inflation going up or down. Can somebody please invent a crystal ball? Until then, over 40 ,000 enterprises have future-proofed their business with NetSuite by Oracle, the number one cloud ERP, bringing accounting, financial management, inventory, HR into one fluid platform. With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions.

1:00:42With real-time insights and forecasting, you're peering into the future with actionable data. If I were a larger organization, this is the product I'd use. Whether your company is earning millions or even hundreds of millions, NetSuite helps you respond to immediate challenges and seize your biggest opportunities. Speaking of opportunities, download the CFO's Guide to AI and Machine Learning at netsuite.com slash ionai. That's netsuite, N-E-T-S-U-I-T-E dot com slash ionai, E-Y-E-O-N-A-I, all run together to get the CFO's Guide to AI and Machine Learning. The guide is free to you at netsuite.com slash ionai.

1:01:39netsuite.com slash ionai.

From the publisher

This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.

NetSuite is offering a one-of-a-kind flexible financing program. Head to  https://netsuite.com/EYEONAI to know more.



Can AI learn like humans? In this episode, Patrick Pilarski, Canada CIFAR AI Chair and professor at the University of Alberta, breaks down The Alberta Plan—a bold roadmap for achieving Artificial General Intelligence (AGI) through reinforcement learning and real-time experience-based AI.

Unlike large pre-trained models that rely on massive datasets, The Alberta Plan champions continual learning, where AI evolves from raw sensory experience, much like a child learning through trial and error. Could this be the key to unlocking true intelligence?

Pilarski also shares insights from his groundbreaking work in bionic medicine, where AI-powered prosthetics are transforming human-machine interaction. From neuroprostheses to reinforcement learning-driven robotics, this conversation explores how AI can enhance—not just replace—human intelligence.

What You’ll Learn in This Episode:

  • Why reinforcement learning is a better path to AGI than pre-trained models

  • The four core principles of The Alberta Plan and why they matter

  • How AI-driven bionic prosthetics are revolutionizing human-machine integration

  • The battle between reinforcement learning and traditional control systems in robotics

  • Why continual learning is critical for AI to avoid catastrophic forgetting

  • How reinforcement learning is already powering real-world breakthroughs in plasma control, industrial automation, and beyond

The future of AI isn’t just about more data—it’s about AI that thinks, adapts, and learns from experience.

If you're curious about the next frontier of AI, the rise of reinforcement learning, and the quest for true intelligence, this episode is a must-watch.

Subscribe for more AI deep dives!



(00:00) The Alberta Plan: A Roadmap to AGI  

(02:22) Introducing Patrick Pilarski

(05:49) Breaking Down The Alberta Plan’s Core Principles  

(07:46) The Role of Experience-Based Learning in AI  

(08:40) Reinforcement Learning vs. Pre-Trained Models  

(12:45) The Relationship Between AI, the Environment, and Learning  

(16:23) The Power of Reward in AI Decision-Making  

(18:26) Continual Learning & Avoiding Catastrophic Forgetting  

(21:57) AI in the Real World: Applications in Fusion, Data Centers & Robotics  

(27:56) AI Learning Like Humans: The Role of Predictive Models  

(31:24) Can AI Learn Without Massive Pre-Trained Models?  

(35:19) Control Theory vs. Reinforcement Learning in Robotics  

(40:16) The Future of Continual Learning in AI  

(44:33) Reinforcement Learning in Prosthetics: AI & Human Interaction  

(50:47) The End Goal of The Alberta Plan  




More from Eye On A.I.

All 266 episodes
#241 Patrick M. Pilarski: The Alberta Plan’s Roadmap to AI and AGIEye On A.I. · 1 h 2 min
Listen in VO