#150: Yann LeCun on World Models, AI Threats and Open-Sourcing

2 Nov 2023 · 56 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode Summary

Episode #150

Yann LeCun on World Models, AI Threats, and Open-Sourcing

Host

  • Craig S. Smith: Longtime New York Times correspondent and host of Eye On A.I.

Guest

  • Yann LeCun: Turing Award winner and one of the founding figures in deep learning, known for his work on convolutional neural networks.

---

Episode Overview In this 150th episode, Craig discusses various topics surrounding artificial intelligence (AI) with Yann LeCun, including advancements in AI technology, the concept of "world models," the open-source debate, and societal threats posed by AI.

Key Topics Discussed

  1. World Models
  2. Definition: Systems that predict future states of the world, enabling AI agents to make informed decisions.
  3. Training: Involves observing the world and learning to predict outcomes from actions taken by agents.
  4. Complexity Levels: World models can have varying levels of complexity, including understanding the agent's own state and the external world.
  1. Computational Demands of AI
  2. Discussion on the energy requirements and processing power needed for training modern AI models.
  3. Difference in computational needs between generative models for video and text.
  1. Embodied Turing Tests and Augmented Language Models
  2. Embodied Turing Tests: Evaluates AI based on its ability to perform tasks typically done by living beings.
  3. Importance of augmented language models to enhance reasoning and tool usage capabilities.
  1. AI as a Societal Threat
  2. Yann LeCun expresses his belief that current AI systems do not pose an existential threat to humanity.
  3. Discussion on the risks of using AI for malicious purposes, emphasizing that AI systems are not yet capable of performing complex destructive tasks autonomously.
  1. Future Directions in AI Development
  2. LeCun outlines the necessity for new architectures and concepts to achieve human-level AI.
  3. The importance of continuous learning in AI systems to adapt to new, untrained scenarios.
  1. Open Source AI
  2. LeCun advocates for open-source AI, arguing it fosters collaboration and innovation while also enabling diverse contributions from the global community.
  3. Concerns about the potential legal implications of open-sourcing powerful AI models.

---

Key Takeaways

  • World Models: Essential for AI’s decision-making abilities, enabling planning and predictive capabilities.
  • Computational Efficiency: The development of new architectures could potentially reduce the operational costs of AI training.
  • Societal Impact: The benefits of AI can outweigh risks, provided that the technology is developed responsibly and ethically.
  • Open Source Importance: Open-source platforms are crucial for ensuring diverse contributions to AI development and preventing monopolization of technology.

---

Conclusion This episode emphasizes the transformative potential of AI while addressing both the societal implications and the necessity for thoughtful, collaborative development practices. With contributions from leaders like Yann LeCun, the landscape of AI continues to evolve, posing both opportunities and challenges for the future.

Follow-Up Listeners are encouraged to engage with ongoing discussions about AI ethics, open-source technologies, and the future trajectory of AI development. The conversation highlights the importance of understanding AI in a broader context, particularly as it becomes increasingly integral to various industries.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Even if you train a system to have a world model that can predict what's going to happen next, The world is really complicated and there's probably all kinds of situations that the system hasn't been trained on and need to fine tune itself as it goes. The question of how we organize AI research going forward, which is somewhat determined by how afraid people are of the consequences of AI. So if you have a rather positive view of the impact of AI on society and you trust humanity and society and democracies to use it in good ways, then the best way to make progress is to open research. AI might be the most important new computer technology ever.

0:35It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud. Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds.

1:28If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together, oracle.com slash IonAI. Hi, I'm Craig Smith, and this is IonAI. In this episode, I speak again with Jan LeCun, one of the founders of deep learning and someone who followers of AI should need no introduction to. Jan talks about his work on developing world models, on why he does not believe AI research poses a threat to humanity, and why he thinks open source AI models are the future. In the course of the conversation, we talk about a new model, Gaia 1, developed by a company called Wave AI.

2:29I'll have an episode with Wave's founder to further explore that world model, which has produced some startling results. I hope you find the conversation with Jan as enlightening as I did. I mean, first, the notion of world model is the idea that the system would get some idea of the state of the world and be able to predict sort of following states of the world resulting from just the natural evolution of the world or resulting from an action that the agent might take. So if you have an idea of the state of the world and you imagine an action that you're going to take and you can predict the resulting state of the world, then that means you can predict what's going to happen as a consequence of a sequence of actions.

3:14And that means you can plan a sequence of actions to arrive at a particular goal. And that's really what a world model is. At least that's the way people have understood the word in other contexts, like in the context of optimal control and robotics and things like that. So that's what a world model is. Now, there are several levels of complexity of those world models, whether they model yourself, the agent, or whether they model the external world, which is much more complicated. And then, so training a world model basically consists in just observing the world go by, and then learning to predict what's going to happen next or observing the world, taking an action and then observing the resulting effect.

4:04An action that you take as an agent or an action that you see other agents taking, right? So that establishes causality essentially. You could think of this as a causal model. So those models don't need to predict all the details about the world. They don't need to be generative. They don't need to predict exactly every pixels in a video, for example, because what you need to be able to predict is enough details, you know, some sort of abstract representation to allow you to plan. Right. So, you know, you're assembling, I don't know, something out of wood and you're going to, you know, put two planks together and attach them with screws.

4:53It doesn't matter the details of which type of screwdriver you're using or the size of the screw within some limits and things like that. There are details that, in the end, don't matter as to what the end result will be, or the precise grain of the wood and things of that type. So you need to have some abstract level of representation within which you can make the prediction without having to predict every detail. right and that's why those J-PAL architectures have been advocating are useful so models like the Gaia one model from Wave actually makes prediction in an abstract representation space there's been a lot of work in that area for years also at FAIR but generally the abstract representation were pre-trained so the encoders that would take images from videos and then encode them into some representation or trained in some other way.

5:47And the progress we've made over the last six months in self-supervised learning for images and video is that now we can train the entire system to make those predictions simultaneously. So we have systems now that can learn good representations of images. And the basic idea is very simple. You take an image. You run it through an encoder. then you corrupt that image, you mask parts of it, for example, you transform it in various ways, blur it, change the colors, change the framing a little bit, and you run that corrupted image through the same encoder or something very similar. And then you train the encoder to predict the features of the complete image from the features of the corrupted one.

6:42You're not trying to reconstruct the perfect image. You're just trying to predict the representation of it. And this is different. This is not generative in the sense that it does not produce pixels. And that's the secret to getting cell supervision to work in the context of images and video. You don't want to be predicting pixels. It doesn't work. You can't put a pixel as an afterthought, which is what the Gaia system is doing by sticking a decoder on it and with some diffusion model that will produce a nice image. But that's kind of a second step. If you train the system by predicting pixels, you just don't get good representations.

7:19You don't get good predictions. You get blurry predictions most of the time. So that's what makes learning from images and video fundamentally different from learning from text. Because in text, you don't have that problem. It's easy to predict words, even if you cannot do a perfect prediction because language is discrete. So language is simple compared to the real world. And, you know, there's a lot written right now about the energy required in the computational resources, GPUs required to train language models. Is it less in training a world model like using iJEPA architecture? Well, it's hard to tell because there is no equivalent training procedure, self-supervised training procedure for video, for example, that does not use JEPA.

8:18The ones that are generative don't really work. Yeah. Yeah. Well, but this architecture could also be applied to language, couldn't it? Oh, yeah, absolutely. Yeah. So you could very well use a JPA architecture that makes prediction in representation space and apply to language. Yeah, definitely. And in that case, would it be less computationally intense than training a large language model? It's possible. it's not entirely clear either. I mean, there is some advantage regardless of what technique you're using to making those models really big. They just seem to work better if you make them big.

9:06So if you make them bigger, right? So scaling is useful. Contrary to some claims, I do not believe that scaling is sufficient. So in other words, we're not going to get anywhere close to human level AI. by in fact, not even animal-level AI, by simply scaling up language models, even multimodal language models that we apply to video. We're going to have to find new concepts, new architectures. And I've written sort of a vision paper about this a while back, of a sort of different type of architecture that would be necessary for this. So scanning is necessary, but not sufficient. And we're missing some basic ingredients to get to human-level AI.

9:59We're fooled by the fact that LLMs are fluent. And so we think that they have human-level intelligence because they can manipulate language, but that's false. And in fact, there's a very good symptom for this, which is that we have systems that can pass the bar exam, but answering questions from text by basically regurgitating what they've learned, more or less by road. But we don't have completely autonomous level five self-driving cars, or at least no system that can learn to do this in about 20 hours of practice, just like any 17-year-old. Yeah. And we certainly don't have any domestic robot that can clear out the dinner table and fill up the dishwasher, a task that any 10-year-old can learn in one shot.

10:53So clearly, we're missing something big. And that something is an ability to learn how the world works. And the world is much more complicated than language. And also being able to plan and reason. basically having a mental world model that goes on that allows to plan and predict consequences of actions that's what we're missing it's going to take a while before we figure this out you were on another paper that talked about augmented language models and in the embodied Turing test was that the same paper, the embodied Turing test can you talk about that first of all what is the embodied Turing test I didn't quite understand that well okay it's a different concept but it's basically the idea that you it's based on the Moravec paradox right so Moravec many years ago noticed that things that appeared difficult for humans turned out to sometimes be very easy for computers to do, like playing chess, much better than humans.

12:10Or, I don't know, computing integrals or whatever, certainly doing arithmetics. But then there are things that we take for granted as humans that we don't even consider them intelligent tasks that we are incapable of reproducing with computers. And so that's where the embodied Turing test comes in. Like, you know, observe what a cat can do or how fast a cat can learn new tricks or how a cat can plan to jump on a bunch of different furniture to get to the top of wherever it wants to go. That's an amazing feat that we can't reproduce with robots today. So that's kind of the embodied Turing test, if you want.

12:53Can you make a robot that can behave, have behaviors that are distinguishable from those of animals, first of all, and can acquire new ones with the same efficiency as animals. Then the augmented data lamp paper is different. It's about how do you sort of minimally change large language models so that they can use tools, so they can, to some extent, plan actions. Like, you know, you need to compute the product of two numbers, right? You just call a calculator and you know you're going to get the product of those two numbers. And LLMs are notoriously bad for arithmetic, so they need to do this kind of stuff.

13:35Or do a search, you know, using a search engine or database lookup or something like this. So there's a lot of work on this right now and it's somewhat incremental. Like, you know, how can you sort of minimally change LLM and take advantage of their current capabilities, but still augment them with the ability to use tools. Yeah. And I don't want to get too much into the threat debate, but you're on one side, your colleagues, Jeff and Yashur, on the other. I recently saw a picture of the three of you. I think you put that up on social media saying how, you know, you can disagree but still be friends.

14:19This idea of augmenting language models with stronger reasoning capabilities and the ability and agency, the ability to use tools is precisely what Jeff and Joshua are worried about. Can you just, why are you not worried about that? Okay. So first of all, this is not necessarily, what you're describing is not necessarily what they are afraid of. They are alerting people and various governments and others about various dangers that they perceive. Okay, so one danger, one set of danger are relatively short term. There are things like, you know, bad people will use technology for bad things. What can bad people use powerful AI systems for?

15:15And one concern that, you know, governments have been worried about and intelligence agencies encounter intelligence and stuff like that is, you know, could badly intentioned organizations or countries use LLM to help them, I don't know, design pathogens or chemical weapons or other things or cyber attacks, things like that. Right now, those partners are not new. Those partners have been with us for a long time. And the question is what incremental help would AI systems bring to the table? So my opinion is that as of today, AI systems are not sophisticated enough to provide any significant can't help for such badly intentioned people, because those systems are trained with public data that is publicly available on the internet.

16:09And they can't really invent anything. They're going to regurgitate with a little bit of interpolation, if you want. But they cannot produce anything that you can't get from a search engine in a few minutes. So that claim is being tested at the moment. There are people who are actually kind of trying to figure out, like, is it the case that you can actually do something, you're enabled to do something more dangerous with sort of current AI technology that you can't do with a search engine? Results are not out yet. But my hunch is that, you know, it's not going to enable a lot of people to do significantly bad things.

16:52Then there is the issue of things like cogeneration for cyber attacks and things like this. And those problems have been with us for years. And the interesting thing that most people should know, also for disinformation or attempts to corrupt the electoral process and things like this. And what's very important for everyone to know is that the best countermeasures that we have against all of those attacks currently use AI massively. So AI is used as a defense mechanism against those attacks. It's not actually used to do the attacks yet. And so now it becomes the question of who has the better system?

17:33Like are the countermeasures, is the AI used by the countermeasures significantly better than the AI used by the attackers so that the problem is satisfactorily mitigated? And that's where we are. Now, the good news is that there are many more good guys than bad guys. They're usually much more competent. They're usually much more sophisticated. They're usually much more better funded. And they have a strong incentive to take down the attackers. So it's a game of cat and mouse, just like every security that's ever existed. There's nothing new there. Okay. Nothing qualitatively new. Yeah. But then there is the question of existential risk, right?

18:23And this is something that both Jeff and Joshua have been thinking of fairly recently. So for Jeff, it's only sort of just before last summer that he started thinking about this. Because before, he thought, he was convinced that the kind of algorithms that we had were significantly inferior to the kind of learning algorithm that the brain used. and the epiphany he had was that, in fact, no, because looking at the capabilities of large language models, they can do pretty amazing things with a relatively small number of neurons and synapses. He said, maybe they're more efficient than the brain and maybe the learning algorithm that we use, backpropagation, is actually better than whatever it is that the brain uses.

19:06So he started thinking about what are the consequences but that's very recent and in my opinion, he hasn't thought about this enough. And Yoshua went to a similar epiphany last winter, where he started thinking about the long-term consequences and came to the conclusion also that there was a potential danger. They're both convinced that AI has enormous potential benefits. They're just worried about the dentures. And they're both worried about the dentures because they have some doubts about the ability of our institutions to do the best with technology. whether they are political, economic, geopolitical, financial institutions or industrial to do the right thing, to be motivated by the right thing.

20:03So if you trust the system, if you trust humanity and democracy, you might be entitled to believe that society is going to make the best use of future technology. If you don't believe in the solidity of those institutions, then you might be scared. I think I'm more confident in humanity and democracy than they are. And whatever current systems than they are. I've been thinking about this problem for much longer, actually. since at least 2014. So when I started FAIR at Facebook at the time, it became pretty clear pretty early on that deploying AI systems was going to have big consequences on people and society.

20:57And we got confronted to this very early. And so I started thinking about those problems very early on. Things like countermeasures against bias in AI systems, systematic bias, countermeasures against attacks, or detection of hate speech in every language, things like that. These are things that people at FAIR worked on and then were eventually deployed. Just to give you an example, the proportion of hate speech that was taken down automatically by AI systems five years ago, no, in 2017, was about 20 to 25%. Last year, it was 95%. And the difference is entirely due to progress in natural language understanding, entirely grew to transformers that are pre-trained self-supervised and can essentially detect hate speech in any language.

21:47Not perfectly. Nothing is perfect. It's ever perfect. But AI is used massively there, and that's the solution. So I've started thinking about those issues, including existential risk, very early on. In fact, in 2015, early 2016, actually, I organized a conference hosted at NYU on the future of AI, where a lot of those questions were discussed. I invited people like Nick Bockstrom and Eric Schmidt and Mark Schreffer, who were the CTO of Facebook at the time, Demisa Sabis, a lot of people, both from the academic and AI research side and from the industry side. and there were two days, a public day and kind of a more private day.

22:31What came out of this is the creation of an institution called the Partnership on AI. So this is a discussion I had with Demi Sesebis, which was, you know, would it be useful to have a forum where we can discuss before they happen, sort of bad things that could happen as a consequence of deploying AI? Pretty soon we brought on board Eric Horvitz and a bunch of other people, And we co-founded this thing called Apprenticeship on AI, which basically has been funding studies about AI ethics and consequences of AI and publishing guidelines about how you do it right to minimize harm. So this is not a new thing for me.

23:13I've been thinking about this for 10 years, essentially. Whereas for Joshua and Jeff, it's much more recent. Yeah. Yeah. But nonetheless, this augmented AI or augmented language models that have stronger reasoning and agency raises the threat, regardless of whether or not it can be countered to a higher level. Right. Okay. So I guess the question there becomes, what is the blueprint of future AI systems that will be capable of reasoning and planning, will understand how the world works, will be able to use tools and have agency and things like that, right? And I tell you, they will not be autoregressive LLMs.

24:06So the problems that we see at the moment of autoregressive LLM, the fact that they hallucinate, they sometimes say really stupid things. They don't really have a good understanding of the world. People claim that they have some simple world model, but it's very implicit and it's really not good at all. For example, you can tell an LLM that A is the same as B, and then you ask if B is the same as A, and it will say, I don't know, or no. Those things don't really understand logic or anything like that. So the type of system that we're talking about that might approach animal-level intelligence, let alone human-level intelligence, have not been designed.

24:53They don't exist. And so discussing their danger and their potential harm is a bit like, you know, discussing the sex of angels at the moment. Or to be a little more accurate, perhaps, it would be kind of like discussing how we're going to make transatlantic flight at near the speed of sound safe when we haven't yet invented the turbojet in 1925. Yeah. Yeah. Like, you know, we can speculate, but, you know, how did we make turbojet safe? It required decades of really careful engineering to make them incredibly reliable. And, you know, now we can, you know, run like halfway around the world with a two engine turbojet aircraft.

25:45I mean, that's an incredible feat. And it's not like people were discussing sort of philosophical questions about how you make turbojet safe. It's just really careful and complicated engineering that none of us would understand.

26:03So, you know, how can we ask the AI community now to explain how AI systems are going to be safe? We haven't invented them yet. No. Okay. That said, I have some idea about how we can design them so that they have these capabilities. And as a consequence, how they will be safe. I call this objective-driven AI. So what that means is essentially systems that produce their answer by planning their answer so as to satisfy an objective or a set of objectives. So this is very different from current LLMs. current LL answers produce one word after the other or one token, which is, which are subordinates.

26:46It doesn't matter, right? They don't really think and plan ahead as we, as we said before, they just produce one word after the other. That's not controllable. The only thing we can do is see if what they've produced, like check if what they've produced, uh, satisfy some criterion, uh, or set a criteria and then not produce an answer, produce a non-answer. If the answer that was produced is inappropriate. But we can't really force them to produce an answer that satisfies a set of objectives. So objective-driven AI is the opposite. The only thing that the system can produce are answers that satisfy a certain number of objectives.

27:31So one objective would be, did you answer the question? Another objective could be, is your answer understandable by a 13-year-old because you're talking to a 13-year-old. Another would be, is this, I don't know, terrorist propaganda or something? You know, you can have a number of criteria like this, guardrails, that would guarantee that the answer that's produced satisfies certain criteria, whatever they are. Okay. Same for a robot. You could guarantee that the sequence of actions that is produced will not hurt anyone. Like, you can have very low-level, you know, guardrails of this type that say, okay, you have humans nearby and you're cooking, so you have a big knife in your hand.

28:13Don't flare your arms. That would be a very simple guardrails to impose. And you can imagine having a whole bunch of guardrails like this that will guarantee that the behavior of those systems would be safe and that their primary goal would be to be basically subservient to us. So I do not believe that we'll have AI systems that can work that will not be subservient to us will define their own goals. They will define their own sub-goals, but those sub-goals would be sub-goals or goals that we set them and will not have all kinds of guardrails that will guarantee the safety. And we're not going to have, it's not like we're going to invent a system and make a gigantic one that we know will have human level AI and just turn it on.

28:59And then from the next minute it's going to take over the world. That's completely preposterous. What we're going to do is try with small ones. You know, maybe as smart as a mouse or something, maybe a dog, maybe a cat, maybe a dog, and walk her way up and then, you know, put some more guardrails. Basically, like we've engineered, you know, more and more powerful and more reliable turbojets. It's an engineering problem. Yeah. Yeah. You were also on a paper. Maybe this is the one that talked about the embodied Turing test on neuro AI. can you explain what neuro AI is okay well it's the idea that we should get some inspiration from neuroscience to build AI systems and that there is something to be learned from neuroscience and from cognitive science to drive the design AI system, some inspiration, okay?

30:04Something to be learned, as well as the other way around. So what's interesting right now is that the best models that we have of how, for example, the visual cortex works is convolutional neural networks, which are also the models that we use to recognize images primarily in artificial systems. So there is kind of information kind of being exchanged both ways because there's one uh you know one way to make progress in ai is to kind of ignore nature and and just you know kind of uh try to solve problems in a sort of engineering fashion if you want uh i found interaction with neuroscience always thought-provoking so you don't want to be copying nature very too closely because there are details in nature that are irrelevant And there are principles on which, you know, natural intelligence is based that we haven't discovered.

31:01So, but there is some inspiration to have, certainly in convolutional networks, inspired by the architectural individual cortex. The whole idea of neural net and deep learning came out of the idea that, you know, intelligence can emerge from a large collection of simple elements that are connected with each other and change the nature of their interactions. That's the whole idea, right? Right. So inspiration from neuroscience certainly has been extremely beneficial so far. And the idea of neural AI is that you should keep going. You don't want to go too far. So going too far, for example, is trying to reproduce some aspect of the functioning of neurons with electronics.

31:44I'm not sure that's a good idea. I'm skeptical about this, for example. So your research right now, your main focus is on furthering the JEPA architecture into other modalities or where are you headed? Yeah, so the long term goal is to get machines to be as intelligent and learn as efficiently as animals and humans. And the reason for this is that we need this because we need to amplify human intelligence. And so intelligence is the most needed commodity that we want in the world, right? And so we could possibly bring a new renaissance to humanity if we could amplify human intelligence using machines, which we are doing already with computers, right?

Read the full transcript

32:38I mean, that's pretty much what they've been designed to do. But even more, imagine a future where every one of us has an intelligent assistant with us at all times. They can be smarter than us. We shouldn't feel threatened by that. We should feel like we are a director of a big lab or a CEO of a company that has a staff working for them of people who are smarter than themselves. I mean, we're used to this already. I'm used to this, certainly. working with people who are smarter than me. So we shouldn't feel threatened by this, but it's going to empower a lot of us, right, and humanity as a whole.

33:22So I think that's a good thing. That's the overall practical goal, if you want, right? Then there's a scientific question that's behind this, which is really what is intelligence and how you build it? And then which is, you know, how can a system learn the way animals and humans seem to be learning so efficiently? And the next thing is, how do we learn how the world works? By observation, by watching the world go by through vision and all the other senses. And animals can do this without language, right? So it has nothing to do with language. It has to do with learning from sensory percepts. And learning mostly without acting, because any action you take can kill you.

34:04So it's better to be able to learn as much as you can without actually acting at all, just observing, which is what babies do in the first few months of life. They can't hardly do anything. Right. So they mostly observe and learn how the world works by observation. So what kind of learning takes place there? So that's obviously kind of self-supervised. Right. It's learning by prediction. That's an old idea from cognitive science. And the thing is, you know, we can learn to predict videos, but then we notice that predicting videos, predicting pixels in a video is so finishly complicated that it doesn't work.

34:38And so then came this idea of JEPA, right? Learn representations so that you can make predictions in representation space. And that turned out to work really well for learning image features. And now we're working on getting this to work for video. And eventually, we'll be able to use this to learn world models where you show a piece of video, and then you say, I'm going to take this action, predict what's going to happen next in the world.

35:09which is a bit what the Gaia system from Wave is doing at a high level, but we need this at various levels of abstraction so that we can build systems that are more general than autonomous driving. Okay. Yeah. And it's my fault, so I won't go over the hour, But is it conceivable that someday there will be a model that you maybe embodied in a robot that is ingesting video from its environment and learning as it's just continuously learning and getting smarter and smarter and smarter? Yeah, I mean, that's kind of a bit of a necessity. The reason being that, you know, even if you train a system to have a world model that can predict what's going to happen next, the world is really complicated.

36:14And there's probably all kinds of situations that, you know, the system hasn't been trained on and need to, you know, fine tune itself as it goes. so you know animals and humans do this early in life by by playing so play is a way of learning your world model in situations that basically won't hurt you but then during life of course you know when we learn to drive there's all kinds of mistakes that we do initially that we don't do after having some experience. And that's because we're fine-tuning our world model to some extent. We're learning a new task. We're basically just learning a new version of our world model.

37:01So yeah, this type of continuous learning is going to have to be present. But the overall power and intelligence of the system would be limited by how much a big of a neural net it's using and various other constraints, computational constraints, basically. You know, you're still young. I'm not sure about that. Well, you're younger than Jeff, let me put it that way. I'm younger than Jeff. I'm older than Joshua. But the progress you've made on world models is fairly rapid from my point of view, watching it. But are you hopeful that within your career you'll have embodied robots that are building world models through their interaction in reality and then being able to?

38:01Well, I guess the other question on world models, do you then combine it with a language model to do reasoning? Or is the world model able to do reasoning on its own? But are you hopeful that in your career, you'll get to the point where you'll have this continuous learning in a world model? Yeah, I sure hope so. I might have another 10 useful years or something like this in research before my brain turns into Deschamil sauce or something like that. 15 years if I'm lucky. Or perhaps less. But yeah, I hope that there's going to be breakthroughs in that direction during that time. Now, whether that will result in the kind of artifact that you're describing, you know, robots that can, like, you know, domestic robots, for example, or self-driving cars that can run fairly quickly by themselves.

39:02I don't know, because there might be all kinds of obstacles that we have not envisaged that may appear on the way. You know, it's a constant in the history of AI that you have some new idea and a breakthrough, and you think that's going to solve all the world's problems. And then you kind of hit a limitation and you have to go beyond that limitation. So it's like, you know, you're climbing a mountain. You find a way to climb the mountain that you're seeing. And you know that once you get to the top, you will have the problem solved because now it's, you know, a gentle slope down. And once you get to the top, you realize that there is another mountain behind it that you hadn't seen.

39:47So that's been the history of AI, right? Where people have come up with sort of new concepts, new ideas, new way to approach AI reasoning, whatever, perception. And then realized that their idea basically was very limited. And so, you know, this inevitably, we're trying to figure out, like, what's the next revolution in AI? That's what I'm trying to figure out. So, you know, learning how the world works from video, having systems that have world model that allow systems to reason and plan. and there's something I want to be very clear about which is an answer to your question which is that you can have systems that reason and plan without manipulating language.

40:41Animals are capable of amazing feats of planning and also to some extent reasoning. They don't have language. At least most of them don't. And so many of them don't have culture. because they are mostly solitary animals. So, you know, it's only the animals that have some level of culture. So the idea that the system can plan and reason is not connected with the idea that you can manipulate language. Those are two different things. It needs to be able to manipulate abstract notions, but those notions do not necessarily correspond to linguistic entities like words or things like that. We can have mental images if you want of things.

41:30Like you do ask a physicist or a mathematician how they reason. It's very much in terms of mental models that have nothing to do with language. Then you can turn things into language, but that's a different story. That's a second step.

41:48So we're going to have to figure out how to do this, reasoning, hierarchical planning in machines, reproduce this first. And then, of course, sticking language on top of it will help. It will make those systems smarter and allow us to communicate with them and teach them things and they're going to be able to teach us things and stuff like that. But this is a different question, really. The question of how we organize AI research going forward, which is somewhat determined by how afraid people are of the consequences of AI. So if you have a rather positive view of the impact of AI on society and you trust humanity and society and democracies to use it in good ways, then the best way to make progress is through open research.

42:36And for the people who are afraid of the consequences, whether they are societal or geopolitical, they're putting pressure on governments around the world to regulate AI in ways that basically limit access, particularly of open source code and things like that. And it's a big debate at the moment. I'm very much on the side. So is Meta very much on the side of open research. Yeah, actually, that was something I was going to ask you, now that you've brought it up. Because I've been talking to people about this, And there is a view that aside from the risks of open source, you know, again, Jeff Hinton saying, you know, would you open source thermonuclear weapons?

43:25Aside from that, the question is to whether open source can marshal the resources to compete with proprietary models because of the tremendous resources required for when you're scaling these models. and there's a question as to whether or not Meta will continue to open source future versions of LAMA or not continue to open source but whether it'll continue to invest the resources needed to push the open source models so what do you think about that? Okay, there's a lot to say about this so first thing is there's no question that Meta will continue to invest the resources to build better and better AI systems because it needs it for its own products.

44:22So the resources will be invested. Now, the next question is, will we continue to open source the base models? And the answer is probably yes, because that creates an ecosystem on top of which an entire industry can be built. And there is no point having 50 different companies building proprietary closed systems when you can have one good open source-based model that everybody can use. It's wasteful and it's not a good idea. And another reason for having open source models is that nobody has no entity as powerful as it thinks it is as a monopoly on good ideas. And so if you want people who can have good new innovative ideas to contribute, you need an open source platform.

45:16If you want the academic world to contribute, you need open source platforms. If you want the startup world to be able to build customized products, you need open source based models because they don't have the resources to build, to train large models. right okay and then there is the history that shows that for for foundational technology for infrastructure type technology open source always wins right it's true of the software infrastructure of the internet in the early 90s and mid 90s there was a big battle between sun microsystems and microsoft to produce the deliver the software infrastructure of the internet, you know, operating systems, web servers, web browsers, and, you know, various server side and client side frameworks, right?

46:10They're both lost. Nobody is talking about them anymore. The entire world of the web is using Linux and Apache and MySQL and JavaScript and, you And even the basic core code for web browser is Open Source. So Open Source won by a huge margin. Why? Because it's safer, gathers more people to contribute. All the features are necessary. It's more reliable. Vulnerabilities are fixed faster. And it's customizable. So anybody can customize Linux to run on whatever hardware they want, right? So open source wins. But same for AI. It's going to be the same thing. It's inevitable. The people now who are climbing up like OpenAI, their system is based on publications from all of us and from open platforms like PyTorch.

47:18ChaiGPT is built using PyTorch. PyTorch was produced originally by Meta. Now it's owned by the Linux Foundation. It's open source. They've contributed to it, by the way. Their LLM is based on transformer architectures invented at Google. All the tricks to kind of train all those things came out of various papers from all kinds of different institutions, including academia. All the fine tuning techniques, same. So nobody works in a vacuum. The thing is, nobody can keep their advance and their advantage for very long if they are secretive. Yeah, except that with these models, because they're so compute intensive and they cost so much money to train, you need somebody like Meta who's going to be willing to build them and open source them.

48:11And that's why when I was asking whether they'll continue, obviously, Meta will continue building resource-intensive models, but the question is whether they'll continue to open source. I tell you, the only reason why Meta could stop open sourcing models are legal. So if there is a law that atlaws open source AI systems above a certain level of sophistication, then of course we can do it. If there are laws that in the US or across the world makes it illegal to use public content to train AI systems, then it's the end of AI for everybody, not just for the open source. or at least the end of the type of AI that we are talking about today.

49:09We might have new AI in the future, but that don't require as much data. And then there is liability. If you believe in the kind of that someone doing something bad with an AI system that was open sourced by Meta, then Meta is liable. then Meta will have a big incentive not to release it, obviously. So the entire question about this is around legal reasons and political decisions. But on the idea of open source winning, don't you need more people or more companies like Meta building the foundation models and open sourcing them? Or could an open source ecosystem win based on a single company building the models?

50:00No, I mean, you need two or three. and there are two or three, right? I mean, there is Hugging Face. There is Mistral in France who's also embracing open source LLM. They're very good at LLM. It's a small one, but it's very good. There is, you know, academic efforts like Lion. They don't have all the resources they need, but they, you know, they collect the data that is used by everyone. So everybody can contribute. One thing that I think is really important to understand also is that there is a future in which I described earlier, in which every one of us, every one of our interactions with the digital world would be mediated by an AI assistant.

50:41And this is going to be true for everyone around the world, right? Everyone who has any kind of smart device. Eventually, it's going to be in our augmented reality glasses, but for the time being, in our smartphones, right? And so imagine that future where you are, I don't know, from Indonesia or Senegal or France, and your entire digital diet is done through the mediation of an AI system, your government is not going to be happy about it. Your government is going to want the local culture to be present in that system. It doesn't want that system to be closed sourced and controlled by a company on the West Coast of the US.

51:33Yeah. Okay. So just for reasons of preserving the diversity of culture across the world and not having our entire information diet being biased by whatever it is that some company on the West Coast of the US thinks, there's going to need to be open source platforms. and they're going to be predominant at least outside the US for that reason. Including China, right? There is all those talks about, oh, what if China puts their hands on our open source code? I mean, China wants control over its own LLM because they don't want their citizen to have access to certain type of information. So they're not going to use our LLMs.

52:18They're going to train theirs that they already have. Yeah. And nobody is particularly ahead of anybody else by more than about a year. Yeah. And China is pushing open source. I mean, they're very pro-open source within their ecosystems. Some of them. There's no unified opinion there. But I mean, it's the same in the West, right? There are some governments that are too afraid of the risks. or are thinking about it and some others that are all for open source because they see this as the only way for them to have any influence on the type of information and culture that would be mediated by those systems.

53:07So it's going to have to be like Wikipedia, right? Wikipedia

53:13is built by millions of people who contribute who are from all around the world in all kinds of languages. And it has a system for sort of vetting the information. The way AI systems of the future will be taught and will be fine-tuned will have to be the same way. It will have to be crowdsourced because something that matters to a farmer in southern India is probably not going to be taken into account by the fine-tuning done by some company on the west coast of the US. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested.

53:55So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud. Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course, nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds.

54:42If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash IonAI. That's E-Y-E-O-N-A-I, all run together, oracle.com slash IonAI. That's it for this episode. I want to thank Jan for his time. If you want to read a transcript of this conversation, you can find one on our website, IonAI. That's E-Y-E hyphen O-N dot A-I. And remember, the singularity may not be near, but A-I is changing your world. So best pay attention.

From the publisher

This episode is sponsored by Oracle. AI is revolutionizing industries, but needs power without breaking the bank. Enter Oracle Cloud Infrastructure (OCI): the one-stop platform for all your AI needs, with 4-8x the bandwidth of other clouds. Train AI models faster and at half the cost. Be ahead like Uber and Cohere.

If you want to do more and spend less like Uber, 8x8, and Databricks Mosaic - take a free test drive of OCI at https://oracle.com/eyeonai

 

Welcome to episode 150 of the 'Eye on AI' podcast. In this episode, host Craig Smith sits down with Yann LeCun, a Turing Award winner who has been instrumental in advancing convolutional neural networks and whose work spans machine learning, computer vision, and more.

Tune is as Craig and Yann explore the intricacies of AI, world models, and the challenges of continuous learning.

In this episode, Yann delves deep into the concept of a "world model" - systems that can predict the world's future states, allowing agents to make informed decisions. The discussion transitions to the challenges of training these models, particularly when dealing with diverse data like text and images. We then discuss the computational demands of modern AI models, with Yann highlighting the nuances between generative models for videos and language. 

He also touches upon the idea of the "Embodied Turing Tests" and how augmented language models can bridge the gap between human-like behavior and computational efficiency.The spotlight then shifts to pressing concerns surrounding the open-source nature of AI models, with Yann articulating the legal ramifications and the future of open-source AI. Drawing from global perspectives, including China's stance on open-source, Yann underscores the imperative for a collaborative approach in the AI space, ensuring it's reflective of diverse global needs.

 

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

(00:00) Preview, Oracle and Introduction

(02:42) Decoding The World Model and Gaia 1 

(07:43) Energy and Computational Demands of AI

(08:06) Video vs. Text Processing & True AI Capabilities

(11:17) Embodied Turing Test & Augmented LLMs

(15:38) Is AI a Threat To Society?

(25:04) Where is AI Development Headed?

(31:06) Interplay of Neuroscience and AI**

(33:33) Yann's Vision, JEPA, and Learning Challenges

(39:05) Yann's Career, AI Progress, and Challenges

(44:47) The Open Source Debate in AI

(55:30) Oracle Cloud Infrastructure

More from Eye On A.I.

All 266 episodes
#150: Yann LeCun on World Models, AI Threats and Open-SourcingEye On A.I. · 56 min
Listen in VO