#174 Tianmin Shu: How World Models are Shaping The Future of AI

10 Mar 2024 · 49 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast - Episode #174 Summary

Episode Title

Tianmin Shu: How World Models are Shaping The Future of AI

Host

Craig S. Smith

Guest

Tianmin Shu, Assistant Professor at Johns Hopkins University

Episode Description

In this episode, host Craig Smith converses with Tianmin Shu about the concept of world models and their significance in both artificial and human intelligence. The discussion highlights Shu's academic journey and innovative work in AI, focusing on the integration of cognitive science into AI development and its implications for future technologies.

---

Key Topics Discussed

  1. Introduction to Tianmin Shu
  2. Background in AI and cognitive science.
  3. Recent academic role at Johns Hopkins University.
  4. Previous postdoctoral experience at MIT and PhD from UCLA.
  1. World Models
  2. Definition: Internal representations that allow entities (both humans and AI) to understand, simulate, and predict their environment.
  3. Origin: Rooted in cognitive science and developmental psychology, particularly theories by researchers like Liz Belkey and Josh Tenenbaum.
  4. Function: Enable simulation of future scenarios based on current knowledge of physical laws (e.g., object permanence).
  1. Integration with AI
  2. World models can enhance AI's ability to interpret and interact within environments.
  3. Potential applications include household environments and social learning where AI systems learn alongside humans.
  1. Large Language Models (LLMs)
  2. Discussion on how LLMs can be integrated with world models.
  3. LLMs provide reasoning capabilities while world models predict future states.
  4. Challenges in using LLMs alone for complex scenarios.
  1. Research Projects
  2. Tianmin's work on creating world models for household environments.
  3. Emphasis on multimodal sensory data (vision, audio, touch) for training agents.
  4. The significance of simulation environments in gathering experiential data for model training.
  1. Implementation Challenges
  2. The complexity of creating a single world model capable of handling the variability of real-world interactions.
  3. Need for domain-specific models rather than a one-size-fits-all approach.
  1. Future Directions
  2. Exploration of social learning in AI, enabling models to learn and cooperate with humans.
  3. Potential for commercial applications in robotics and virtual assistants.
  1. Commercial Viability
  2. Current status of research and its transition to real-world applications.
  3. Promising developments in indoor navigation robots and household automation.

---

Key Takeaways

  • World Models are essential for the development of AI that functions similarly to human intelligence.
  • The integration of cognitive science into AI research is crucial for creating more intuitive and interactive systems.
  • The development of agents that can understand and predict human behavior is still an ongoing challenge that requires further exploration.
  • The synergy between world models and large language models presents a potential pathway for advancements in AI reasoning and interaction capabilities.
  • Social learning and the ability of AI to learn from humans represent significant areas for future research and application.

---

Conclusion The episode concludes with a reminder of the evolving landscape of AI and the importance of understanding its implications for society. With the integration of advanced models and cognitive principles, the next frontier in AI research promises to transform human-technology interaction significantly.

---

Links

  • [Eye On A.I. Twitter](https://twitter.com/EyeOn_AI)
  • [Craig Smith's Twitter](https://twitter.com/craigss)
  • [1Password Trial Offer](https://www.1password.com/eyeonai)

---

Feel free to reach out to the podcast for more insights and discussions on the topic of artificial intelligence and its future!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So I think a lot about not only building artificial intelligence, but also how human intelligence can inform us to build better AI. And I think that the idea of building a model that actually originally comes from cognitive science. And then especially I have a kind of strong interest in the social aspect of AI and also human intelligence. So how we can build human level kind of social intelligence so that we can allow systems to understand human and really cooperate with human in the same way that we can understand and cooperate with each other. We have our thoughts in our minds, and we think about those thoughts in the form of language.

0:40We can also write down those thoughts and use our language to communicate with others about our thoughts. Hi, I'm Craig Smith, and this is Eye on AI. In this episode, I speak with Tianmin Xu, a rising figure in the AI research community, soon to join Johns Hopkins University with an impressive background that includes a postdoc at MIT and a PhD from UCLA. Tinman works at the intersection of AI and cognitive science with a particular focus on societal aspects of both human and artificial intelligence. Our conversation revolved around the concept of world models, a term rooted in developmental psychology and cognitive science, and their integration into AI.

1:32Tianmin elaborated on how these models, coupled with large language models, can enable the creation of agents capable of interpreting and interacting with the world in human-like ways. He sheds light on his projects involving world models for household environments and the potential of multimodal sensory data in their training. The episode also delved into the challenges and future trajectories in AI, especially in the realm of social learning, where AI models learn alongside of and from humans. Tiamen discusses research driving the future of AI and human-robot interaction, and I hope you find our conversation as insightful as I did.

2:21At home, at work, we all know one person who's password challenged. Sticky note reminders, emailing passwords, reusing passwords, using the word password as their password. Because data breaches affect everyone, you need one password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach costs millions of dollars. 1Password secures every sign-in to save you time and money. 1Password lets you switch between iPhone, Android, Mac, and PC with convenient features like autofill for quick signups.

3:14All you have to remember is the one strong account password that protects everything else, your logins, your credit cards, secure notes, or the office Wi-Fi password. One password generates as many strong unique passwords as you need and securely stores them in an encrypted vault that only you have access to. I use 1Password, and you should too. 1Password's award-winning password manager is trusted by millions of users and over 100 ,000 businesses from IBM to Slack. It beat out 40 other options to become Wirecutter's top pick for password managers. Plus, regular third-party audits and the industry's largest bug bounty keep 1Password at the forefront of security.

4:03Right now, my listeners get a two-week free trial at 1password.com slash IonAI. So why don't you start, Tianmin, by introducing yourself, and then I'll ask questions from there. Give a little of your educational background and where you're working today and what you're working on. So I'm actually going to join Hopkins, the CS department, and also the Carlson department as an assistant professor. But previously I did my postdoc at MIT, and before that I did my PhD at UCLA. So I'm working at the intersection between AI and cognitive science. So I think a lot about not only building artificial intelligence, but also how human intelligence can inform us to build better AI.

5:00And I think that the idea of building a model that actually originally comes from cognitive science. And then especially I have a kind of strong interest in the social aspect of AI and also human intelligence. So how we can build human level kind of social intelligence so that we can allow citizens to understand human and really cooperate with human in the same way that we can understand and cooperate with each other. And again, I think that's something that's one aspect is very challenging right now. And also it's not something that big companies are working on. and I think it's also something that will really benefit from having a world model, both in the sense of understanding how humans understand the world and also how we can understand humans that live in this world.

5:55Okay. Can we start, first of all, by talking about world models? I've had a few episodes on world models, But maybe you can, when did the idea of world models emerge? I mean, I've been following Jan LeCun's work, but I know it's much broader than that. When did that idea emerge? And can you give us just a very brief history of its development? So just to determine, I don't think that the history I can tell is probably not the full history. There's going to be some people at some point talking about world model that I don't know about. But the kind of history I'm familiar with is coming from developmental psychology and also cognitive science.

6:45So in developmental psychology, there's this theory about quodnarch. So it's actually a theory proposed by Liz Belkey, who is now a professor at Harvard. So she had this theory about, well, so since first, we have some cool knowledge about the world and about other agents that we can use as a foundation for developing more sophisticated intelligence that we can develop later in life. And one of them is to understand the physical world. So for example, we have some on physical common sense, like object permanence, we oftentimes support, you know, if you put something on table, it won't fall. So this kind of basic kind of physical common sense allow us to actually develop more sophisticated intelligence.

7:35And we have acquired those physical common sense very long, maybe even through evolution. Now it's unclear how much of it is actually come from. evolution, how much it comes from experiences living in life. Now, there's also this idea about having an internal world model that can allow you to simulate what could happen in the future. And you can use that kind of simulation to not only reason about future physical states, but also plan yourself better through this kind of simulation. And that comes from a lot of, I guess, work from my actually PD, not PD advisor, postdoc advisor, Josh Tenenbaum.

8:23So he had this idea that, you know, human has this internal intuitive physics engine. So we all know physics engine, which allow you to simulate what could happen in the physical world. Now, human, in our minds, we can have intuitive physics engine that allow you to simulate what could happen in the physical world, just like how you can simulate what could happen in the physical world using a game engine. Now, using that, they show that you can simulate physical events and use that kind of simulation results, adding some kind of noise on top of that. And you can use that to approximate how humans judge physical events pretty well.

9:04So they did experiments where they show a tower of blocks, and they asked people to judge whether the block will fall or not. If they are fall, which way they are fall to. And then it turns out that if you actually simulate that using a game engine, and ask a little bit of noise on top of that because people have imperfect kind of simulation. So a kind of physical simulation plus a noise or some kind of uncertainty, that's a possibility human judgment really well. And I think there's also recent kind of collaboration with a neuroscience at MIT, Nancy Canvisher. They show that there's some kind of evidence in our brain that maybe our brain is using actually a game engine to simulate what could happen in a physical world.

9:56Yeah. And that engine to simulate, what's the computer code or algorithmic architecture behind that? I think one way to understand it is that you can write down code to describe a genetic process. So in this case, you can describe a genetic process as a world model. So, you know, you sample objects, you kind of simulate force, and then the journey model will tell you after you apply this force, what will happen next, will be the state of these objects you just sampled. So, you can write down this journey process as a program, and then when you run this program, you will sample this sample from the journey model.

10:53And it will tell you what's the distribution of the future states of these objects after you run the simulation. So that allows you to first of all simulate the agent model of the world, and second, it gives you attribution. And what you can do with attribution is you can and conduct probability inference. So to say, if you want to predict where these objects will fall in the next few steps, you can use that simulation, that distribution to tell you what is the likelihood of these objects falling into this place in the next few steps. You can also use this simulation to do planning, for example, waste uncertainty.

11:42So I think probably this object is going to fall this place, but maybe also in that place based on my simulation, based on the division coming from this property program. Now, given this uncertainty, how I can plan myself in order to fall and maybe catch these objects. Well, that's like one implementation, right? You use a physical simulation and that allows to sample feature states. And that can be implemented as a property program. I see. And then in your talk, you were talking about combining world models and large language models. And the language model would provide the reasoning and the world model would provide the predictions of future states.

12:34and then the language model would would pick which action for which to arrive at a future state is that right um well we kind of export um it's got different ways in which people have tried to use language model to implement as like a word model um so the the most i guess direct use is say, let's prompt Lens model with the current state, the action, and ask Lens model what could happen next. So this is the basic idea of a world model, right? You want to predict what happened given the current state and given our action. And then I think that the paper you mentioned is the reasoning as planning paper, the red paper.

13:20So the idea behind that is that why not we try to use language model to do this, to imagine what could happen after taking action at the current step. It works to a certain extent. I think in some simple scenarios it can work. I don't think at least poor language model is robust enough to serve as a world model on its own for more competitive scenarios. Okay. And then this most recent paper and the workshop at NeurIPS was about combining the language model with a world model. So the world model predicts the future states and the language model provides the reasoning. Did I understand that correctly?

14:02I don't think that's, I guess, the complete idea. So the idea is that we want to have experienced world models and also experienced Asian model that's built up on world model. And then using Asian model and world model, we can conduct model-based reasoning. Now, we want to explore the idea that how language model could be as a backend to implement world model and agent model. And of course, like the red paper, there has been efforts that try to use language model directly as vulnerable world model and agent models that can work in certain scenarios but not in more complex scenarios. Now, we want to explore, for example, how it can maybe enhance the model to serve as a better world models and agent models.

14:52And second, we want to also explore ways in which we can not only just use language, but also use other modalities, like the vision, even like touch or audios to really build multi-model world models. So obviously that's not going to be the kind of typical language model you see nowadays. Yeah. So can you talk a little bit about that? I mean, how you introduce the language model into this architecture with a world model in order to produce an agent? So I think I would say there are probably like two ways you can introduce this. One way, that's not coming from my own work, that's from my college at MIT.

15:42So they introduced the idea that use language as interface to translate language discussion about scenarios into a properties program so that you can conduct properties inference using that properties program. And again, properties programs serve as a joint model of the world and also about other agents. So this idea comes from a very classical idea in cognitive science called language of thought. So we have our thought in our mind, and we think about those thoughts in the form of language. We can also write down those thoughts and use our language to communicate with others about our thoughts.

16:24So then if you do the reverse process, we are given language about scenarios. and also a question about the scenarios, how you can conduct recently using internal thoughts, not as dirty as language reasoning, but you want to convert language into an internal thought and reasoning over that thought. So for example, you can describe a language, you can use the language to describe a physical situation. Like I just told you, you know, there's maybe a tower of blocks on the table. On the bottom, you have a small block. In the middle, you have a larger block. On top, you have an even larger block. Like an axial, can you imagine how stable this top of blocks can be, right?

17:15And I also described a different scenario. On bottom, you have a very large block. On the middle, you have a smaller kind of block. On top, you have a teeny tiny block. Then you can also imagine what could happen next. So notice I have shown you zero image about this, but you can, based on my language discussion, try to understand what I'm trying to describe here. So how we can convert that language into something like a thought process that allows to simulate what could happen next. So I think the colleagues at MIT, they come up with this idea that you can use language model to translate this language distribution into code.

18:01And so basically a program that will basically become a physical simulation. So the scenario I told you, after the language model, you'll get a code that allows you to simulate the physical situation. And you can simulate what could happen next. And that this again allows you to do model based kind of reasoning about the physical world. So I think that's a very neat idea. And it's also a very good use of language model because again, I don't think language model itself has all the knowledge about the physical world, but language model is really good at writing code. And we all know that. So, but, and then again, And it kind of translates, kind of do this inverse process where you can convert language into a thought.

18:49That's really interesting. And then how, and you guys have done this. This isn't theory. You've built these models to work together. What is the agency layer? I mean, that's the reasoning and the planning or decision-making. Then how do you add the agent on top of that? Right. So agent is modeled as a decision-making process, basically, in this framework, where you have your work model, your internal work model that you can use to see what could happen. you'll have your own goal and goal has a condition of goal you have a reward so your reward will try to help you to reach the goal but also in the process minimize the cost and then you have your belief about the world because an agent only has partial observation of all of the physical environments you don't know everything so you have a belief about what the actual physical state is so again belief is like a distribution of the physical state Now, condition on this, then you can try to make a plan that can, I guess, based on your first of all, your belief about how the actual vertical states look like.

20:16And then also try to maximize your rewards that allow you to reach the goal after doing simulations about different kinds of plans. Yeah. So I'm trying to think of it in, so that's an RL reinforcement learning engine in effect. Well, that doesn't have to be reinforcement. That can be a model-based planning. So the idea behind the model-based reinforcement or model-based planning is that you don't need to do a lot of trial and error in actual physical environments. You can do so using a simulation. You can simulate beforehand what could happen after taking different kind of actions. Then you can search for the best set of actions that allow you to reach the goal without actually trying out in a physical world.

21:09At home, at work, we all know one person who's password challenged. Sticky note reminders, emailing passwords, reusing passwords, using the word password as their password. Because data breaches affect everyone, you need 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach costs millions of dollars. 1 Password secures every sign-in to save you time and money. 1 Password lets you switch between iPhone, Android, Mac, and PC with convenient features like Autofill for quick sign-ups.

22:02All you have to remember is the one strong account password that protects everything else. Your logins, your credit cards, secure notes, or the office Wi-Fi password. 1Password generates as many strong, unique passwords as you need and securely stores them in an encrypted vault that only you have access to. I use 1Password, and you should too. 1Password's award-winning password manager is trusted by millions of users and over 100 ,000 businesses from IBM to Slack. It beat out 40 other options to become Wirecutter's top pick for password managers. Plus, regular third-party audits and the industry's largest bug bounty keep 1Password at the forefront of security.

22:51Right now, my listeners get a two-week free trial at 1Password.com slash IonAI. Right. That last part sounds like a little bit like AlphaGo, where you simulate various courses of action and then you search for the optimal solution and implement that solution. Is it similar in architecture to AlphaGo? I mean, the AlphaGo, it's not using a world model. Okay, so what is the core component of AlphaGo? It's the multi-color tree search, right? So that's the planning algorithm. That is not IELTS. That is planning. And the reason why you can do planning in that case is that you know the rule of the game already, right?

23:48So you know exactly this will be the next stage of the game after taking certain moves. And that will allow you to do forward simulation. You can think ahead for many, many steps and get the better plans by doing these kind of imagination. Now, you mentioned that there doesn't seem to be one model. There is a one model. The one model is very simple. It's just the rule of the game. And then you know exactly what could happen next. Now the more, and I think MCDS kind of for planning, always is very kind of intriguing because it's very general in the sense that if you have a reward, if you have one model that allows you to do forward simulation, in theory you can do anything, right?

24:39If you do enough simulation, you can get pretty good planning. And if you add a little bit learning like AlphaGo, you can also make the search even faster. Now the downside, the challenge is that in the real world, you don't have that perfect word model. But if you do build a word model for certain domains, then you can plug in that word model to do a monoculture search or any kind of model-based planning. Yeah. How far have your implementations of this gone? I mean, you've built a model with all these elements. Does its efficacy depend on how large the model is, as is the case with language models?

25:36Is it the larger the model, the better performance? Or is this more a matter of just optimizing hyperparameters and stuff like that? Right. You mean like building, actually learning a role model that can be used for, I guess, learning a role model that can be used for a very broad set of domains and not just specific domains. I think we are not remotely there yet. The reason is, for example, there's just so many different kinds of things you can model in the world. You have so many different objects, and even if you think about one specific object, they have so many different kinds of properties you can try to model, like the materials, the textures, the weight of objects, the shape, and all these different kinds of properties.

26:32right um and depending on your test you might need to model the objects in different uh using different reputation of the objects so for instance if you only want to say um carry uh some box around you'll probably only need to know the weights and the rough shape of the box you don't need to know um very detailed kind of representation of the the box if you want to manipulate fluids, for example, I don't think shape and weight will be enough. You need to actually model fluid as some kind of particles and how the particle will flow and will interact with different kind of surfaces. If you want to say how I can manipulate some kind of deformable objects, so say if you want to cut some vegetables, then you also need to know how the vegetables will interact with the knife.

27:28And then how they will maybe, in some cases, maybe a row on the cutting board, so you need to somehow hold it. So there's a different kind of representation for this, like if you only think about single objects. Now, if you think about border case where you have many, many different objects in the world, that's not only having different kind of properties, but also beings also in their dynamic. So the states are also changed by the forces in nature, but also other people. So for example, objects may be moved by another agent while you're not present. So now you need to also kind of imagine where the object could be.

28:14So one model will become very, very complex in the real world. Now, how we can build that, I don't think there's a single answer yet. We have world models that can work pretty well in certain domains. So I mentioned that for robots manipulation for them, but we can learn how some kind of substance or methods, objects will change, how they are stable change, given certain actions of the robots. And using that, you can try to manipulate robots for certain tasks. I think you mentioned Gaia. So Gaia is another great example of a domain-specific kind of world model. Of course, you can use that to maybe build self-driving cars, but you cannot use that for, say, object manipulation for robots.

29:07We have other cases where, for example, you can learn a world model for...

29:15like some kind of abstract kind of knowledge. So there are some world model that's represented as a knowledge graph. So say if I do this, what happens? Like say, if I want to take a flight, I need to buy a ticket first. So some kind of abstract kind of common sense knowledge. Now you can use that to do certain things. So say like maybe help you to make a plan for vacations. But again, that's because it's representing it as like an abstract knowledge, you cannot use that for critical manipulation. So I would say I don't have a great answer about like say building a single world model that can represent anything in the world.

29:58I also don't think at this moment it's the best strategy to try to build such a model. It's probably better to say a world model that can do well in a range of specific domains. And also you can easily generalize that to similar domains and fine tune it with a small set of data. But I don't think it's the best strategy yet to build one single world model that can do everything. Right. And the models that you've built are trained or tuned for what domains? Right. So specifically, we've tried to build world models for household domains. So we have built at MIT, in our group at MIT, we built a household simulation called Visual Home.

30:50So we can build different kinds of apartments. You have different kinds of objects in an apartment you can interact with. You can put many different kinds of agents, simulated agents, in those environments that, again, you can interact with. So in those simulations, we can generate embodied experiences as training data to fine-tune this, for example, language model or different type of model as a world model. So this is a very different from, say, training a language model on text and hoping you can get some knowledge about the words from those tests. Because, again, just intuitively speaking, for humans to understand the words, I don't think you can just read books and you can, for a few months, then what could happen.

31:34There are many details about the words that won't be described in text. So say, what will happen if I push this button of the dishwasher? Or what will happen if I open the cabinet door, put another object into the cabinet, and then go away, couple days later, come back, how many objects will be there in the cabinets? Those kinds of knowledge are very common to us, but you won't read it from the books. But again, those are kind of common sense knowledge that even kids can have. But the idea here is that if you build embodied agents in those simulation environments, ask them to do different kinds of tasks with different agents, or even just do random exploration, they'll get these experiences.

32:21Your guests say, so how many objects will be in the cabinet after I did this sequence of actions? What will happen after I push this button of the dishwasher? What will happen if I say, we haven't done that yet, but let's say what will happen if I talk to someone and then tell them that I'm pretty happy today. Let's hang out maybe in the afternoon. will be their response. But if you can simulate all this and gather off this experience data, then we can train a model that can really understand how the world works or even how people will react to actions. And for training in a specific domain, is that done through video or still images or text?

Read the full transcript

33:15What is the data that you're using to train, for example, a household? Technically speaking, you can use any kind of data that you can get from the environment. So we have the vision data for sure. You can have the guancho state of the objects and agents. You can translate those guancho states into language. So say how many objects are there inside these cabinets. There's some simulation that allow you to also simulate audio. So you can say, if I drop this ball onto the ground, it will be the sound that I can get. You can even simulate touch in some systems. So I think actually in the same group as MIT, they had the collaboration with another group that works on multi-model sensory.

34:07So they have this nature paper in recent years where they build this graph that can allow you to simulate, allow you to get touch sensory. So now imagine if you have some kind of VR set up where people can use this kind of graph interweights with different kind of objects and give you even the touch sensors. Now you can use those different kinds of multi-model sensory data to train your world models. We haven't gone that far yet. So the work I tell you about, we are using the conscious state of the objects and agents and then translate that into language. So we can more easily train language models to understand the world.

34:58We have also done another project we have done this. So again, we are using states that we can extract from the images. So particularly, we represent a state in the environment as a scene graph. So scene graph is the structure representation about the world. Each node is objects, each edge connecting two nodes, representing the spatial relationship between two objects. So if you have a cabinet, you have cabinet nodes. If a cup, you have a cup node. And then if the inside cabinet, there will be an inside edge connecting the cup and the cabinet. So this kind of representation can be extracted from images.

35:40But it's a pretty appealing kind of repetition compared to pixels because from pixels you don't get this semantic knowledge about the physical states. Now, you don't need to represent it into language, but you can just use the symbolic remuneration, the SynGraph, as the board state. And again, you can train that model on top of those symbolic remuneration to predict full-length what will happen next. Yeah. During that training process, I spoke to somebody a year or so ago who was working on a training system that was hyperspeed. So 100, I think it was a million frames a second or some credible rate.

36:33Do you have to train in real time or can you train at hyperspeed? How long does a system need to absorb an image or a frame of a video? So I think that depends on how data hungry your model is. And so say if you want to train model-free RL, I think you mentioned this. So usually for training model-free RL, you want your system to be as fast as possible because it is very data hungry. You have to try many, many different actions just blindly. only just blindly, so that you can get enough data or enough training signals to obtain your policy. If you haven't tried enough, you probably cannot even get any positive reward at all in some very complex task.

37:26However, I think for training world model, on the other hand, I think the speed is actually not the biggest issue. you can just that's an algorithm not algorithm even just like any kind of embodied agents in a simulation that it was for try different things in the environment for a period of time you're going to get a lot of data already because again it's different from training RL agents because RL agent has the rule function that's defined for a specific goal you try a lot of things you get a lot of observation data but none of those probably at the beginning are related to your task. So you don't get any positive reward at all.

38:09However, for training world model, you train a lot of random things in the world. A lot of random things will happen. And those random things happening is actually all the useful training signal for your world model. If I push something, something will be moved. If I push something down, something will be in a different location. If I walk from one room to another room, I know my location will be changed. and my observation also will change. So that gives you many, many training signals you can use. So you don't need to, I think, compared to at least one of the IAEA agents, you don't need to try too many steps to get useful data.

38:48So you've trained in a household environment for that domain. Is this all in simulation, or is this outside of your area of research? Have you tried controlling a robot with this kind of a language world agent model? I guess part of my research is also in human-robot interactions. So down the line, we do want to evaluate all these models in the real world. Recently, I haven't done that yet. But there are other groups that have done that. So there's, I think we actually mentioned this paper in our tutorial. So there's a very recent paper from AI2, AI Research Institute in Seattle. So they show that if you want to train the policy, a robot policy for indoor environment kind of navigation, so say try to find a TV remote for me, or try to train some policy to do simple, very simple object manipulation, say pick up one glass and give it to me.

40:02Now, and then they train it actually on raw pixels. So they don't actually train on Gwangju's world state. So give them a pixel that we can see in front of you, and then give a command, say, like, find this object for me. You want the robot to actually navigate and search through an apartment and find an object for you. Now, previously people think there could be a very big domain gap between simulation and the reward. In fact, I think many robotics researchers are actually against using simulation precisely because they think no matter how good simulation is, they're not gonna be real. They're gonna be 100 % real.

40:43And particularly if you think about the kind of policy we want to build here, it's crazy, right? it's actually low pixel inputs and they actually map it to some actions. Now the pixel rendering from simulation will be definitely not gonna be 100 % real compared to real world images. However, they show that if you have enough data, you actually do not need any kind of real world data to fine tune the model you're training in simulation. And they can deploy in real world apartments and that's the real physical robot to do the same kind of test. So I think that's very promising because the quality of the current simulation is already very good.

41:26That's one. And second, it allows to generate a lot of data, a lot of data you cannot get from the real world. And if you have enough data, you can get very good policies out of it. And I think it also presents maybe even the third kind of So when you have a lot of simulations, you have household simulations, you have traffic simulations, you have those kind of objects manipulation simulation for robots, you have fluid simulation, you have all kinds of physical, fixed engine kind of simulation. You have all these simulations people have built. You have a model that can understand, learn like general knowledge about the world from all these different simulations.

42:11Then those kind of knowledge are very useful for the real world. Even though like for individual simulation, they don't represent the whole world. But you can combine knowledge from different kinds of simulation together. Then you can have, I guess, like a forward picture about how the physical world works. This is all still in the research phase. But how long do you think before, I mean, this is one kind of agent, there are a lot of different strategies to build agents. How promising do you think this method is as opposed to others? And how long do you think before somebody refines this and puts it into a product in the real world?

43:02Right. Well, I guess like you mentioned Gaia. So I don't know exactly how they are using their simulation for their products. But, you know, they are startup companies. I don't think they're going to build anything that's just purely for research that have no real value at all. And then I think also we have things again, like I mentioned these indoor kind of robot navigation kind of tasks, right? Those are very useful, right? If you want to build robot systems that can help you to do hustle tasks, one of the most fundamental tasks you're going to ask a robot to do is to find objects you're going to need for whatever the thing you want to do in-house.

43:43Now, down the line, if you train a robot to do more complex tasks using a similar kind of approach, I think it will be a pretty kind of compelling kind of product for you. And that's definitely going to be more useful than Roomba, for sure. And I think another thing that could be very useful is not just embody agents like robots or self-interpreting cards, but also just, let's say, web agents. So think about web interface or any kind of software interface as also a role model. You can try to figure out how this interface can work after you've taken some actions, and you can model it as a kind of virtual model, so to speak.

44:27Then if you can learn that, you can also build any kind of system agents for web interfaces for softwares, etc. And so do you think that within a year or two years or five years, there will be, whether or not they're embodied or virtual, there will be, people will be working with agents built on this architecture in various domains? I am not great at prediction, but I think embodied agents are always very hard to build. Like for example, you would think one of the most basic kind of tasks for robots, which is grasping an object. I think some robotics actually previously said that whenever you start to work on grasping, you can have a whole career ahead of you.

45:23It's just a really hard problem right now. However, if you think about simplified tasks that doesn't require a very compact physical manipulation, like say maybe like web interfaces or some kind of virtual assistants that can help you to do some tasks, I think that could happen easily. I think there are already people trying to build such products as well. And I think also maybe even, let's say for the body agents, for very, very kind of, not very, but somewhat rejecting environments. So say like, you know, like in, I don't know, like warehouses or factories. Again, there are already products like that, But I think with better kind of AI models, you can have more powerful kind of robot co-workers in those more rigid environments.

46:23Yeah. But do you know of any commercial enterprise that's taking your research and implementing it? Right now, I don't think so. I don't have that knowledge. Yeah. Yeah. Yeah. It's fascinating research. So I see, do you refer to it as language agent LAW, law? Right. So language model, agent model, and world model. World model, yeah. Yeah. I mean, I see that LAW. I don't know if that's a term that you use. Okay. So what's next in your research? Well, like I said, building world models and also agent models, I think that's a very, I guess, fascinating direction for me and also for many, many people now.

47:17And hopefully, more people will find it fascinating and work on that. But I think another thing that's really interesting is social learning, which was also talked a little bit about in the tutorial. So you have model that can learn on its own but think about human human actually learning how things with each other or from other people right um and we can also teach knowledge we learned to other people so how we can actually build models that can or agents that can also learn with humans learn for human or learn from humans i think that's another kind of direction i really want to for gone. Okay.

47:55Well, we're over 45 minutes, so this is probably enough. I really, really appreciate your time. And on such short notice, it was a fascinating tutorial. And I'll watch your work going forward. Yeah, thanks so much. I had a lot of fun talking to you today. That's it for this episode. I want to thank Tianmin for his time. If you want to read a transcript of today's conversation. You can find one, as always, on our website, eye-on.ai. In the meantime, remember, the singularity may not be near, but AI is changing our world, so pay attention.

From the publisher

Join host Craig Smith on episode #174 of Eye on AI as he sits down with Tianmin Shu, Assistant Professor at Johns Hopkins University in both the Computer Science and Cognitive Science departments.

 

In this episode, Tianmin unravels the fascinating concept of world models and their intersection with artificial and human intelligence. Discover how these models, rooted in cognitive science, offer a blueprint for understanding our environment and enhancing AI's ability to interpret, predict, and interact within it.

 

Explore the intricate dance between AI and cognitive science as Tianmin shares insights from his rich academic journey from UCLA to MIT, leading to his innovative work on social aspects of AI and the development of agents capable of human-level cooperation and understanding.

 

Dive deep into the discussion on the integration of world models with large language models, and how this synergy could revolutionize AI's predictive capabilities and reasoning, paving the way for more intuitive and interactive systems across various domains, especially in household environments.

 

Don't miss this riveting exploration of the next frontier in artificial intelligence research and its potential to transform our interaction with technology.

 

Remember to rate us on Apple Podcast and Spotify if you're intrigued by the insights shared in this episode!

 

This episode is sponsored by 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach costs millions of dollars. 1Password secures every sign-in to save you time and money.

Right now, my listeners get a free 2-week trial at: 1password.com/eyeonai

 

Stay Updated:

 

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

More from Eye On A.I.

All 266 episodes
#174 Tianmin Shu: How World Models are Shaping The Future of AIEye On A.I. · 49 min
Listen in VO