In short
Eye On A.I. - Episode #159: Peter Chen: Building the AI-Driven Future of Robotics
Podcast Overview Host: Craig S. Smith Guest: Peter Chen, Co-Founder and CEO of Covariant Focus: The episode explores the intersection of artificial intelligence (AI) and robotics, highlighting the advancements and future potential in these fields.
Episode Highlights
Introduction
- The episode is sponsored by Oracle, promoting Oracle Cloud Infrastructure (OCI) for AI needs.
- Craig introduces Peter Chen and explains Covariant's mission in robotics.
Peter Chen's Journey in AI (03:10)
- Originating from China, Peter developed an early interest in programming and AI.
- He pursued a PhD in AI at UC Berkeley, focusing on:
- Reinforcement Learning: Training models to learn from their actions.
- Generative Models: Co-creating a course on generative AI in 2017-2018.
Philosophies and Influences (09:53)
- Peter discusses the foundational philosophies from his time at OpenAI, emphasizing:
- The importance of foundation models.
- The use of generative models on large datasets.
- Reinforcement learning as a tool for teaching AI agents.
Covariant's Foundation Model (20:48)
- Covariant aims to create a universal foundation model for robotics.
- Peter elaborates on the Three Pillars of building this model:
- Data from the Internet: Videos and images for initial training.
- Synthetic Data: Simulated data to teach AI about various scenarios.
- Robotics Data: Real-world interactions of robots to gather precise data.
Technical Insights (27:36)
- Architecture Insights: Covariant employs a mix of:
- Convolutional neural networks (CNNs).
- Attention mechanisms (similar to transformers).
- Graphical neural networks for structured tasks.
Adapting AI Models (33:20)
- Emphasizes the need for AI models to adapt to diverse hardware.
- Explains that models use configurations to instruct specific robotic tasks based on hardware capabilities.
Future of Robotics (35:01)
- Discussion on the evolution of robotics:
- Current capabilities in industrial settings.
- Vision for future use in more uncontrolled environments.
Real-World Applications (38:55)
- The potential for AI-powered robots to perform tasks in uncontrolled environments.
- The importance of continuous training from operational robots.
Warehousing and Automation (42:11)
- Vision for future warehouses:
- Automation and robotics enhancing productivity.
- Human roles evolving from repetitive tasks to supervisory roles.
Challenges and Innovations (45:51)
- Hardware limitations vs. AI advancements.
- The potential for step changes in robotics as compute and data scales increase.
Key Takeaways
- Foundation Models: The concept of using a single robust model to power diverse robotic applications.
- Continuous Learning: Emphasizing the need for robots to learn from real-world interactions to improve their capabilities.
- Market Potential: As robotics technology advances, the integration of AI into various industries will likely expand significantly.
Conclusion Peter Chen's insights reveal a promising future for AI and robotics, with the potential for transformative changes in how industries operate. The episode emphasizes the importance of understanding both AI and real-world physics to develop effective robotic systems.
---
Additional Resources
- Follow Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
- Follow Eye on A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
- Learn more about OCI: [Oracle's Test Drive](https://oracle.com/eyeonai)
*For a complete transcript of this episode, visit [eye-on.ai](https://eye-on.ai).*
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Learning a role model of other humans, other human drivers, pedestrians, cyclists, behavior gives them a much better ability to build a driving AI. So absolutely makes sense. Like the same idea of understanding the environment through a lot of data so that you can better anticipate like what actions you should take also makes sense in our robotics world, in particular robotic manipulation worlds. In robotics, inference, cost, and speed are very important. In robotics, like you need your robots to be acting all the time. That means you have to really optimize your AI models to give output very quickly and so that robots can take actions continuously.
0:41AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs. OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course nobody does data better than Oracle.
1:33So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash ionai. That's E-Y-E-O-N-A-I, all run together. Oracle.com slash IonAI. That's Oracle.com slash IonAI. Hi, I'm Craig Smith, and this is IonAI. In today's episode, we delve deep into the world of AI and robotics. with Peter Chen, co-founder and CEO of Covariant, an industrial AI robotics company. Peter talks about building a universal foundation model that can operate different kinds of industrial robots across three continents.
2:40We also discuss the role of world models in predicting the outcomes of actions and their potential to generalize across various robotic applications. Whether it's navigating the complexities of warehouse automation or envisioning the future of robotics, Peter's insights are as transformative as the technology he's helping to create. I hope you find the conversation as engrossing as I did. Craig, it's great to be here. Thank you so much for having me. My name is Peterson. I'm one of the co-founders as well as CEO of Covariant, a company that is focused on building foundation model for robotics.
3:19A bit of a personal history about myself, I was born and raised in China. I got into programming and computer science in a very young age. And it always fascinates me like how much intelligence you can get by programming software, like which is basically instructing a computer to act intelligently. But there's always something left. What if the computers can learn from data as opposed to you have to instruct every single bit of the intelligence rule? So that's what got me into PhD after my undergrad at UC Berkeley. I started my PhD in AI with Professor Peter Bill at UC Berkeley, focusing on two areas that have both become extremely hot today.
4:09One area is reinforcement learning, like the art of having machine learning models produce actions. And some of these actions lead to good consequences. Some of these actions lead to bad consequences. How can you have models learn from its own mistakes and successes? So that's reinforcement learning. Another area of my PhD focus was generative models. So my PhD advisor and I actually co-created and co-taught the first graduate level generative AI class at Berkeley back in 2017, 2018, a long time ago. And obviously, we have seen how those fields have really taken off in the last couple of years.
4:55and those ideas that were once very academic, cutting edge and unproven have become a lot more commonplace concept that people interact with, right? Like if you go to ChatGPT, like it is both a generative model and a model that is aligned by reinforcement learning. So it's amazing to see that arc of transformation of something that was cutting edge, unproven ideas to something that now is everyday and is still continuing to accelerate the AI's development. Another bit of my personal background was I also spent time early on at OpenAI. I joined OpenAI fairly early when I started working out at OpenAI.
5:36OpenAI didn't even have an office. We were working out of Greg Progman's apartment at the time. And obviously, OpenAI has done amazingly well and have really powered a lot of the AI revolution that we are seeing today. but maybe the thing that I want to call out that's super interesting to me especially as I reflect back is some of the core philosophies that OpenAI started out with are really the same set of philosophies that are powering the success of OpenAI that we are seeing today. So if I were to summarize the early research philosophies at OpenAI in its first one or two years of existence as a company like it was this belief of a foundation model scaling up big model on large diverse data sets this belief in generative model using genetic model to absorb a lot of unlabeled data unstructured data and third is reinforcement learning like this ability to teach agents or models the ability to take actions in the world those philosophies like heavily influenced me and heavily influenced what we do at CoVariant.
6:48And also it's the same driving forces of what Power OpenAI is assessed today. Fast forward to how we founded CoVariant. So a couple of the co-founders at CoVariant left OpenAI in late 2017 to start CoVariant. We started CoVariant really with much of the same thesis of what powers the current success of large language model. We believe in single model. We believe in single large model that is a foundation model, which means it's a model that is trained on multiple types of tasks. It can leverage the transfer learnings across multiple kinds of tasks and have emergent behavior. And so it generalizes to new tasks better, but it also performs better at any specific task than a bespoke model that is only trained on that task.
7:42And we had an incredibly strong conviction that this foundation model for robotics has to be the way to go to solve robotics problems. Because, like, obviously we have seen the success of this foundation model approach for language. But the reason that it makes even more sense for robotics is that there's only one physical world. unlike in the language where you're trying to compress the whole world of human knowledge which includes many things that have nothing in common with each other like what is the soil composition on the moon versus how do you play chess okay like both of these are knowledge that you can find on the internet but they have absolutely nothing in common with each other and you're trying to compress all these things into one model but if you think about building a foundation model for robotics like the robot could have different bodies, and the robots might be doing different things.
8:39Like it might be interacting in different kinds of environments. However, like all of these robots live and operate in the same physical world. And so it makes a lot of sense to be one foundation model that can learn from all of these different robot experience and really understand physics and understand how do you control robots to move in the world around us. So that was a little bit of the founding story of CoVarian. And fast forward to today, CoVarian has commercialized the first robotic foundation model that's ever built. So essentially one single model that is powering robots, working in production in customer environments in three different continents, dozens of different robot hardware bodies that same single foundation model is powering and solving problems in a lot of different industries.
9:29We're starting out from warehouses, but our long-term goal is to build this foundation model that can solve robotic manipulation problems in general, like across multiple other industries. So that's a quick introduction about myself, where I come from, like the philosophies and the technical ideas that have heavily influenced us and where we have taken it so far. Yeah. I have a couple of questions. When you were doing the first course on generative AI at Berkeley, was that in response to the development of the transformer algorithm? Or did you integrate the transformer algorithm into that course?
10:13Because that has really accelerated and is today the core of generative AI. All right. Yeah, I would say Transformer was not a key focus of that course. So when we think about generative AI, there are, you can think about it as the two major components of it. So one major component is, what is the model architecture? So that is the, how is your brain structure? So obviously, a better brain structure can allow you to learn better. right and so transformer is a really flexible brain structure that can allow you to absorb a lot of knowledge patterns from the data there's another side of how do you train a generative model that is the like how do you teach it right like so even if you have a very flexible brain structure like how do you how do you actually give that that learning like think of it as the the curriculum teaching methods like so in this case like it would be the statistical models that you use, different versions of it would be the most popular one now, obviously, is diffusion, autoregressive model of next token prediction.
11:26And then a little bit earlier ago, there will be GANs, VAEs, and these different models that are the statistical representations that you impose on the world, which you can see as methods of how you teach the models. So I would say the initial class focus more on like what are the different statistical models that you could use as much like as opposed to a more focus on like what is the underlying brain structure like but you can like swap and and and plug in with these kind of things like so for example the idea of diffusion like you can implement diffusion um with both a convolutional neural nets you could implement it with a transformer based architecture and like obviously the currently the more popular one like stable diffusion is implemented with a combination of convolutional neural nets as well as transformers.
12:17You can take really the best of both worlds. Yeah, the current, I mean, there's been a lot of talk, at least in the literature, in the last few months about using large models, large language models or pre-trained transformer models as agents. And then more recently, you and I spoke about this the other day, Jan LeCun and Alex Kendall at a company called Wave AI are working with world models that learn causality directly from sensory inputs, not through the filter of language. And that world model idea speaks to me because it seems much closer to how the brain learns initially. In the covariant foundational model, can you talk about the architecture and how it works?
13:42I know that it's certainly proprietary, but, and then talk about these new developments with large language models and world models and whether you're integrating those ideas or whether you see promise in them. Yeah. So those are really good questions. Maybe we would tackle them separately, like kind of the idea of role model and then also the idea of agents in the language world. So first of all, this idea of learning a role model for any kind of what we call embodied agents makes a lot of sense. So if the goal of an agent is to understand the physical world and take actions in it, and your actions have consequences, then you should have understanding of what that consequences are.
14:40As opposed to just blindly trying things and say, oh yeah, pushing button A is better than pushing button B. I mean, that kind of works if you have a lot of data, but that's kind of like a very naive understanding of the world. Like if I push a lot of button A, it tends to give me better outcome than pushing button B. Okay, then I do pushing button A more. That's kind of like not a very sophisticated understanding of the world. And it's also not very generalizable. Like what if I, like instead of presenting button A and B to you, I give you a keyboard and you need to type in like pass key and it does different things.
15:12You really cannot take any of the learnings that you get from, oh, pushing button A is better. Like if now you suddenly have a different way of interacting in the world. And so the general idea of building a role model is, can we build agents that really understand the environment and have the ability to anticipate what are the consequences of the actions that it takes? And this idea has many different incarnations. With what Alex and the team is doing at Wave, a lot of it is about anticipating other agents' behavior. If you drive a car, slowly edging it to a pedestrian, most often people would try to step away from the car if they notice the car is approaching them.
16:05And if they don't step away, that means maybe they didn't notice that the car is approaching them. So there's some interesting interaction that you can learn by anticipating other agents' behavior in that physical world, which is absolutely the most core problem to solve in self-driving. It's like this kind of multi-agent interaction and the behavior that you need to generate from there. And so by learning a world model of other humans, other human drivers, pedestrians, cyclists' behavior gives them a much better ability to build a driving AI. So absolutely makes sense. And the same idea of understanding the environment through a lot of data so that you can better anticipate what actions you should take also makes sense in our robotics world, in particular robotic manipulation world.
16:54So in our case, the world that you're learning is not what's in another human being's mind, but it's more like what happens to physics. If I pick things up in different ways, what is more stable? What is less stable? If I throw things away in a certain manner, where would it land if I want to carefully tuck things together? And so, again, you have two ways that you can build this AI. Like one way they can build this AI is I just do a lot of blind random trial and I see what happens to work and I just keep doing that. Or you could actually learn a sophisticated understanding of, okay, like if I pick things up this way, this is a pretty stable way of grasping a certain item.
17:39If I pick things up another way, oh, wow, this is like very unstable, very precarious way of picking up an item. but I can, by anticipating what would happen in the physical world, it gives the AI a much better ability to act in it. So we are a strong believer in this idea of a role model, like essentially this AI that can learn about the environment and also anticipate what would happen to it. And then another thing that it's not just a more, like you said it's not just a more sensible way for an intelligent entity to learn because that's like more of like how human would learn you don't just randomly try you anticipate like how things what would be the consequence of the action that you take but in addition to that like there was an really amazing property about role model that we believe is under talked about and which is this idea that if you formulate the right role model, that unified all robotic applications.
18:43So you could have, for example, if you think about a robot that is folding laundry versus a robot that is packing a customer order in a warehouse. At a service level, there's kind of nothing in common with these two robots. like one robot is trying to carefully think about, OK, like how do I pick up a deformable piece of a T-shirt? And how do I flatten it? And how do I fold it? Another robot in a warehouse would be thinking about, well, I need to pick up the item. I need to find where the barcode is. And I need to scan the barcode. From a policy or from an action perspective, there's really nothing in common between these two robots.
19:25But what is in common about them is there's only one physical world that's powering them. right? So if you're learning a role model that understands, if I interact with the world in a certain way, what would happen in the next few seconds physically? That concept is universal. Right? And so what makes this role model idea so powerful is it gives you this formulation or quite like an interface that is the same across all robots, Like no matter what kind of tasks they are doing, no matter what kinds of hardware they are using, it's the same. Like it's the same role model because there's only one physical world.
20:05Like now suddenly you find a way to really scale up the data that can go into training a robotic foundation model. Like for all of these foundation models, like one of the, like obviously you need the right model, you need the right algorithm, but you also need a lot of data of the right kind. So the role model actually opens up the possibility of training large foundation model that's learning on a lot of data, because now you can pull the data from many different robots together. And it doesn't matter what kind of tasks they are doing and the environment that they're interacting. There's the same set of physical principles behind them.
20:43And that's what the model can learn. And the initial training, that's sort of the ongoing training, but the initial training, you can simply train from video. Is that right? That shows the laws of physics or causality and that sort of thing. And then there's ongoing training as the network computers learn. Yeah, so we believe video is a super important format for this. Like how can you, like there's just so much data that's encoded in video. But what we have found is that pure video is also not sufficient. Like for example, like if you just go to YouTube and just watch a whole bunch of videos, you only get a very partial view of the world.
21:42So there are, let me just point out two things that are missing. What if I just go to YouTube and watch a lot of video? The first thing that is missing is, in a lot of the cases, you don't really know what are the actions that are taken, because you're just passively observing things. And so you don't really know what are the actions that are taken. And then the second thing is, in a lot of these videos out there, you also don't have a very detail... It also lacks the very small details that are really important to robotics. So for example, what is the actual velocity of a certain thing down to a very fine degree of precision?
22:28That information is kind of hard to infer from a video. But if you are controlling a full robot system, you actually can get it maybe from the motor encoder, but you can get the information in a lot more precise way. In a lot of robotics cases, you do need a pretty precise understanding of the world that video doesn't fully communicate. So both the lack of understanding of what are the actions that were taken in videos, as well as the position that is required, makes what we call videos in the wild a useful source of data, but it's definitely not a sufficient set of data. And then where do you get, is that why that data is supplemented with data from robots operating in different scenarios?
23:25Or are there other kinds of data? Can you use synthetic data, for example? Yeah. So let me maybe answer the broader question. So when we think about how do you build a robotic foundation model, a truly universal AI that can be powering any robot hardware to do any arbitrary things to a very high level of autonomy. We believe the data recipe for that are three pillars. The first pillar is what we talk about, essentially data on the internet. So video data, image data on the internet. Second thing about it is synthetic data, like generated data that are not, they may not look exactly like the real world, but they contain useful structure about the world that can teach the AI.
24:16And you can get lots of interesting combinations of known factors or variations through simulation through this kind of synthetic data. Like we believe that's very important. But these two things, like these kind of data in the wild on the internet and synthetic data are both very useful learning sources. But in our experience, they are not sufficient. They still lack the kind of actual detailed interaction with the world, like understanding cause and effects, understanding them to a very high degree of fidelity. Like those are not present in these two data sets that we talk about. Like to give an example of like where the synthetic data breakdown, it's very difficult to simulate contact and anything that deforms.
24:59And so those are the kind of places that like, OK, I can use simulated data to simulate something that's rigid or maybe doesn't involve a lot of contact. But as soon as you involve that, like your simulation's quality or the precision of it would quickly decrease. So in the end of the day, what we have found is that the third bucket of the data that you need is robots interacting with objects in the real world at scale. So these are like the three data buckets that go into training a robotic foundation model. So data on the internet, in the wild, synthetic data, and then large volume of robotics data interacting with the real world in production.
25:45And this is really the core focus of Covariant ever since we started the company. We strongly believe in the idea that in order to build the best robot AI, you need to have the most amount of robotics data that are the highest quality, which is why we really focus on, one, solving customer problems, like making sure we build a technology that is not just interesting lab demo, but it's something that actually works reliably 24-7 in an industrial environment. And the robots are so reliable, so autonomous, and deliver at such a level of throughput that our customers' facilities just completely depend on them.
26:28And once you have that, you have robots out there that are generating data at an incredible rate, because we are deploying these type of robots into industrial warehouse facilities that process amazing volume. So once you actually make these systems generate commercial value, you can collect tremendous amount of data while you generate value for your customers. So that's being very customer value focused, being very production deployment focused. It's like one thing that we have really focused on as a company. The second thing that we really focus on as a company is collecting the right kind of data.
27:09So it's not just about, oh, get the robots out there. And then as they get used 24-7, a lot of data gets generated. We also spend a lot of time thinking about what is exactly the right kind of data that you need to collect from the fleet of robots out there so that you can actually enable learning. And there are lots of deep research thinking and iterations that go into it. And it's still something that we are obviously very actively iterating on. Yeah. And the architecture of your world, of your foundation model, is it like a JEPA architecture that Jan LeCun talks about where the model is encoding the data into a higher representation space?
28:04then operating in that space to make predictions or, or, or is it, uh, more along the lines of, um, you know, a, a generative pre-trained transformer model where everything's being tokenized and you're, you're predicting the next in a series. well I guess that wouldn't be a world model but do you combine these things? Yeah there are definitely different schools of thoughts on how do you represent this type of thing like one way to do it is you can say there's some explicit latent representation of the world that I learned and it's in that some kind of latent description of the world that I learned to predict.
28:58I learned causality. There's another version of the world, which is if you think about a large language model that is just transformer looking at all the previous words and I predict the next word, there's never an explicit representation of a latent structure somewhere. There's not saying, oh, I encode all of my previous words into some latent space and I decode the next word from it. but instead you just have this large structure that looks at everything that you have seen and then auto-aggressively predict the next one. I wouldn't say... I think the jury is still out in terms of which one will be a more likely successful structure and we can probably draw some biological inspiration or maybe draw some inspiration from how we work like and you say well like it seems more efficient to operate in some kind of latent space like because like then you're operating in this um your reasoning about the world in a more abstract way as opposed to well looking at every single pixels and then trying to think about like what should the next pixel be um i would say like both are avenues that we look at um but we don't believe if there was like one clear winner at this moment.
30:21But you have a foundation model in operation. So what's the architecture? Can you describe how that foundation model works, whether it's transformer based? Yeah, just.
30:43So I think there are really two questions in there. One is what does role model look like, and what are the other parts of foundation models and the specific architectures that are used there? So the specific role model that we have, I would say it's more similar to the latent type of representation. Like there's a more compact, more abstract representation of the world that's operating on.
31:13and we believe like that likely would continue to be the case like just because like operating in pure pixels of images and videos it's a pretty wasteful representation to look at like when you think about oh if I don't grab something in a stable way and then it drops like in your head you don't try to predict where each pixel would go right like you have this high level notion of oh something would drop from it So we think likely that would be more successful, but we are pretty open-minded about it. In terms of the second question of, what are the specific model architectures that are used in our foundation models?
31:53It's a pretty wide set of architectures. It's not a pure transformer-based architecture, so it's not all attention blocks throughout. like it's a combination of it's a combination of convolution attention and in some places more structured type of attention like graphical neural nets like for specific places that make sense so i would say like the key insight there is in robotics inference cost and speed are very important, right? Because unlike in a, maybe a different world, like where it's okay for my like next sentence prediction to be a little bit slow in robotics, like you need your robots to be acting all the time.
32:44Like that means like you have to really optimize your AI models to give output very quickly. And so the robots can take actions continuously. And because of that, that we spend a lot of work on not just blindly following the biggest, most expressive architecture, but really trying to use the domain understandings that we have about robotics and really optimize the architecture to be more latency sensitive, like to be more compute budget sensitive.
Read the full transcript
33:19And you have to adapt this approach model to whatever hardware it's running on. And we talked before about the hardware constraints in robotics.
33:36Does the world model, your foundation model, does it for each specific robot that it's controlling, the robot has a goal or a policy? and does the model have to take into account the hardware configuration? Certainly. Yeah. So think about in the large language model world, it's very common to use system prompts to configure the character, the tone, or the styles of a language agent. In our world, you would use, think of it as the equivalence of prompting to basically instruct a robotic foundation model to know, well, what kind of hardware am I using? And what are the things that I can do with my current hardware body?
34:42So think of it as one base model. But then on top of that, you add configurations or prompting that actually instruct the model on, okay, you're now using this kind of hardware body, and this is what you should do with it.
35:01You've been working with this foundation model as a world model since the founding of Covariant. Has the research moved significantly from when you started? Excuse me. From when you started? Oh, certainly. For example, yeah, I'm reading a lot now about LLM-based agents and all of this world model. I've followed Ian Lacoon's research, and it's really progressed a lot. Certainly. Like there are many things that are accelerating at a very fast rate. One core thing is compute. Like when we started six years ago, like the amount of compute that you could have access to is very different from the amount of compute that you can have access to today, right?
35:58So as compute's availability goes up, you can train bigger models that have more expressivity. And then the second thing is data scale, right? I mean, when we started as a company, we have no production customers. Now we have robots running autonomously 24-7 on three different continents. So the data that we can generate is just at a completely different scale. And then the last one is the field has also moved a lot very quickly. So if you think about a lot of the really scaled up transformer architectures, How do you do large scale training? How do you do image generation by diffusion? Like a lot of things have happened in the last couple of years that also enable us to build more sophisticated foundation models and world models, right?
36:49Because I think the key thing to recognize is that a lot of these ideas that we are talking about, like whether it's foundation model or world models, they have many different levels of potential expressivity, right? So for example, the most rudimentary form of role model might only be able to allow you to predict whether I have successfully grasped on a certain item or not. That is also a role model. But it's just a role model that's restrictive to understanding whether I have successfully grasped an item. A much more expressive form of role model could be, well, if I have a cylindrical object in front of me, and if I push it a little bit, it would roll around.
37:33and if it's like on a slanted surface in my rollback, like that is a much more sophisticated kind of role model compared to like a role model that is only Taylor specific to one like small use case, right? So I would say like because of the three forces that have happened, like the more compute that is available, the more data that is being generated by our production fleet and the AI fields advances have together allow us to build progressively more powerful foundation models and part of that would be the role model. And these progressively more powerful models can allow existing applications to perform better, but also open up the possibilities for newer type of applications.
38:17So I would say like in the world of building foundation models for robotics, we are seeing a very similar trend to what we are seeing in the large language model worlds. Like you look at the difference between GPT-2 to GPT-3 to GPT-4, there are remarkable differences, right? Like as you scale up compute data techniques, like you get greater capabilities out of it. Like, even though like you can argue they are all the same idea of transformer plus next token prediction. But as you like, as you do those things, like you get qualitatively different results and it enable all those of many to more applications.
38:54Yeah. The covariance robots, even though they're different form factors, operating controlled environments, relatively controlled environments, and the degree of randomness or unpredictability is within a fairly narrow margin. how long do you think before a world model controlling a robot or a foundation model, it wouldn't only be a world model, can really operate in a real world environment, unstructured, uncontrolled? It's a really good question. I wouldn't call the current environment fully structured, like the robots, like where the robots are operating in. If you think about all the objects that we see and manipulate with in our day-to-day life, they go through a warehouse at some point.
39:58So if the covariant brain-powered robots are operating in all warehouses, it's building a pretty sophisticated understanding of how you... manipulate things, like even in the fully unstructured world. But I think maybe the question is getting more at when can we have robots that move around and are like kind of fully in the wild as opposed to in like confined space in industrial processes that are high volume. I think that mostly is going to become a hardware question as opposed to an AI question. I actually believe hardware is going to be the long pole in the tent as opposed to whether you can build the AI layer that can navigate in a less structural world freely.
40:57So there are a lot of companies and efforts working on this, like humanoid robots, like maybe humanoid with wheels, maybe humanoid with legs, or maybe different kinds of form factors that actually allow you to navigate freely in the space and also do useful things with it. I think there were a lot of interesting hardware problems to be solved there. Yeah. And on the warehouse, well, for example, on the hardware side, we were talking about Wave AI. The car is a robot and it's well refined. So how long do you think before world models can be applied to autonomous vehicles with such success and stability that we can use them on the roads today?
42:00Or is that really a regulatory issue? I think this question is much better answered by like the self-driving car experts. Okay, well, let me ask a warehouse question. And I know we're coming up to the end of the hour. But I remember when I first met you guys and we were talking about the advantage of automated warehouses. And one thing that kind of fascinated me is, you know, as you said, they operate 24-7. They can operate in low light environments. They don't need, you know, the air quality that humans need. I mean, do you have a vision for, and as you said, almost every object that we touch has come through a warehouse.
42:50Do you have a vision for warehouses of the future that maybe will be these vast underground spaces where robots are working tirelessly in the dark and a human only has to go down periodically when there's some glitch or some hardware problem? Yeah, I mean, I think it's more going to be a continuum. It's not going to be fully light. What is the extreme of that? The extreme of that is you can think about the whole moon is colonized by robots and there's no human being on it. But there's a space factory there that just keep pumping out goods and they get sent to the Earth autonomously. I mean, that's one end of the extreme.
43:44It's like truly, truly lights out like no human touch point. I think that's pretty far away. I think ultimately we will get there. But what we believe is a more gradual adoption. And you can already start adopting this type of technology today. And for a long time, this form of technology would be adopted in the form of human augmentation. As opposed to, I have to stand there and pump out 500 units of goods a day. I can now supervise 10 robots that are producing 5 ,000, 6 ,000, 7 ,000 units of goods today. And I have a much more engaging job of, well, I'm actually overseeing 10 robots and looking at where they get stuck.
44:32Like, how can I arrange the upstream goods coming in, organize it in a way that allow me to get more output out of this fleet of robots and figuring out how can I unblock the robots as quickly as possible, as opposed to just doing the same motion again and again, like eight hours of a day. And this form of augmentation we see as a way to solve the labor challenges that our customers have. So it allows them to do much more with a much smaller pool of labor, as well as have a much better human experience for people that are working in this vitally important societal infrastructure. Like now you actually have a job that is a lot more fun, engaging, and also way more productive than before.
45:24and over time like as the technology becomes better like you can see that ratio like ships improving like maybe like now it's like one person overseeing a fleet of 10 robots in the future it would be 50 100 like at some point like you would have like whole factual robots and maybe just like one person walking around in it and maybe at some point in the future like we'll get to that oh like there's like two billion space factories on the moon and like no one needs to touch that yeah uh okay well let me ask one more question because when we first uh were talking uh during the pandemic uh there was a lot i i don't know whether it was you or or one of your partners i had asked uh whether uh robots could operate uh a chicken uh processing factory i mean certainly there are some there is some automation there but there was so much uh death from covid in these uh on these production lines uh but the the hardware wasn't versatile enough at that point to be able to deal with something as as soft and and uh floppy as a chicken carcass is that are we on the road to that and then my final final question is do you feel you talked about a continuum are are we close to a step change in robotics or are we still on uh just a pretty steady incline So I don't actually have context on the chicken processing question, so I would maybe not answer that specifically.
47:17But on the idea of, are we making progress on manipulating more dexterous objects, like manipulating deformable objects in a more dexterous way? The answer is definitely yes. So when we look at our systems and the capabilities that they have, it's definitely understanding the physical world in more and more nuanced way and in more and more expressive way. And so we're all on a way there. And then in terms of the questions of whether we see a step change in robotics,
47:52I would say the most honest answer is we don't fully know. There are some forces that are working for it, and then there are some forces that are working against it. So let me tell you the forces that are working for or the huge inflection acceleration. Like the forces they're working for is we are getting to the point that you can really scale up compute and data to train really large robotic foundation models. And we have seen this kind of step change, like this phase transition of capabilities in the language world as you scale to a certain size of model, compute, and data. And we expect the same thing to happen.
48:32We expect the foundation models that power these robots to get significantly smarter, to get significantly more general. I mean, I cannot comment on what would happen outside of covariant, but at least at covariant, we are seeing that coming. How do you scale up data, compute model size significantly to get a much smarter model? So the forces that are working against it is the adoption still needs to go through hardware and it still needs to go through an enterprise adoption process. This is not like ChatGPT that you can largely turn on the faucet and maybe not turn on the faucet, but go to your website and then you can use it.
49:15So the adoption process would be slower, but I would say in terms of the capability leapfrog, Like we think we are very close to a very clear phase transition. AI might be the most important new computer technology ever. It's storming every industry and literally billions of dollars are being invested. So buckle up. The problem is that AI needs a lot of speed and processing power. So how do you compete without costs spiraling out of control? It's time to upgrade to the next generation of the cloud, Oracle Cloud Infrastructure, or OCI. OCI is a single platform for your infrastructure, database, application development, and AI needs.
50:04OCI has four to eight times the bandwidth of other clouds, offers one consistent price instead of variable regional pricing, and of course nobody does data better than Oracle. So now you can train your AI models at twice the speed and less than half the cost of other clouds. If you want to do more and spend less, like Uber, 8x8, and Databricks Mosaic, take a free test drive of OCI at oracle.com slash ionai. That's E-Y-E-O-N-A-I, all run together. Oracle.com slash IonAI. That's Oracle.com slash IonAI. That's it for this episode. I want to thank Peter for his time. If you want to read a transcript of the conversation today, You can find one on our website, eye-on.ai.
51:12And remember, the singularity may not be near, but AI is already changing your world. So pay attention.
From the publisher
This episode is sponsored by Oracle. AI is revolutionizing industries, but needs power without breaking the bank. Enter Oracle Cloud Infrastructure (OCI): the one-stop platform for all your AI needs, with 4-8x the bandwidth of other clouds. Train AI models faster and at half the cost. Be ahead like Uber and Cohere.
If you want to do more and spend less like Uber, 8x8, and Databricks Mosaic - take a free test drive of OCI at https://oracle.com/eyeonai
In episode #159 of Eye on AI, Craig Smith sits down with Peter Chen, the co-founder and CEO of Covariant, in a deep dive into the world of AI-driven robotics.
Peter shares his journey from his early days in China to his pivotal role in shaping the future of AI at Covariant. He discusses the philosophies that guided his work at OpenAI and how these have influenced Covariant's mission in robotics.
This episode unveils how Covariant is harnessing AI to build foundational models for robotics, discussing the intersection of reinforcement learning, generative models, and the broader implications for the field. Peter elaborates on the challenges and breakthroughs in developing AI agents that can operate in dynamic, real-world environments, providing insights into the future of robotics and AI integration.
Join us for this insightful conversation, where Peter Chen maps out the evolving landscape of AI in robotics, shedding light on how Covariant is pushing the boundaries of what's possible.
Stay updated:
Craig Smith's Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview and Introduction
(03:10) Peter Chen Journey in AI
(09:53) The Evolution of Generative AI and Transformer Models
(12:21) The Concept of World Models in AI
(14:03) Building Robust Role Models in AI
(20:48) Training AI: From Video Analysis to Real-World Interaction
(23:10) The Three Pillars of Building a Robotic Foundation Model
(27:36) Architectural Insights of Covariant's Foundation Model
(33:20) Adapting AI Models to Diverse Hardware
(35:01) The Future of Robotics: Progress and Potential
(38:55) Real-World Application and Future of AI-Controlled Robots
(42:11) Envisioning the Future of Automated Warehouses
(45:51) The Evolution of Robotics: Current Trends and Future Prospects




