In short
Podcast Summary: No Priors - Building the Factories of the Future with Covariant CEO Peter Chen
Episode Overview In this insightful episode, co-host Sarah Guo interviews Peter Chen, the co-founder and CEO of Covariant, a robotics startup aiming to revolutionize manufacturing and logistics through advanced AI. The discussion focuses on the development of adaptive AI models capable of learning and completing tasks in the physical world, the roadmaps for Covariant, and the potential future of robotics.
Key Participants
- Sarah Guo: Co-host and startup investor, founder of Conviction.
- Peter Chen: Co-founder and CEO of Covariant, former research scientist at OpenAI.
Episode Structure and Highlights
- Introduction to Peter Chen (0:00 - 0:58)
- Background in AI research at OpenAI and UC Berkeley.
- Focus on reinforcement learning, meta-learning, and unsupervised learning.
- The Future of Robotics AI (0:58 - 3:00)
- Importance of robotics in driving AI forward.
- The combination of robotics with advanced AI as a vehicle for data collection and model improvement.
- Transition from Research to Commercialization (3:00 - 5:46)
- Discussion on why Covariant chose to commercialize AI robotics despite existing limitations.
- Emphasis on building AI models that learn from large datasets generated by real-world applications.
- Incremental Development Approach (5:46 - 8:13)
- Argument for an incremental approach in developing AI capabilities.
- The importance of real-world solutions to gather actionable data for model enhancements.
- Current State of Manufacturing Robotics (8:13 - 12:21)
- Overview of conventional robotics in manufacturing, highlighting limitations.
- Discussion on the need for adaptive AI to handle diverse tasks.
- Practical Use Case: Put Wall (12:21 - 15:45)
- Explanation of the "put wall" use case in e-commerce fulfillment.
- Importance of adaptability in sorting and handling unique items efficiently.
- Covariant’s Vision and Roadmap (15:45 - 18:42)
- Future plans for the Covariant Brain (foundation model) and scaling operations.
- Discussion on the types of customers Covariant aims to serve.
- Grounding Concepts in AI (18:42 - 25:47)
- Importance of grounding abstract concepts in the physical world.
- Comparison of current AI capabilities in understanding physical interactions.
- Scaling Laws in Robotics (25:47 - 29:21)
- Overview of how scaling laws apply to Covariant’s data collection and model training.
- Discussion on the predictability of improvements through data scaling.
- Driving Thesis of Covariant (29:21 - 32:54)
- Emphasis on the belief that the future of robotics relies on extensive data collection.
- Exploration of the limitations of relying solely on simulated data.
- The ChatGPT Moment for Robotics (32:54 - 35:12)
- Discussion on what the "ChatGPT moment" would mean for robotics, emphasizing generality and reliability.
- The need for high-quality data to enable smarter robots.
- Future of Robotics and Safety Considerations (35:12 - 37:02)
- Vision for a robotics-augmented future in warehouses and manufacturing.
- Safety protocols in place for industrial robots.
Key Takeaways
- Incremental Development: A phased approach to robotics development is necessary to align with market needs and capabilities.
- Grounded Learning: Effective AI models require extensive real-world data to operate effectively in physical environments.
- Future Visions: The robotics industry is expected to evolve towards greater generality and reliability, akin to advancements seen in AI language models, with commercial applications leading the way.
Conclusion The episode concludes with a reflection on the future of robotics and AI, emphasizing the transformative potential of adaptive robots in manufacturing and logistics. Listeners are encouraged to think about the implications of AI advancements and how they will shape various industries in the years to come.
Additional Information
- Follow the Show: @NoPriorsPod
- Feedback: Email show@no-priors.com
- Subscribe: Available on major podcast platforms (Apple Podcasts, Spotify, etc.).
---
This structured summary serves as a detailed guide to the episode, encapsulating the main themes and insights discussed while facilitating deeper understanding for the audience.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Hi, listeners. Welcome to another episode of No Priors. This week, I'm joined by Peter Chen, the co-founder and CEO of Covariant, a robotics startup that is developing AI robots. Before he started Covariant, Peter was a research scientist at OpenAI and a researcher at the Berkeley AI Research Lab, where he focused on reinforcement learning, meta-learning, and unsupervised learning. He is a prolific publisher and now a founder. I'm so excited to have you on today to talk about what's going on in robotics. Welcome, Peter. Thanks, Sarah. It's great to be here. There are many exciting reasons to be here.
0:40One is I have been a frequent listener of the podcast. And the second one is just because of the name, like I just have to be on this show. So it's great to be here. Right. Let's go establish some priors for everybody in a very unknown landscape. Right. Can we start with just why you were drawn to robotics and the beginning of your research journey? Yeah. When I was working on research at both UC Berkeley as part of my PhD and at OpenAI, there were two topics that were particularly exciting to me. One topic is, as you have introduced, unsupervised learning. How can we build models that learn from vast amount of data?
1:19And we now more colloquially know this as generative AI, because we train these large models on large amount of text, images, videos, and you learn from them in an unsupervised manner. That topic has always been very interesting to me because if you want to train very capable AIs, you want to have a lot of data. And where you can get a lot of data is through this kind of unsupervised data set. And then the second topic that was really interesting to me was reinforcement learning. It's not just building models that understand, but building models that can make decisions. and reinforcement learning teach these models to make decisions by having them make trials and errors and learn from the better decisions and do less of the worst decisions.
2:06And robotics is just such a great combination of these fields. In order to build really capable robots, they need to really understand the world in a very, very robust way. And they are not just passive agents that just understand text or what's in an image. They actually need to take actions in the real world and the consequences do matter. And so we found robotics to be such a great way to both utilize the advances in AI, but also we think of it as a way to also propel AI forward. This is where you get the grounded data. This is where you get that embodied data of not just AI that is trained on browsing the internet, but AI that is trained with physical interactions with the world.
2:49And so we also believe robotics would be a key way to advance AI. That makes sense. You were at places that are great places to do research. Why did you decide to start a commercial company? It's a really good question. I mean, there are a lot of companies that are funded by prior PhDs that are kind of the classic journey of there's a technology that was built in a lab environment and it got to enough level of maturity that we should start to commercialize it in the real world. That was kind of not the journey of Covariant. When we started Covariant, there was not AI that was good enough to make robots do useful things commercially.
3:29And so it was not a classic journey of technology development in academia and then transition to a commercial landscape. The key insight that we had at that time when we left OpenAI in 2017 to start Covariant was the future of AI is going to be the future of foundation models. These models that are truly multi-task, learn from large amounts of data, and as such, be more generalizable, they can solve new tasks more easily, and are also more capable at every single one of the tasks because of the transfer that you get across tasks. We just had early conviction that that was the path to build AI.
4:08And that is also going to be true for the physical world, for robotics. But there's one big problem, which is you have no data set to build robotics foundation model. Like there's no data set that you can build this AI that understands the physical world and take actions in the physical world. And so in order to build this foundation models for robotics, you really have to build a company that can collect data to do it. And the only way to collect enough data is to build fleets of robots that are actually creating value for customers so that you can collect those data in production. Because even if you try to scale up data collection in a lab environment, there's a limit on how much you can do that.
4:53In that perspective, we strongly believe in the Tesla approach, like where they have the most self-driving car data, because they ship a great car that people want to drive and a good enough entry-level autopilot that people are willing to use it. And they're creating value for their customers, like customer use their products. And those data that they collect can allow them to build much more capable models and AI. And so why we left OpenAI and academia to start Covariant is very much this belief that in order to build foundation models for robots, you have to have a lot of data. And in order to have a lot of data, you have to build autonomously working systems for customers.
5:35And the only way to do that is to build a company to serve those customers. Yeah, there's a really interesting tension if you're trying to build a, let's say, AI capability that doesn't exist yet, because there's no model that is good enough of how much you invest in that upfront versus deliver the product that already exists in the world, right? Like you could just go build a bunch of robots and deploy them en masse. Or, you know, if we draw a analogy to the prior generation or current existing generation of autonomy companies, like we were, you know, I involved early in my prior role in Aurora and Nero, and then I was a personal investor in Kodiak, right?
6:16Like a lot of these companies you were trying to build a brain as an alternative to the Tesla approach. And I think the the economics of collect as you go is getting very, very compelling just in terms of how expensive it is to try to sequence it the other way. Yeah, like this definitely needs to be a incremental approach. Like you have to just find like the right sequence of what is the technology events that I want to build now that enable enough of a products that I can deliver, which then in turn allow you to build more capable models that then in turn like a larger service of area. And And this is like, I mean, we have seen this play out in the non-robotics world as well, right?
6:59Like if we think about OpenAI and Thorpec, Cohere, a lot of these big language models players, like the models that they have are not fully general language models yet, right? But they are good enough that can solve a large section of problems that it's worth productionizing them, getting commercial value out of it, which then in turn allow you to build the next incrementally better system. And I think of it as the same kind of road mapping exercise that you have to do in autonomy. You cannot just go straight to the full general physical AGI at the beginning. You have to build something that represents a justifiable R &B spend as well as timeline that you can justify.
7:48But that allows you to build something that is valuable that you can ship to customers. And from that process, you get more data, you get more learning that then in turn allow you to build the next generation model. So we think of it as very much an iterative approach and having real products and having real customers allow you to ground that approach as opposed to just be in a philosophical debate of like how we build this super, super general thing that is very far in the future. then i think the right way to start actually be to ground the conversation and kind of like the application landscape can you walk us through the sort of limitations of robotics in warehousing and manufacturing that are commonplace right now and how much intelligence these robots have robots are extremely common nowadays like so what we typically work on are robotic arms so think of these are six axes, seven axes, robotic arms that can do very flexible movements.
8:42They are super precise, they're super fast and super doable and very cheap. Lots of factories around the world have robots. But the challenge is like 99 plus percent of the robots that are deployed in the world are dumb robots. These robots are pre-programmed to do the same thing again and again. And they don't really have any kinds of intelligence that can adapt to new circumstances, communicate with people and change what they do on the flight. And so think of robotics that exist today are extremely rigid. And so really the problem that we are solving is we're not trying to make the existing dumb robot use cases better, right?
9:22Like we're not trying to say, oh, instead of manually programming this robot, you could just have an AI that program that robot we're not talking about that like we're really talking about like opening up a couple orders of magnitude more use cases where the robots actually need to be smart like they need to adapt what they do based on the scenario that is presented to them right so like the good way to visualize this is on one hand like think about a robot for example in a tesla factory that is handling a car body. Okay, this is a very incredible feat of engineering that can move multi-ton object very fast, very precisely, but it's just doing the same thing again and again.
10:06And then imagine another robot in an e-commerce warehouse that has hundreds of thousands of unique items that it has to distinguish, pick up, and pack carefully into a box that gets shipped to you, that's a very different kinds of diversity that we're talking about. And so when we think about building AI for robots, when we think about building foundation models for robots, we're thinking about really lifting robotics as a category from this former category of just being able to do repeated things to this category of really being able to handle diversity of environments, changes in the environments, and being able to understand what's around it and make intelligent decisions and actions to handle a diverse set of circumstances.
10:54And we think this would enable really a whole different wave of robotics that is not how robotics is used today. And for Covariant specifically, we are starting from logistics and warehouses as an industry that we focus on. So think of it as the explosive demand that is driven by the growth of e-commerce. there's a lot of complexities that's been injected into the logistics and supply chain. And at the same time, coupling that with demographics change, changing immigration landscape makes fewer and fewer people want to do this kind of warehouse jobs, like drive an hour and a half to the suburb and then have to work through the midnight.
11:39These are not the kind of jobs that people want to do. And our customers have extremely high turnover rate, like an average warehouse that we serve have typically more than 100 % year-over-year turnover rate. And so these are the type of places that we have an extreme shortage of people that want to do those kind of jobs. And yet at the same time, there are no prior robots that can solve pick, pack, ship in warehouses because traditional robots are just machines that do the motions that you're programming to do repeatedly. But here you actually need systems that's actually adaptive and do it at a very high level of reliability.
12:19Can you describe how we should imagine the physical? You obviously have Covenant Brain, but then you have the physical instantiation. What's a put wall, just for our listeners? Yeah, so a common use case that we have for our customers is what we typically call a put wall use case. A put wall is a term that is used in e-commerce fulfillment. which is when you click a button to buy something online and then a box throw up to your door and you might wonder, well, how is that done? Well, there's a complex set of operations that's happening in the background and a put wall is one step of that. And this step is typically used to sort a mix of customer orders to different customers.
13:05Let's say both you and I have ordered a new generation of iphone right and then like a robot would be sitting there and picking up one iphone and say oh this one should go to sarah and this one should go to peter if you think about like what that robot needs to do like the robot needs to have an incredibly great ability to grasp items without damaging it and have the accurate ability to identify what is the item and then route them to the appropriate customer like in this case like either you or me And so put wall, you can think of it as a sortation mechanism. You can think of it as a physical router that exists in the world.
13:42So instead of thinking about network router that sends digital packets around, you can think about put wall as a physical router that sends goods to different places. Is it fair to say that identification and routing are more solved problems than grasping? I would say identification and routing is a typically more, considered more solved problem than grasping. Like, because if you, there are other, like, more mechanical way to solve those problems. Like, you can design a piece of conveyor that, like, if you always put an item to the same place, then you can route it to a design location. And so, like, that becomes mostly a mechanical problem.
14:23And anything that is a mechanical problem is typically more solved. And so that is very much true. I would say out of this grasping identification and routing, definitely the grasping part involves more AI. But as we build more advanced AI and bring it into a more traditional field like robotics, what we actually find is that even in the identification step, even in the routing steps, there are a lot of ways that AI can make more traditional mechanical systems smarter. For example, a classic way to do identification is through scanning the barcode. But where's the barcode? How do you scan the barcode?
15:02Well, that's actually something that AI can inform it. And oftentimes, humans can identify an item without even scanning the barcode because you can read the packaging. You can infer what is in there. And that is also something that AI can help. And so while it is true that there are some steps of the problems that can be solved by more traditional mechanical and robotic systems, What we have found is that once you have a very flexible AI, you can actually rethink a lot of the processes. You would make something that was previously impossible possible, like grasping. And then you can also improve a lot of the other steps of the processes that were previously possible.
15:40But now you can do them in a more intelligent way. Is the next step of expansion that you are excited about for covariance variants still within pick and pack or are there other tasks within warehousing and logistics that you think are really interesting to expand into or you know there's other forays into different robotic applications like you know humanoid robots like the tesla optimist or other industrial applications yeah um a couple like starting at a very highest level right when we think about the covariant brain, this foundation model that we are building, we are not building it just for warehouse applications.
16:20We are not just building it for pick-and-place applications within warehouses. So definitely, everything that you're talking about, it's very exciting to us. So both applications outside of warehouses, as well as applications to newer hardware form factors like humanoid robots and so like that definitely is the long-term path for us i would say like in the very immediate future as a company we are focused in the manipulation space of warehouses just because there are so much demand and there's so many different kinds of use cases that exist in the warehouse domain already because a warehouse for a apparel company is very different from a warehouse for a cosmetics company which is very different from a warehouse for a meal prep company.
17:09And across all of these, you actually have very different manipulation skills that you need and very different kinds of data that you can collect to train the foundation model, and also very different large markets that we can tap into. But we are very intentional in how we build the models in a way that makes sure it's generalizable and so we can actually extend into new domain. And one more comment on the humanoid question. I think that would be one of the most exciting advances in robotics is to make humanoid as a form factor possible, right? Because our world is designed around human bodies.
17:46So humanoid is the universal hardware form factor that can be drop into any place in our world. And so we really cannot wait for the human noise to be commercially and also technologically available. Because when that platform is available, that is really the best mechanism for us to deploy covariant brain, this foundation models, to go to more places more quickly. Fortunately, we are not relying on it. Even by using the existing industrial robots, hardware, we can build a scaling business. We can continue to bootstrap and build incrementally more capable models. But if when it comes, like that would be a really big acceleration for us.
18:35One more question on the sort of application or maybe just the covariance side before I would love to talk a little bit more about the research is, can you give our listeners There's a sense of you're five years into Covariant. Like how big is the team? You have robots in the production. What are your types of customers? Yeah, so Covariant is about 200 people company and we are extremely international. I would say roughly half of our customers are in Europe, half of our customers in North America. And we have robots deployed across three continents at this point and more than 10 countries. And what is really remarkable, all of these customers, all of these different robots are networked together.
19:22It's one single foundation model. And everything that they learn come back and make this central model better. And our customers are typically large retailers, large e-commerce brands. And essentially anyone that runs a large distribution centers or a network of distribution centers would likely choose Covariant. as their model that power their physical world. Amazing. Can we talk a little bit just about the research? And I think the first thing I'll ask you to explain as just a very high-level concept is what the concept of grounding and understanding of the real world or foundation models that understand physics and objects interaction, like what that means or how that's missing today.
20:11Yeah. So grounding is this interesting idea of, like, if you just read the text on the internet, like, you learn a lot about abstract concepts, right? But they could be, like, purely symbolic. Like, you might read, apple is delicious. Okay, I have this association that, okay, like, something that is apple could be delicious. And if I ask for a delicious thing, you can say apple is a delicious thing. But that is very symbolic. Like, that has, like, no actual grounding in our physical world. Like, what does an apple look like if i give you an image of an apple can you recognize it uh and can you recognize like the different other physical properties of an apple uh and so like the first thing that you want to do is like grounding is to to ground all these symbolic abstract concepts into something that is real that is physical um and there are actually a lot of advances of this like even outside of robotics um that's happening already like we have a lot of multi-model model that exist in the world.
21:16If you go to GPT-4V, you're actually given an image and then it can answer something for you intelligently about what's in the image. So GPT-4V has grounded these type of multi-model language models already have an understanding of those grounded concepts. So where does it get those grounding from? It gets those grounding from essentially the image and text pairs that happen on the internet, right? Like if you look at an Instagram image, like it might have a set of captions along with it. So we can train this kind of multi-model models with a combination of those data, right? Like after you have seen enough of the Instagram image of an Apple and enough of people tag them as Apple, then after you have trained on a large amount of such data, you start to get that grounding.
22:12You start to pick up that associations. So that's like, I would say, outside of robotics, like how typically grounding happens and how you typically get this kind of multimodal understanding that understands beyond just pure symbolic concepts, but actually has an understanding of how it gets associated with the real physical world, typically manifested through an image of the real world. And if I think about just the concept of an apple is in many videos on YouTube, they are kind of round, they are affected by gravity, they have some mass. Like what's missing from those captioned images and videos when you talk about like the data that's missing that you need to go collect for robotics to improve?
22:55Yeah, so there are a couple of aspects of it. So like obviously this kind of internet scale data is very useful. Like you can already pick up a lot of association and grounding with the physical world. But there's still a lot of things that's missing, right? So for example, when you think about this kind of naturally occurring text and image pair data, they are typically about high-level concepts. They're typically not about something that is very precise. So for example, when I presented an apple to you, you don't typically describe the precise shape of the apple. Is this a very round-shaped apple?
23:32Is this a very full apple? You might use some high-level concept to describe it, but there's really nothing that describes it, say, down to sub-millimeter level precision, which is kind of like the level of precise understanding that you need to interact with the viewable. You don't just say, well, there's kind of an apple there, but there might be up to a two centimeter difference in understanding of where the boundary of that apple is and how should I do it. And so here's the first dimension of things that is missing, which is there's really no precise grounding. There's no precise understanding of the physical world that's naturally occurring on the internet.
24:13So that's one of the first things that you'll find, kind of the departure of robotics foundation models from other general multimodal foundation models. It's this idea of precision. You now actually need to understand things to a much higher level of precision that don't otherwise exist in this kind of data set. And so that's one big thing. And then another really big thing is this ability to understand effects of your own actions. And a large part of this is just because there are not a lot of robots that are doing interesting things in the world. and so like there are not a lot of data sets that are in the format of robot does something and you know the outcome of it like is this a good way to pick up something like if i move an item too quickly like would it damage it if i press like for example a tomato like what is the force that is appropriate that that is possible like you don't have a lot of these kind of um action and outcome pairs um that exist in the world like the closest thing to that is probably on the youtube you have human doing those things.
25:23But then there was a research question of like, well, can you have a robot that learns from just watching a human does it? And you don't actually fully know like how hard does a human press on a tomato or like how you precisely size something. So you're still lacking a good amount of the data that like completes this feedback loop. Do you have some sense of like how or if scaling laws apply for you? Like, do you know how many robots you need to deploy or how much data you need to go collect to get to certain levels of improvement? Or can you try to predict it now? So I would say the most technical definition of scaling law does apply.
26:00And we have seen it apply in this domain. And it's somewhat not surprising because if you think about the scaling law in the most technical sense, which is if you scale up data and you scale up your model capacity and you scale up the compute that you throw at it, you'll get lower loss function, like training loss function. out of it. And we have seen this play out across so many different domains, like more than just language model, that is not surprising. I think the question that you're asking is probably the more, not the most technical definition of scaling law, but it's the general definition of scaling law, which is, as you scale those up, would you get emerging capabilities out of it?
26:43Like, would you kind of like get something that's like modeled as orders of magnitude smarter in some loose definition of it, which is kind of the thing that we see from the large language model world, like when you go from GPT-3 to GPT-4, when you go from Cloud1 to Cloud2, you kind of see this step change, improvement in reliability, in generalization that you get from it. So I assume that's probably what you're asking. Yes, do you believe in some emergent? So I would say we see some element of it, but it is something that we rely less on. And here's where I think there was a really interesting, crucial distinction between a full general model that is designed to solve everything in the world to what I think of as a domain-specific foundation model, like in our case, like solving robotic manipulations.
27:37So in a full general model, like for example, like GPT-5 that you wanted to solve everything in the world, then you have this problem of essentially out of domain generalization. Like when we say, like, as you scale it up, like, do you get something that is much smarter out of it? Like we are not saying like whether GPT-5 would fit the training data better. Like we are saying, like, if you give a scenario that is completely outside of training data, like how well does it work? and that is where you kind of like need to rely on this strong form of scaling law. But you kind of don't need that when you are in a more restricted domain like robotics because like you actually could have so much data coverage that your test scenarios are just part of your training scenario.
28:28So to some degree, like we actually don't need to rely on this strong form of scaling law to whole. for us to build really valuable technology out of it. And so I expect something similar like that would happen, would follow the similar trend that you see in the language world. But at the same time, we don't require it. We know that as you get more customers, as you get more data, these systems would get better. And especially if you have targeted data coverage for specific domains, for specific customers, they would be guaranteed to get better. So to some degree, whether you believe robotics can scale or not, it's a simpler bet.
29:12It's just whether you can get data of that domain. And if you can get it, then you can for sure that you can fit it. Last question in this research area. Is there a specific scientific insight or bet that Covariant has made? Or should we think of this as not at all trivial, but a full stack play with the right people, very well prepared engineers and scientists doing the relevant data collection that doesn't exist today that will support increased robotic intelligence versus, let's say, like a architectural better or whatever it is? Yeah, it's like the architecture has changed like maybe five times already.
Read the full transcript
29:52It has gone through significant transformation every year. I don't think you can be married to any single specific architecture in a field that is moving so quickly. But there is one unique bet that we are placing. So that one unique bet is we believe the future of robotics would be built by whoever that has most robotics data. And essentially the whole company is built around that thesis. And you can say, what is an alternative belief? An alternative belief would be, can we just solely rely on simulation? We actually don't need much real world data. That would be a different philosophical bet on it.
30:34We also use simulation, but we think of simulation as more of a way to augment the data, not as the way to replace everything. There are lots of smart Tesla and ex-Tesla people where Tesla has been a, I guess, big proponent of high quality simulation, including for, you know, training data generation. Right. Where are the gaps or why do you believe that's that's insufficient? So when we think about simulation, it's actually somewhat different for different kinds of autonomy domain. So when you think about simulation in self-driving car, we are really mostly thinking about systems that hopefully don't physically interact with each other.
31:14Like if two cars get in contact with each other, that's a really terrible thing. And so the simulation there is more about simulation of multi-agent behaviors, like avoidance of contact. But if you think about manipulation, if you never contact something, that's also a big problem because then you actually don't do any work. And whenever you involve contact, simulation of those things become very, very difficult. Like items that can deform, like the contact dynamics is incredibly challenging. And so those are where simulation becomes very difficult. Like it's when it involves contact, complex dynamics.
31:52And then there's the second thing that makes simulation difficult is like I mentioned earlier that a typical customers that we serve like may have 100 ,000 distinct objects in a warehouse. So if you want to fully recreate that in your simulation, that is actually more work than just learning a system that can deal with the real world. So there's a vacation problem. In order to specify the real world in your simulation, that actually might require more data or more work or whatnot. And that being said, we believe in learn role model. We believe in foundation models that can learn from the real world and you can simulate new scenarios of what would happen if you do things differently.
32:36But I think of that as like different from the classical simulation that I referred to earlier, which is program-based and you are just hard coding the rules of reality and then building agents that learn from the mechanical interpretations of the rule of realities that you encode in your simulator. So for our last couple of minutes, should we zoom out and talk a little bit about the future? Yeah. So you have said we're sort of pre-Chat GPT for the robotics industry? What is the ChatGPT moment for robots? What do you imagine? The ChatGPT moment for robots, you want AI that is as general as ChatGPT.
33:12So you would be able to throw robots into any arbitrary new scenarios and you'll be able to learn how to deal with it very quickly. But in addition to that, which is kind of like what ChatGPT allow people to experience is you can ask it arbitrary problems and then they can solve to some degree. to you. So you want the same kind of generality. But in addition to that, what you also need is really high reliability because you really don't want robots that only succeed in the task that you ask it to do 70 % of the time. And then there might be 30 % really catastrophic outcomes that come with it. So I would say the bar for the chatGPT moment for robotics is higher.
33:56You need to solve the generality, which is the same kind of problem, but you need to solve it with high level of reliability. And this is where one of this concept that we talked about earlier comes in. You really need large amount of high quality data to densely cover this robotic fuse that you want. And so that would be what I think about as the model side of the chat GPT moment for robotics. And then you also need to think about the hardware portion of it, right? Like even if you have a robot AI that is very smart, unless you're just interacting with this robot AI in some metaverse digital 3D world, you still need some hardware body for robots.
34:38And before humanoids are fully widespread, I think we will see that the chatGBT of moment for robotics being articulated in the industrial settings earlier than in the commercial settings. like because those are the places that can actually justify the hardware investments because the hardware is being used 24 7 as opposed to like home robots that might only be used two hours a week like that's a very different ROI from the hardware piece that you need to put in it what does the like warehouse or factory or um logistics center of the future look like like lights out no humans I don't think it would be fully lights out and no human at least in the near future but I think of it as would be very robotics augmented.
35:26So think of one person would be able to oversee 10, 20, 30 robots. So instead of one person have to manually do all those work, you actually work with a fleet of robots. So think of it kind of as a physical co-pilot type of setup. You just get this large amplification of what one person can do. But most likely it wouldn't be completely lights out. Like you will still have people there. I think this form of expression of AI like would probably be true not just for robotics, but many other fields of AI as well. I realized you just said industrial applications first from an ROI perspective, that makes sense.
36:08But do you have a guess or a hope for what the first form or use case for intelligent robot that your average human, like your consumer interacts with? If I have to guess, it probably would be a home robot that don't involve much manipulation. So think of it as like a home robot that might be like a Roomba. It can follow you around, like you can talk to it. So like it has that navigation of movement aspects of it, but not necessarily the manipulation aspects of it, like not actually manipulating the physical world around it. I think that would be the most technologically feasible version. So think of it as similar to Amazon's Astra robot, like this kind of like cute robot that has two wheels that can follow you around and someone calls it, it can go there.
36:54And so like I think that type of form factor would probably be like when we would see it earlier. Robotics AI work, it triggers a lot of concern around safety in both like the short term practical sense and in sort of the AGI breaking into the real world sense. How do you think about safety at Covariant? We have a simple carve out to this question, because we focus on industrial applications. And well, all industrial robots have a set of safety rules that they need to conform to. Because it's not just AI can be dangerous, manual programming can be dangerous. You could program a robot to do dangerous things already.
37:35And so there's a really robot set of rules around, you have to put safety cages around robots. And if you don't have safety cages, you need to have certain kinds of certified controller that makes sure a robot doesn't do anything that's dangerous to the surrounding equipment, people. And so from that sense, because we're just following the same rules, any kinds of robots that we build and deploy are by definition safe or by construction safe. But that is very different from when you say, well, what if we hook up an arbitrarily expressive agent into a home robot that also has... How do you limit that to be safe?
38:16It's much harder. Just similar to if you just hook up a language agent to give it arbitrary Python code execution capability and arbitrary ability to access the internet. It just becomes very difficult to say, well, how can you make sure it doesn't do anything dangerous? And that's where the alignment problem comes in. And then that's where there's a lot of this good safety research comes in. But we have a simpler carve out, like at least for the near term in this kind of industrial applications. Peter, what advancement in AI research or application outside of robotics are you most personally interested in?
38:49Looking backward or looking forward? Looking forward. I can only look forward. I think the same kind of events that we have seen in last year, like we would see at least the same more order of magnitude of them in the coming year. Like it's just, if you look really behind like all these advances in large language models, image generations, they are still using relatively primitive technology. Like, so like if you, especially large language models, like they are mostly still trained just on next token prediction, like which for people that study reinforcement learning, we call it behavior cloning, which means you're just asking the AI to clone the behavior of another agent.
39:31And that is one of the most primitive way possible to train this type of systems. Because if you're just mimicking something, there's a natural ceiling on how good you can get on that. And then there was just so many other proven toolboxes that we have not deployed yet that I would say progress is guaranteed in everything that we have seen so far. And I'm super excited about that. And I'm also super excited about the open source movement continuing in the AI world, like where a lot of these advances make available to a broad set of communities that can continue to build on it and experiment with it.
40:13And so I think it will continue to be a very exciting year of AI progress. Okay, then looking backward and forward at the same time, last question is your favorite sci-fi book with robots in it, realistic or not? It's not a book, but I really like Westworld. Okay, great. Westworld, the future comes. Peter, thank you so much for joining us on No Priors. Until next time. Thanks. Find us on Twitter at NoPriorsPod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.
40:56Thank you.
From the publisher
Building adaptive AI models that can learn and complete tasks in the physical world requires precision but these AI robots could completely change manufacturing and logistics processes. Peter Chen, the co-founder and CEO of Covariant, leads the team that is building robots that will increase manufacturing efficiency, safety, and create warehouses of the future.
Today on No Priors, Peter joins Sarah to talk about how the Covariant team is developing multimodal models that have precise grounding and understanding so they can adapt to solve problems in the physical world. They also discuss how they plan their roadmap at Covariant, what could be next for the company, and what use case will bring us to the Chat-GPT moment for AI robots.
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @peterxichen
Show Notes:
(0:00) Peter Chen Background
(0:58) How robotics AI will drive AI forward
(3:00) Moving from research to a commercial company
(5:46) The argument for building incrementally
(8:13) Manufacturing robotics today
(12:21) Put wall use case
(15:45) What’s next for Covariant Brain
(18:42) Covariant’s customers
(19:50) Grounding concepts in Ai
(25:47) How scaling laws apply to Covariant
(29:21) Covariant’s driving thesis
(32:54) the Chat-GPT moment for robotics
(35:12) Manufacturing center of the future
(37:02) Safety in AI robotics




