Ep 161: Generalist CEO Pete Florence on the Robot Revolution & Bringing General Intelligence to the Physical World

24 Aug 2026 · 37 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Generalist CEO Pete Florence argues robotics is entering a “GPT-3 era,” where foundation models trained on massive physical-interaction data can be rapidly adapted to new tasks, enabling general intelligence in the physical world. He claims robots can learn from scalable egocentric data collection (robots shouldn’t sit idle), reach “competence” in minutes and “mastery” with very high reliability, and unlock faster scientific experimentation and reindustrialization.

Guest backgrounds

Pete Florence is co-founder/CEO of Generalist; PhD at MIT; previously a senior scientist at Google DeepMind on robotics and large-scale multimodal learning. Interviewers are Joe and Vivek (not further identified in the transcript).

Key claims

scaling laws exist in robotics; Generalist Gen 0/Gen 1 show reliability (99%+), speed, and improvisational intelligence; emergent ambidexterity and tool generalization appear after training; data creation/training/deployment loops are the competitive moat.

Notable examples

MIT grad-school GoPro “robot hands” exercise making an iced latte; brushing a cube into a bowl; ambidexterity from training only with the right hand.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Robot Revolution Begins

0:00 to 1:18

Exploration of the rapid advancements in robotics and AI applications.

“It feels like we are in GPT-3 era for robotics models.”

Pete's Journey in Robotics

1:29 to 2:22

Pete Florence shares his path from MIT to DeepMind and his experiences.

“And then, yeah, you know, long story short, eventually made my way to do my PhD at MIT.”

Epiphany on Robotics Dexterity

2:22 to 4:33

Pete discusses his realizations about data-driven learning in robotics.

“Um, I think, I think there was many different points, Joe.”

The Role of Large-Scale Learning

4:33 to 5:32

Exploration of large-scale multimodal learning in robotics.

“the rest of like, how do we train the models?”

Defining Embodied Intelligence

5:32 to 7:24

Discussion on what embodied intelligence means in the context of robotics.

“But any type of intelligence that actually, you know, isn't just existing purely sort of digitally, but actually has some type of a physical embodiment in the world, right?”

The Tongs Experiment

7:24 to 10:46

Pete recounts a personal experiment to illustrate robotic manipulation.

“you know, that model would probably be capable of.”

Intelligence vs. Mechanical Complexity

10:46 to 12:56

Examining the balance between intelligence and mechanical design in robots.

“Like we were talking about this problem where you walk around most robot labs and everybody's robots are just sitting still.”

The Current Robotics Wave

12:56 to 14:00

Understanding the current state and potential of AI in robotics.

“That is a very nicely way to say it, Vivek, yes.”

Advancements in Robotics and General Intelligence

14:00 to 14:44

Explore the rapid advancements in robotics and their commercial viability.

“Like now we're, you know, in the world of, you know, all these coding models and all these coding agents that are just like rapidly advancing.”

Leaving Google for Generalist

14:44 to 15:32

Learn about the motivations behind leaving DeepMind to build Generalist.

“I mean, general intelligence for the physical world, you were probably one of the most research-rich places at DeepMind.”
Show all 22 chapters

Building a Full Stack Robotics Team

15:32 to 16:40

Understand the need for a specialized team to tackle robotics challenges.

“crack the like general intelligence for the physical world problem.”

Training for General Intelligence in Robotics

16:40 to 18:09

Discover the training methods for developing general intelligence in robots.

“The robotics, as we know today, is focused on specialties.”

Emergent Properties of Intelligence

18:09 to 19:29

Examine how models exhibit emergent properties beyond their training.

“And then we now have a model that we can very quickly adapt to, you know, whatever specific type of task you want to do.”

Speeding Up Training Processes

19:29 to 20:31

Learn how the training duration for robot models is significantly reduced.

“use tools in a way that was also not trained for the task.”

The Data Challenge in Robotics

20:31 to 22:20

Discuss the unique data requirements and challenges for robotics models.

“So one of the very simple but interesting theories here is that LLMs, everyone, Anthropic, Gemini, OpenEye, everyone has access to massive amounts of data online.”

Measuring Performance in Robotics

22:20 to 24:10

Explore methods for evaluating the performance and correctness of robotic models.

“There's lots of different types of data.”

Innovations in Gen 1 Robotics Model

24:10 to 26:35

Understand the advancements made with the Gen 1 robotics model.

“So first, just like, you know, what are the models and how do we measure them?”

Future of Robotics and Market Dynamics

26:35 to 28:01

Discuss the implications of robotics advancements for the near future market.

“So that also starts to indicate really being at this kind of level of capability where you can take a new task and you can get the robot to do it without just an overburdening amount of data going in.”

Achieving Competence and Mastery in Robotics

28:01 to 30:03

Explore the significance of competence and mastery levels in robotic tasks.

“which means that we get a robot that's kind of good at a task, which is actually a really important level of capability, right?”

Advancements in Science through Robotics

30:03 to 31:45

Learn how robotics can revolutionize scientific research and efficiency.

Future of Robotics in Daily Life

31:45 to 34:28

Discover the potential impact of robots in various everyday tasks and industries.

“And then like the breakthroughs that could come out of this, it's hard to predict.”

Reindustrialization and the Role of Skilled Labor

34:28 to 36:28

Understand the relationship between skilled labor and robotics in reindustrializing the economy.

“We do have, obviously, millions of businesses in America, small businesses, and you can imagine any one of those people running one of those coming up with ideas that would delight us and serve us in new ways.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00It feels like we are in GPT-3 era for robotics models. We now have a model that we can very quickly adapt to whatever specific type of task you want to do. I understand you have some sort of epiphany around robotics and dexterity. I did my PhD at MIT and then I was at Google D-Mind for four and a half years. You need to have data to learn stuff and everybody's robots. We're just sitting still. We just train a model on hundreds of thousands or millions of hours of data of every single type of possible physical interaction in the world we can think of. What does it mean to apply AI to the physical world?

0:31What's possible right now? Maybe that wasn't five years ago. Things that maybe would have taken decades and large numbers of people. And now we can just iterate on that very quickly. We can just do 10 times more science per year than we could before. And the breakthroughs that could come out of this, it's hard to predict.

0:54Pete Florence is the CEO of Generalists. He's about to change everything about what's possible with robotics. He was an MIT PhD, a superstar at DeepMind, and a generalist. They're now changing what's possible with robots. I thought this was going to be maybe something that mattered in the 2030s based on what I've seen now in the last year. And from our chat with Pete, it's really clear this stuff's going to matter really, really soon. In the next couple of years, robots are going to be able to do everything. Let's learn about it from Pete. We're here today with Pete Florence, the co-founder and CEO of Generalist.

1:21Also my partner Vivek. Really excited to chat about this kind of the really fun stuff going on in the physical world and AI. First, Pete, let's welcome to the show. Let's talk about your background first. Where'd you come from? Where'd you grow up? Thank you, Joe and Vivek. Thanks for having me on the podcast. Awesome to be here. So I grew up here in the Bay Area. It's a fantastic place to be. And then, yeah, you know, long story short, eventually made my way to do my PhD at MIT. That's where I started working on AI in robotics and, uh, been doing it ever since. And you were involved in business role a little bit before your PhD and then afterwards now as well, right?

1:52Uh, that's right. Yeah. I worked at a startup for one year before I went to MIT. I learned a ton there. There was a serial entrepreneur who had done five companies. He either, uh, you know, uh, took public or, or, or, uh, exited and, um, learned a lot there. But then, then I went and did my PhD and I really just, you know, focused on, on research and becoming the best possible researcher I could. And, uh, you know, really, really enjoyed my time there. And I understand you have some sort of epiphany around robotics and dexterity and grad school. Tell us about this. What was the insight? Um, I think, I think there was many different points, Joe.

2:27Um, you know, if, if you go back a little bit, right, like, uh, just a handful of years ago, um, the idea of having a robot that had a very generalized capability in terms of manipulation or what's called dexterity, right? That was very far off feeling at the time. And everybody was like, there was a whole field of researchers that were trying to figure out how do we start to crack this problem? I think if there's any type of like overall epiphany that I feel like I had many, many times over and over again was that it felt like everybody was trying to approach like, how do we bring machine learning into robotics?

3:03How do we bring machine learning into the hardest problem in robotics, which was, I think, kind of broadly felt in the field that dexterity has been kind of the holy grail for the last handful of years. And the place that everybody was coming to it from was just assuming that we'll never have a lot of data for robotics. And we need to do everything we can to try and design our system smarter or like figure out tricks along the way to try and get robotics started without a lot of data. Um, and that was like always kind of, um, like a kernel of a thing that I was always very interested in is how could we figure out how to get on a path where we really had massively scalable learning for robotics.

3:43Um, one type of epiphany that like the thing that I would always have is that if you would walk into a robotics lab, like whether it was at a university or if it was, you know, some of the industrial research labs that started to pop up working on robot learning, you would walk into a robot lab but most of the hours of the day the robots would just be sitting there still doing absolutely nothing they try and do something to at least learn from moving around you think that yeah like it's it's so obvious but like you need to have experience you need to have data to to learn stuff um like that that's like one of the it's how people learn stuff too it's how people learn exactly play games it's how kids learn it's it's how like everything in nature learns and it's how machine learning works you need to have data to learn stuff and and everybody's robots.

4:27We're just sitting still. And it just felt like we have to get on this path where we're actually moving and physically interacting with the world at scale. And then we'll figure out all the rest of like, how do we train the models? How do we do the architectures? How do we do everything else? And you didn't jump, I guess, right from MIT to building a business. You were a senior scientist at DeepMind, right? On the robotics team. What were you doing there? I did my PhD at MIT. And then I was at Google DeepMind for four and a half years. A lot of my research was focused in what I would broadly say is, how do we bring large-scale multimodal learning into robotics.

4:56So, you know, it was the, during the time when I was at Google where like large language models became a thing and a large vision language models became a thing and like multimodal learning in general was starting to become possible. Um, and a bunch of my research was actually both like in some of the fundamentals of that area, but also specifically like, how do we bring that into robotics? How do we, how do we get on this path? Like, like, you know large language models has was just taking off for text and other domains how do we have the same type of capability uh in in robotics and like one of the good things that we saw back then with with language models was transfer learning across domains because of one large pre-train uh what were the first signs of that that you saw on the embodied intelligence side of things and what is embodied intelligence great okay let's get some definitions here so embodied intelligence, I would say broadly, like any type of intelligence, which we can assume that we know what intelligence means, but we can keep going deeper if we want, turtles all the way down.

5:58But any type of intelligence that actually, you know, isn't just existing purely sort of digitally, but actually has some type of a physical embodiment in the world, right? Now, that's what we mean with embodied intelligences. Like we can actually choose how to physically observe the world and choose how to physically act in the world, right? And then Vivek, 100%, the sort of the arc of pre-training, like that sort of broad notion, has been one of the biggest sort of massive narratives in machine learning research over the past decade or two. And it started in very humble ways. Back in early 2010s, people would use what was called ImageNet pre-training.

6:42So it was at the time a large data set, a famous one that people would train vision models on. And then people would see a little bit of benefit from starting on training on ImageNet and then using that same model that had already been trained on ImageNet to train on other types of vision tasks. Of course, skipping over many other things that happened, But large language models really brought pre-training in a much bigger way into the fore. And everybody started to understand that if you take a model, you train it on every single possible piece of text you can find on the Internet, then now you have something that is very capable at many other things than you could, quite frankly, than you would have imagined just, you know, that model would probably be capable of.

7:28And having a similar type of thing in robotics, there's been lots of attempts along the way. So some of the first ones, again, I would say happened in very humble ways, including like, you know, literally using ImageNet pre-training. But some of the things where it really started to upshift, similar to like large language models becoming real, some of that was, you know, some of the research that we did back at Google where we would literally take a large language model and we would turn it into the robot brain. which this is like very simply stated but honestly at the time was like a kind of a crazy idea to do because everybody was just trying to make robot brains in many other ways or if we go to the problems to solve in robotics there's problems of intelligence of robot brains and there's also problems that people are solving with the actual embodiments and the end effectors uh and there's this big debate of like what contributes more to like solving dexterity and solving manipulation is it the end effectors or is it the brain and the first time we met you you told us this very fun story of your days at MIT with tongs.

8:30You want to tell, I actually went and tried this with my own hands afterwards. You want to tell the audience the story? Sure. So yeah, what Vivek's referring to here, back in kind of the middle of graduate school for me, as I was thinking more about this problem of, you know, how do we start to think about scaling robot learning? How do we also get some intuition on like, yeah, what really is the blocker. So this, this little exercise that I did was I was sitting at my desk in, in, in grad school and I strapped a GoPro to my head. So the kind of like very early, like egocentric data collection.

9:05So I strapped a GoPro to my head and then I wanted to emulate as if my hands were just very simple, like robot grippers. So what I did is I grabbed, um, and it, if you know, like woodworking, like there's these clamps that people use are called Irwin quick grip clamps. I grabbed a couple Erwin quick grip clamps into my hands and I just walked around the lab and tried to see if I am just emulating myself as a robot with robot hands and I have these very simple grippers like what can I accomplish in the world and it's just a little kind of like thought experiment exercise and um so with the GoPro in my head I recorded myself walking around and making myself an ice latte and like not like not an easy mode making yourself an ice latte like in hard mode making yourself an ice latte where I went to the freezer and I grabbed, um, uh, like ice cubes and I poured them out into a mug and I closed and then I spilled a bunch of stuff and I cleaned it all up.

9:59And then, then I also think I, it was one of the ice cube trays where I was using the last ice cubes and my wife had trained me that like, if you're emptying the ice cube tray, you have to refill it. So I refilled the ice cube tray and I put it back in the freezer. Exactly. Um, and then I, it was a little Nespresso machine and I had to like open a drawer and I had to open a new box of Nespresso's and I had to dump them out and I had to like grab the little Nespresso capsule and like, like hard mode, making yourself an iced latte. And it was all totally fine. And, um, like we learn a couple of things from that, like exercise.

10:33One is that like pairing human level intelligence with even very simple, uh, you know, as you said, Vivek end effectors, like you can really accomplish a lot in the world. Um, and then it also like painted one path to like, Like we were talking about this problem where you walk around most robot labs and everybody's robots are just sitting still. It painted a very obvious path to like if we could just have tons of people that are effectively kind of sensorized, walking around, collecting this type of egocentric data, we could very quickly start to accumulate just vastly more data than we ever have had in robotics before.

11:08And then once you start to like have that initial kindling of what could then be like a self-perpetuating sort of ability to just keep going. Once we have that initial kindling of like a sufficient amount of robotics data, then we feel like we would just be off to the races and we'd be able to figure out the models and everything else. This was unintuitive to me at first because I thought maybe you needed like really good hands to do things. And you've shown you probably don't. For most tasks, you only need just two on each side. As a side note, why do we have five fingers then? Do we evolve to have backups maybe?

11:39Is that part of it? Because we fight each other or something? Because you probably can do most things with two or even three fingers on each hand, right? Yeah. And, you know, five finger hands like humans have are fantastic. Right. They are very mechanically complex. Like, I think over time, the field will be able to manufacture reliable and cost cost effective five finger hands. But it's also, you know, the main blocker the entire time has been the intelligence, not so much like the sort of mechanical ability. Most things you could deal with. And so if we can keep things simple and we can use that to accelerate our ability on the research side to figure out the model intelligence recipe, then like then we feel like we'll be in a place to, you know, figure out all the rest of the things and including all the more mechanically complicated hands, et cetera.

12:27But Joe, it's a great question. Why humans have five finger hands? I mean, I've studied a little bit sort of evolutionary biology. It's a complicated, complicated world out there. Human hands are fantastic. There's a lot of other end effectors in the world, though, that nature has produced that make very competent beings. So anyway. It seems like the marginal returns right now for more intelligence are higher than the marginal returns to mechanical complexity. That is a very nicely way to say it, Vivek, yes. And you can scale intelligence for a bit. There's more to do. I mean, the current model is Gen 0 and Gen 1.

13:04We can talk about that. But they're pretty smart and you can get smarter. Exactly. Before we dive into generalists, like where we are now, let's just help understand the robotics wave as it stands. Like, what does it mean to apply AI to the physical world? What's what's possible right now? Maybe there wasn't five years ago. Like what's what's going on at a high level? Sure. I would say like very high level, Joe. It feels like to give an analogy, it feels like we are in the kind of GPT-3 era for robotics models, you know, in terms of the development of language models. Right. So we're starting to have models like Gen 1 that feel like they are getting to breakthrough levels of this is a very general model, but it's starting to be commercially viable for very simple types of applications.

13:50If people remember GPT-3, when it came out, it was mind-blowing to a lot of researchers that would interact with the thing. But then at the same time, it really wasn't good enough to be commercially viable for most types of applications. Like now we're, you know, in the world of, you know, all these coding models and all these coding agents that are just like rapidly advancing. Of course, there's like massive applications. So back then, GP3, there was it was clear that it was headed somewhere. It was very simple types of applications like writing marketing copy for ads and those types of things that that were starting to be viable.

14:22And we feel like it's a similar type of era right now for robotics where the general models, they are getting better at a very significant pace. And they're starting to cross into these levels kind of like the GPT-3 era where we're having these general models cross into commercial viability in types of applications that haven't been possible before for robotic intelligence. Okay, let's talk about generalists. I mean, general intelligence for the physical world, you were probably one of the most research-rich places at DeepMind. Why did you leave to build Generalist? Deemine was fantastic. Google Deemine was fantastic in a lot of ways.

14:57I think like I have many lovely things to say about my time at Google, like the 20 percent culture of like encouraging you to go to do things outside of what's supposed to be your nominal job. like the just like very scale-pilled engineering culture that goes back to not just like machine learning but even just like you know in the you know you know coming of age of the internet just building these massive scale systems that supports have you know supported search for decades like that is all awesome stuff um but it overall felt like there was this window of opportunity to go crack the like general intelligence for the physical world problem.

15:35And, uh, there's a few different things that just felt like we could move significantly faster if we were to build a team from scratch that was just like shaped for the shape of, of the problem. Right. And, and part of that is, um, you know, robotics is a very full stack type of challenge. And we wanted to build a team where we have the right types of expertise and just like amazing leaders across every single part of the entire robotics full stack as well as massive scale AI. And then also, quite frankly, too, like an amazing part of building a frontier team in robotics. There's like all these operational challenges with scaling data.

16:18There's also all of these very interesting questions that we haven't yet. We certainly hadn't yet a couple of years ago cracked and how we get these models out into the world. How do we how do we deploy them? How do we figure out the right types of interfaces to be working with the right types of partners to just get these models everywhere? And it felt like we could iterate just significantly faster in a team that we could build from scratch outside of Google. The robotics, as we know today, is focused on specialties. Mostly it's all narrow tasks that are hard programmed, right? This is how everything works.

16:47Like, how do you go about even training something to have general intelligence? Why does that make sense? Good question. I think the general idea is very simple in essence, right? That if you take a model and you give it access to every single type of possible physical interaction in the world that you and I could imagine, and you let it train on all of this, and then you figure out the right ways to not just have the model kind of understand everything, but you also get the model just like when we train chatbots to align them to like, what do we want this incredibly knowledgeable model of text to do?

17:26when we ask it to help us solve questions. Similarly, we have to do that same type of process for these robotics models. So they understand the physical world in a very important way, but then we also figure out how to get them to, when we can ask them to solve some type of new problem for us, and it can apply all of its broad-based knowledge to figure out how to come up with ideas to accomplish what we've asked them to do. That was all sort of very, very, you know, kind of abstract language there. What this looks like more concretely, I would say, is like we just train a model on hundreds of thousands or millions of hours of data of every single type of possible physical interaction in the world we can think of.

18:09And then we now have a model that we can very quickly adapt to, you know, whatever specific type of task you want to do. So when we do this with LLMs, we've obviously seen what seem like emergent properties of intelligence that are not just tied to predicting the next letter. There's some way in which the transformers and everything have emergent intelligence. Is this the case with Generalist 2 in the physical world? We are definitely starting to see that. So some examples that we've seen of very emergent type of capabilities for these models. One that we continue to see, which has been quite surprising, is that we can take the model trained on everything, the raw pre-trained model.

18:52And then we can train it on a new task. And for that new task, we might only ever use the right hand when we are demonstrating that task. But then we will often see that the models have emergent ambidexterity. And we can put the model in a spot where, like, it would just make more sense for it to use its left hand and it just decides to use its left hand like that's that's one example it's very concretely like easily pointed like we can easily point to like this task that we asked it to do we only ever showed it how to do it with its right hand but it already knows how to do it with its left hand and that is just purely learned from the data that's one example um another one is um the ability to have some like ingenuity around how to use tools in a way that was also not trained for the task.

19:46So basically we can ask the robot to, uh, to do a certain type of task where we've only trained it with one type of tool. We can give it a different type of tool and it can know, Oh, I know how to use this tool. I know what the task is supposed to be. And it has that like generalization ability to figure out, Oh, this is a new tool on a new task, but I know how to combine them to solve the problem. I'm allowed to say like when you brush the cube into the bowl, that was pretty impressive. Is that, is that secret? Should we tweet? No, I think we can talk about that. It was very cool. Very simple idea, but to be able to figure out how to move the thing and brush it in, and it's clearly figuring out itself after being trained for minutes, I guess you'd say.

20:17Because it used to take months or weeks to train these things. You're training them a lot faster now, right? Yes. We are now starting to get into the modes where we can train a model on a new task in minutes. Let's talk briefly about the amount of data required and the competitive dynamics. So one of the very simple but interesting theories here is that LLMs, everyone, Anthropic, Gemini, OpenEye, everyone has access to massive amounts of data online. They all can use this data, which is truly massive. It's like quadrillions of things or whatever with lots of logical consistencies. And now you don't have that for a bias.

20:54You have to create it yourself somehow. Your competitors have to create it yourself somehow. Now, it seems like this leads to a more winner-take-all, or at least a smaller number of winners, because it's so expensive. Let's say it takes billions of dollars to get to the point where it's really, really great. That's harder for everyone to get to, I'd assume. So how should we think about this market and that challenge? A lot of the similar dynamics for language models are here, in that you'd need a significant amount of compute. You need an incredibly talented team to optimize every single layer of the training and model stack.

21:27But then, yes, as you're saying, Joe, the data situation is also fundamentally very different. And then also the hardware as well, too. And to a certain extent, hardware, there are aspects of hardware that get commoditized. But the full stack knowledge of how to combine the hardware and the data and get the best performance out of the models with all these things working together, that requires a very sort of interdisciplinary set of talent across the entire team. There might be a huge multiplier. If your hardware is better, you're getting 10 times more value out of the data or something like that.

22:03Something like that. And definitely, I think data is such a concrete example of like, yes, if you're going to make a significant robotics foundation model, um like there's you don't have any chance of just downloading uh enough data from the internet to do that you really have to go find some way to access it um and then the other thing too is that um you know it's not just a game of like get a bunch of data now i have a bunch of data now i can train my model not even close to that right so like what we have been doing for a couple years is closing the loop between large-scale data creation, training the models, figuring out what type of data is giving the biggest boost for the models, and closing and closing and iterating on that loop.

22:57Data is not just data. There's lots of different types of data. There's lots of different types of capabilities that different type of data can help provide for the model. Very similarly to you know, a lot of the labs have these very significant exercises and like, you know, all the data providers and like, we were talking about like finding like PhD level experts to like give the right reasoning traces for like, you know, frontier math questions. And I think it's a little bit different in robotics, but like similarly, there's a lot of knowledge that goes into how do we create the right type of data to create the right types of intelligence.

23:39Yeah. And then when we're thinking about maybe we can just talk a little bit about the two models that you guys have talked about publicly and announced, like, what are the sets of things that like Gen 1 can do? And when we think about measuring how good Gen 1 is or like any of any of these robot models, like how do you measure correctness and dexterity? It's a little bit different than language models where you might have preference data or you might have code where it's verifiable. Correctness is a bit more fuzzier here. Yeah. Great points, Vivek. So first, just like, you know, what are the models and how do we measure them?

24:14And then and then, yeah. And then, yeah, more on the sort of like broader sort of measuring and benchmarking robotics. So our first model we announced back in November, our Gen Zero model, it was a very significant moment overall for the field. It was really, it was the first model to show scaling laws really exist in robotics in really in any significant way, I would say. We're very scaling law-pilled, so this is why we reached out. Part of why I've always enjoyed spending time with the AVC team is sometime after we put out that model, which was in November, I remember seeing a post which was dated October from Vivek, you and Alex on the team about like, we really believe that scaling laws will happen in robotics.

24:59So, you know, we've always been very aligned in our thinking. So that was November. Gen 0 was a great model. It was an important sort of scientific moment. But then, you know, fast forward to April, Gen 1 is just better than Gen 0 in every way. So we can kind of skip what Gen 0 is capable of. But Gen 1 starts to be capable of, really, I think the important thing was being able to hit like very significant levels of performance. We use the phrase mastery to combine like high reliability, like 99 % plus reliability, speeds that are sort of not embarrassing, which is uh it's kind of hard to accomplish for for most of the field right now uh without quantifying it we'll just leave it at not embarrassing speeds uh high reliability um and then this other aspect which um we think is also like really important for for the future of of these types of models and robotics which is this type of improvisational intelligence which is basically um it's a separate thing in and of itself it flows back into giving higher reliability and higher speed, but it's just a separate thing in and of itself where you can put the models into like pretty unexpected scenarios and they can creatively come up with like, oh, I've really never seen this before, but I kind of know what to do.

26:11Um, it's a certain type of generalization and it's one that we think is really, really important, especially for like, you know, real production scenarios where like things in the world change, somebody moves a chair or somebody, you know, changes some part in your workflow and like the models need to know what to do. So anyway, Gen 1 had very, I would say, significant step up in these types of capabilities across the board. And also just another sort of marker of performance here was that Gen 1 was able to, across several different tasks and then many other tasks that we showed since, hit these 99 % plus levels of reliability on just one single hour of robot data.

Read the full transcript

26:50So that also starts to indicate really being at this kind of level of capability where you can take a new task and you can get the robot to do it without just an overburdening amount of data going in. I'd imagine there's ways of measuring progress in this space and there's ways of measuring the things the way you're talking about it. There's also just like, is it commercially useful is probably a really important measurement as well. right yeah if you would have asked me like in like last year frankly i would have said well yes i do think these things are going to be really important to the world in the 2030s and now uh just understanding what's going on it seems like they're going to be important to the world in the late 2020s which is which is surprising to me and i don't know what we say say about this but it seems it seems like like how do you quantify like how i mean these things took like months to train and then they took like weeks and days and out and i guess or take hours or minutes for different types of tasks?

27:45How should you be thinking about how these are kind of moving towards a point where they're clearly going to be useful? I think there's kind of two different capability frontiers that we're mostly focusing our research efforts on today, Joe. So one is in the, we use the phrase competence, which means that we get a robot that's kind of good at a task, which is actually a really important level of capability, right? And so what we talk about internally is rapidly getting to levels of competence. So call that like 50%, 70%, 80 % type of success rate. So not like production levels. But the faster we can get to like competence levels for taking a general model and having it just very quickly learn a new task, we think that that is a very important type of frontier to work on.

28:40And it also is actually we can very quickly iterate on that, in part because it actually doesn't take that much time to measure in the real world if the robot is competent or not. You can basically just try it on the task. A couple minutes later, you have a sense, like, have we achieved competence? But the other really important frontier is the more what we say this like mastery levels of performance. So like really actually getting the right numbers of nines of reliability, like, you know, 99.99, however many more nines percent reliability and these levels of speed and improvisational intelligence so that we could actually, you know, deploy these models at scale in production, doing all these things.

29:20Make stuff cheap. There you go. Make stuff cheap. And better. Make stuff better. Make stuff that we could never make before. That's fair. I do like the idea of just being able to like create goods. like you have energy and you have materials and then you have goods for people if you're thinking about for like the middle class or like for poor people like suddenly they should be able to afford like a lot of stuff they couldn't before that's one side of it but then the other side of it you're saying is like there's actually things we can never have made that are really important for humanity what are examples of that i think um the more obvious things yes are like just taking the the things that already exist in the world and kind of making them better but i think that the things that like we, we, we can't even make it or like if things that we can't even dream about yet, those types of things, I think that that, like, I see that very much happening in part, wherever you, you find that like the ability to just like fundamentally advance progress in things like science or, or, or other fields where just like, there are physical limitations on how quickly we can iterate and progress based on like people or other types of things in the world needing to physically interact with the world right so science i think is a great example and one that's like like close to home for me actually before i started working on robotics i was doing like research and uh sustainable energy i was trying to make solar cells that could be like you know orders of magnitude uh cheaper to scale than silicon this is what i was working on like over a decade ago before i started maybe way more efficient solar cells you could make that you couldn't make otherwise or something like this just as one example and like i would find that like as i would just try and like think about how could i make my research work in that field like 10 or 100 times more efficient like i just needed like 10 or 100 times more me like physically running all these experiments in the lab of course like different types of scientific automation have has existed for for decades but very generalized capabilities to um like be in a science lab and use your hands to pull off all the manipulations that are needed to run all these scientific experiments that has always been actually like one of the inner bottlenecks of just like fundamentally advancing science And that's one where it's like, okay, now we can just do 10 times more science per year than we could before.

31:45And then like the breakthroughs that could come out of this, it's hard to predict. I mean, that's like historically been like postdoc student bottlenecked or PhD student bottlenecked, right? Like it's part of the bet here could be like you could turn every lab into a high throughput lab and you don't have to buy specialized like a high throughput furnace or a high throughput pipetting machine. you wouldn't have to like spend hundreds of thousands of dollars in capex for these high throughput machines and like have a general purpose embodied intelligence with like cheaper arms and cheaper hands just doing all these things.

32:18You just have infinite free nerd stuff. Yeah. I mean there's different excuses there. I think that's cool too. We could have your solar panels powering it and I'll go to the mountain and we'll have a thousand palaces for my friends and I built into the mountain by the robots and we'll have your cheap energy and like free body scanners, you know. So it'll be both. It'll be great. You can do all the things. But seriously, in 10 years, I think it's almost for sure we should have a lot of advanced robotic intelligence in the world. What are you most excited about broadly for the world with tons of these robots doing things?

32:49I feel like there's a few layers to this question. One, obviously, people like to use the phrase dull, dirty, dangerous. There's a sort of first layer of obvious things that everybody agrees we would want a robot to do. things that are like, you know, dangerous for humans or just like, we just don't want people doing those types of things anymore. Um, so that's like just, just one simple layer. Um, then of course, like if we have robots that can do all the stuff that like, you know, maybe it's not dull, dirty dangerous, but like, we just don't want to do it anyway. Like, you know, of course people often think of like, you know, laundry dishes, et cetera, and your house, but like kind of every different part of the economy where like we, people just don't want to do, um, these types of tasks and we can have robots that just help us be way more efficient and way more capable.

33:34I think that is certainly one area, but cleaning up the city or the, or the streets or the highways or anything, whatever, whatever people think of, like, um, these are types of things that like, we, we can imagine today, but like, we just don't have a robot that, that, that there's probably a lot more stuff we can't even think about yet. Yeah. It's, it's really this, this last category of just like stuff that we, we, we, we can't even really think about today. That I think that, you know, when you have individuals that are able to just like have a little, you know, little fleet of robots that can help them like, you know, just build things that maybe would have taken decades and large numbers of people.

34:15And now we can just iterate on that like very quickly, similarly to how like coding tools have just like massively increased the rate of software engineering. I think that everything physical can have a similar type of effect. We do have, obviously, millions of businesses in America, small businesses, and you can imagine any one of those people running one of those coming up with ideas that would delight us and serve us in new ways. 100%, yeah. And it's also on a critical path to reindustrialization. This is a big theme right now in our economy. A lot of us talk about we probably don't have the entire labor pool to do everything we need to do to reindustrialize, but we do have some exceptionally skilled labor.

34:50Like, is there a way that the skilled labor robots work together to do more? Like, how do you think about that? 100 percent. I think like the amplifying effect of, you know, having somebody who has skilled knowledge in a certain type of trade. And now they can just have a fleet of robots next to them that they can like have the robots do some of the things that that that they can do. and how they can just have their business be 10 times larger or do 10 times more different types of things than they could do before. I think that is very much a type of vision that we're very excited about. Sounds like a much wealthier and more prosperous world.

35:32Yeah, again, to draw the parallels of software engineering. Now that everyone who is either technical or non-technical has the ability to be a digital product creator at their fingertips, The bottleneck on that has moved over to review and intelligent review by an expert. So like more software engineers are doing code review than they're doing the generation of code. In the same way, there's probably a parallel here where anyone should be able to assemble or create or dexterously manipulate things that they wouldn't be able to do before. And the expert ends up being on the review and management bottleneck of all these agents doing cool things.

36:15I think it's a fascinating idea, Vivek, and we're going to figure it out. And there's a lot of things that we have to do to figure out how to really make that happen. You've got to get there. It's a huge amount of money for journalists, and it sounds like you're well on the way to creating this optimistic future. Thanks for joining us. Thanks for having me, Joe.

From the publisher

Until recently, robots were constrained to narrow, specific tasks. Think assembly lines, medical devices, or household vacuum cleaners. But now all that is changing. One robot trained by Generalist's foundation model can do everything from folding clothes and washing dishes to sorting objects and repairing other robots. How are Pete Florence and the Generalist team bringing AI to the physical world? Why is intelligence, not hardware, the key unlock for robotics? And how could general-purpose robots unlock 10 times more scientific progress in the years ahead?Pete grew up in the Bay Area, earned his PhD at MIT, and became a star scientist at Google DeepMind. In 2023, he left to co-found Generalist, which builds robotics foundation models to help bring general intelligence to the physical world. Its latest model, GEN-1.5, can now train robots in a matter of seconds on simple tasks such as folding laundry and sorting objects. For this conversation, we're also joined by 8VC Partner Vivek Gopalan. We begin with Pete's entrepreneurial journey and his epiphany in grad school — walk into any robotics lab and the robots were standing still most hours of the day. Pete unpacks why robots, like humans, learn through data and experience, and why intelligence, not hardware, is the next frontier for robotics. At DeepMind he specialized in bringing multimodal learning (language models + vision models) into robotics, and he's now bringing that to the next level with robotic world models at Generalist. We dive into robotics' GPT-3 moment, the emergent capabilities Generalist is seeing — including self-taught ambidexterity — and the data moats that will shape the industry. Finally, we explore what a world of abundant robots means for science, reindustrialization, and American prosperity.

00:00 Episode intro

01:30 Bay Area to MIT and DeepMind

03:55 Pete's epiphany: robots sitting still

08:30 Why intelligence, not hardware, is the next frontier

13:20 The GPT-3 era of robotics

14:55 Why leave DeepMind?

16:50 How to train robot brains

18:30 Emergent ambidexterity

20:40 LLMs vs robot world models

23:50 GEN-0, GEN-1 & scaling laws in robotics

27:28 10x more science per year

32:50 Skilled trades, reindustrialization & future of robots



This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit blog.joelonsdale.com

More from Joe Lonsdale: American Optimist

All 113 episodes
Ep 161: Generalist CEO Pete Florence on the Robot Revolution & Bringing General Intelligence to the Physical WorldJoe Lonsdale: American Optimist · 37 min
Listen in VO