In short
Whether “world models” are the key unlock for Physical AI/embodied robotics, and how they help generality, long-horizon planning, and data collection (including in zero gravity).
Guests (backgrounds)
- Boris Soffman, co-founder/CEO of Bedrock Robotics; raised $270M Series B; builds AI for construction/earthmoving machines.
- Jeff Haack, co-founder/CTO of Odyssey; raised $310M Series B; focuses on world models for general-purpose robotics.
- Ethan Barajas, co-founder/CEO of Icarus Robotics; working on Series B; builds autonomous robots for orbit/ISS missions (Voyager Technologies; robot “Joy”, mission “Joyride”).
Key claims
- World models are causal multimodal systems that learn/predict interactable futures over long horizons; they’re a scalable alternative to pure simulation.
- Embodied AI’s core bottleneck is generality and getting enough diverse data; rare, safety-critical scenarios require synthetic/adversarial generation and careful validation.
- World models are mainly a development/simulation tool; robots still execute learned behaviors (transformer-based end-to-end policies), with autonomy “earned” via staged deployment (teleop → primitives → longer tasks).
Notable examples
- Odyssey’s “Star Child” (multimodal video+audio world model) and “Prowl” (regret-driven optimization using Minecraft to provoke world-model failures).
- Bedrock moved 65,000 cubic yards of dirt on one project; compares scaling to Waymo’s multi-city data.
- Icarus discusses ISS cargo handling: adaptive sliding controllers, human-in-the-loop, and teleoperation due to comms delays; laser comms ~25 Gbps/100ms vs S-band ~1 GB/day.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroducing the Guests
0:45 to 1:16
Meet the experts discussing the advancements in physical AI.
“It seems that AI-driven robotics are all the rage.”
Market Sentiment and Trends
1:16 to 3:05
Discussion on the excitement and challenges in the physical AI market.
“So please join me in welcoming to the show.”
Global Perspectives on AI
3:05 to 4:32
Insights from Boris and Jeff on the AI climate in different regions.
“I think the general sentiment is like feels still full of really deep excitement.”
Generational Perspectives on AI
4:32 to 7:00
Ethan discusses the divide in AI understanding between generations.
“I think there's a lot more stuff that we need to build.”
Challenges in Self-Driving Technology
7:00 to 10:00
Exploring the issues faced by self-driving technologies in real-world applications.
“You have the exact same feature space that everything got to use.”
The Role of World Models
10:00 to 12:00
Defining world models and their importance in AI applications.
“but in many respects, I think it's one of the best ways we have of using the sort of the digital AI phrase used to really help fund and help that development of solving this broader foundational intelligence.”
Simulating Reality in AI
12:00 to 14:01
Discussion on the implications of world models for various industries.
“The way that I think about this, Jeff, is that LLM's model language and world models model reality.”
Challenges of Embodied AI and Data Utilization
14:01 to 17:58
Learn about the challenges faced in embodied AI and the importance of data for model training.
“No, yeah, it's something that we actually face day-to-day right now.”
Understanding Adversarial Scenarios in AI
17:59 to 20:07
Explore how adversarial scenarios are created and their significance in AI training.
“but actually still be pretty exposed when it comes to the rare cases.”
Advancements in World Models and Their Applications
20:08 to 28:00
Discover the latest advancements in world models and their applications in AI.
“This means as you're predicting frames forward, you're getting both pixels coupled with an audio wave at the same time.”
Show all 27 chapters
Exploring World Models in Robotics
28:00 to 33:20
Learn how world models serve as critical tools for developing robotics in autonomy applications.
“when you deploy physically your world is your simulator in some sense and so you basically just apply your final output in the final model you train.”
Challenges of Autonomous Operations in Zero-G
33:20 to 37:10
Understand the complexities and challenges faced when deploying robotics in zero-gravity environments.
“Let Boris respond to that, then keep going.”
Data and Communication in Space Robotics
37:10 to 42:00
Discover the implications of data collection and communication technologies for space robotics operations.
“And I'm curious when we're going to get Moon 1 from you guys.”
Data Challenges in Space Missions
42:00 to 42:51
Discussing the challenges of data retrieval in space missions.
“And so this is a massive increase from what currently is state-of-the-art.”
Data Accumulation in Earth Moving Projects
42:51 to 45:29
Exploring the impact of data accumulation from large-scale earth moving.
“Boris, we're talking a lot about, you know, where the rubber meets the road here, or I guess where the robot meets the zero G, but you're definitely out there in the real world.”
Public Perception of AI and Optimism for the Future
45:29 to 47:50
Discussing the need for a positive conversation around AI advancements.
“And these are subsets of projects that get into the many millions of cubic yards of earth that become the opportunity.”
Comparing AI Evolution to Historical Transformations
47:50 to 50:46
Drawing parallels between AI advancements and historical technological revolutions.
“of the astronomic positives that have happened in every single technological transformation that are actually harder to predict.”
Space Technology's Impact on Daily Life
50:46 to 53:13
Highlighting how space research benefits everyday technology and health.
“Like, to me, those got me so excited about the possibilities of going faster and further.”
The Need for Better Communication in AI Development
53:13 to 56:00
Discussing how to improve communication about AI's benefits and impacts.
“random positioning machine to try to replicate zero G because we can't.”
The Challenges of AI Regulation
56:00 to 57:26
The discussion highlights the difficulties of regulating AI technologies and the need for communication with the government.
“It's almost like saying we're going to government control software.”
NVIDIA's Role in Physical AI
57:26 to 58:45
The hosts debate NVIDIA's recent announcements and their implications for the physical AI landscape.
“And I was like, that's not at all how we're going to build the coalition here to keep this from becoming a bipartisan issue that's going to get crushed.”
Complex Systems and Robotics
58:45 to 1:00:45
The speakers discuss the complexities of integrating robotics with physical systems and the importance of safety.
“Now, Boris, there was also something from NVIDIA called Halos for Robotics, which is full-stack open robotics safety system.”
Evaluating Autonomous Systems
1:00:45 to 1:02:38
The conversation delves into the challenges of evaluating and testing autonomous systems in complex environments.
“I think autonomous driving is sort of a little bit ahead of the curve compared to some of the other fields, just because it's a bit more of a mature industry, a little bit more defined problem at some level.”
Precision in Robot Control
1:02:38 to 1:04:36
The hosts discuss the precision required in robotic control, especially in environments shared with humans.
“It's actually harder in your case with bedrock because if you think about like modern equipment, it's often hydraulically driven, right?”
Humanoid Robots vs. Purpose-Built Machines
1:04:36 to 1:09:23
The discussion compares humanoid robots to purpose-built machines and their respective advantages in different environments.
“Are there still tolerances you can play with or is it like millimeter precise requirements?”
The State of Physical AI
1:10:07 to 1:11:30
Discussing the current state and future predictions for physical AI technology.
“And I'd say it sounds like it's a little bit longer because it got a little bit of a mess with, like, kind of, you know, government, but now they've got to, you know, to go with it.”
Industry Predictions and Challenges
1:11:31 to 1:12:37
Exploring the potential timeline and challenges for advancements in robotics and AI.
“And the GPT-3 moment for robotics, future or past?”
Transcript
Automatic transcript. May contain errors.0:00It's the year of physical AI. 80 % of the world's GDP is still grounded in physical industries. There's people that are using cloud code and codex and co-work on their own, developing things that we've never seen before. Can make 100 ,000 and employ them everywhere and democratize labor. And when you look at the earth and you say, okay, half of the earth's GDP is people doing things, it's human labor. There's a very compelling argument there. Are we still approaching the GPT-3 moment for physical AI or have we already gone past it? I think we're so early. I think it's a far, far time away. Thanks to our friends at PayPal, the exclusive sponsor for This Week in AI.
0:36Try the payment and growth platform that's trusted by millions of customers worldwide. PayPal Open. Start growing today at paypalopen.com. Hey, everybody. Welcome back to This Week in AI. My name is Alex. Now, it's the year of physical AI. It seems that AI-driven robotics are all the rage. People are worried that hardware will become the only moat possible. No one has enough compute. and companies ranging from digging in the dirt to flying in orbit are pursuing different approaches to understanding and interacting with the world. So on today's episode, we're sitting down with the CTO of a world model startup, the CEO of a company automating construction equipment, and the CEO of a company that wants to take autonomous robots up into orbit.
1:15I really want to understand how quickly the physical AI market is developing, which type of AI model is the best suited to bringing AI into the real world, and what the economics of AI in the world, or perhaps AI of the world look like. So please join me in welcoming to the show. It's Boris Soffman, the co-founder and CEO over at Bedrock Robotics. Boris, good to have you here. Good to be here, Alex. Thanks for having me on. And Bedrock just raised a$270 million Series B, so shout out. We also have Jeff Haack, the co-founder and CTO of Odyssey. Jeff, how you doing? Doing well, thank you. Thanks for having me.
1:47$310 million Series B, just pipping what Bedrock put together. And then we have Ethan Barajas, the co-founder and CEO over at Icarus Robotics. Ethan, welcome to the show. And where the hell is your series B? We're working on it, Alex. We're working on it. But thanks so much for having me. No, I'm so glad you guys are here. I want to start with something a bit broader, though, than going super deep. Because when I look out around the tech market today, it feels like in the last couple of weeks, the vibes have changed. And I think it's fair to say that we're all AI bulls here on the show. Don't mind a data cinema when I see one.
2:22But it seems that public sentiment has shifted. It seems that the markets are choking a little bit. And so I'm just curious if this vibe shift that we're seeing outside of the industry is also taking part inside the world of AI and robotics, or if this is just stuff that's more on the media and public opinion side of things. And Boris, I thought we'd start with you. You look at the last couple of years, it's felt like it's just been a nonstop march outwards in terms of energy and excitement. There is something foundationally true about the fact that you could do things with physical AI that just were just infeasible five years ago.
2:54And so it's very exciting because like 80 % of the world's GDP is still grounded in physical industries. You have tremendous opportunities in transportation and industrials and manufacturing and all these areas. There'll be ups and downs. I think the general sentiment is like feels still full of really deep excitement. there's just also a kind of a noisiness and a you know kind of like a mass to the messaging and so we haven't sensed too much of a reprieve in it it's still incredibly hard to find talent it's really you know hard to like put all the pieces together compute still expensive all these things are still present but I fully expect there to be some ups and downs in the near future.
3:34Okay that makes sense and that tracks what I'm hearing in the states but more Jeff you're over in London. What's it like over there? What's the temperature? What's the level of excitement versus fear, if you will? Literal temperature is sweltering. It's a heat wave in Europe.
3:52We sort of exist in sort of the, what I probably call it, a sort of AI research epicenter, where there's a lot of focus more on sort of the core fundamental AI research. And I'd say the general sentiment is very strong and quite bullish. There's sort of two macro themes that a lot of new labs and the large labs are also focusing on, firstly around this idea of recursive development of AI, and then also in world models. And both have a pretty major presence, and at least from what I've seen, the sentiment is very bullish on both. So I think it's early days, and as much as we've seen a march up progress over the past year, I don't think we're done, is maybe the way to say it.
4:31Glad to hear that. I think there's a lot more stuff that we need to build. Speaking of which, Ethan, we're going to talk about your company in a couple of minutes, but I hate to say this, as the designated young person on the show today, what's the vibe like out there amongst people that can't rent cars? Yeah, still get that 500 bucks extra every time. I know. It is literally ageism. I don't mean to put you on the spot, but I'm curious because I feel like Gen Z being super digital is probably imbibing a bit more of the discourse here, if you will. So it's quite funny, actually, because like myself in the position I am in the tech sphere that I live in, I'm so far in this bubble.
5:10And then when you interact with other people outside of this tech sphere, it's way more polarizing than you would actually think. And, you know, even with my friends from my hometowns, whenever I make it back and get to talk to them, you know, they're not implementing it in their daily lives. And this is actually something that's crazy. We're still so early. You know, we sit on two ends of the spectrum. One end of the spectrum is we kind of push the frontier in this very novel domain, which is zero G. You don't necessarily see world models built for it. You don't see lots of data of teleoperation collected for it and expert examples.
5:40You just don't see that. And when we try to talk to even some of the folks at NASA, it's a crazy conversation. Safety, it's not Cartesian control, model predictive control. It's this unknown value still that they've just never seen. But then when you talk to some of the people on the implementation side, even terrestrially, some of the stuff that we're doing still in the tech sphere, it's still super exciting. You know, this is the thing that people jump for and it's more of the mainstream what you come to expect people talking about in the same excitement than you expect from the mainstream. It's very different on these two other sides.
6:10And I think when you look at, you know, our demographic, it's even more split than ever right now of people that are saying, oh my God, AI, like you can use it for what? And physical AI, what does this even mean? I've had to have that conversation so many times of what is physical AI that's not an LLM. And then on the other end, there's people that are using Cloud Code and Codex and co-work on their own, developing things that we've never seen before. I'm a huge Codex guy. And my spouse recently, for the first time, used AI to do a thing. And I was like, oh, we live in two different worlds inside of our house.
6:44And she was like, yes, honey, that's your world. And I was like, okay, fair enough. But it does surprise me. Boris, on the point about people being more excited about physical AI, perhaps than purely digital AI. Are you seeing that same split that Ethan just discussed? Yeah, well, there's like a feeling that there's a, like digital AI just matured a lot. Like it moves fast. You have the exact same feature space that everything got to use. You had infinite data on the internet. And so you started to see this snowball that was shocking, impactful, very, very exciting. On the physical side, it's like, it's genuinely harder.
7:16You don't have the same sort of access to data. You have uniqueness in the hardware and the sensors and everything, But the opportunities are tremendous. And you see it in, you know, just how much Waymo is starting to progress and, you know, the value it's created and, you know, the way it's starting to scale. You start to see the opportunities in assembly and, you know, the sort of things we're doing. So I think there's an excitement that it won't necessarily have the exact same shape as like Anthropic or OpenAI did in these kind of horizontals with things wrapped around it. But you can actually go and tackle these problems that have astronomical value to them.
7:52And what's interesting is that it's like incredibly expansionary. It's not like a replacement of labor in a lot of cases. It's that you can actually start to do things that were massively supply constrained or just weren't possible before. I think we're not even scratching the surface of how transformative that can be because oftentimes AI is thought of as a cost-cutting measure. In reality, it can be incredibly expansionary if it's actually applied in the right way. And you see that more in the physical world in some ways, although you've seen that in coding and other aspects as well on the digital side.
8:20It's also famously expansionary when it comes to your budget line items that you have to then report to the CFO. Yeah, good to be AWS in this world. It's great to be anthropic in this world. They put out a new model. They double the price. Everyone goes, can we please have some more? I mean, it's just best place in the world to be. Jeff, Boris just brought up Waymo, and I was going to get to this later on, but let's do it now. I've been a little bit surprised to see the issues with self-driving cars as they've scaled recently. And I've talked to a number of people in the self-driving world, you know, Wabi and Wave and so forth.
8:53And I've talked to them about world models and kind of where they use their technology and so forth. But I thought no matter what underlying technology self-driving cars would use, that as they scaled actually in the physical world, getting off simulation, getting out of the world model, they would get such rich data that we would see improvements. But instead, if you look in the last month or two, the headlines about self-driving progress, I mean, you know, Waymo doing the major recall about construction zones. And I think there was an issue with school zones. And then Tesla is in trouble for several different things.
9:21Is there a problem here or is this just the standard kind of like, you know, AI meets the real world, there's going to be some teething problems? I think it speaks to the sort of core problem that sort of all embodied AI faces, which is really one of solving generality. And solving sort of an open-ended world problem is really, really, really difficult. I mean, I respect that was sort of one of the core theses of what led to Odyssey, this idea that you have to, Boris, you mentioned, for example, needing to scale data. And it's just hard to get a lot of data. The same is true regardless whether it's construction vehicles across different types and different environments.
9:51It's an immense challenge. So in many respects, or the bet we're making is that wall models are that unlock for generality for general purpose robotics. The good thing is it's not just an embodied AI problem, obviously as applications there, but in many respects, I think it's one of the best ways we have of using the sort of the digital AI phrase used to really help fund and help that development of solving this broader foundational intelligence. Certainly our bet. the fact that you would bring world models to the table as maybe a sort of analogy. Think of this as like the 18 years you spend learning how the world works before you then to go and do 30 hours of driver training.
10:27There's also a statistical element to it where like you drive 200 million miles, you're kind of poking statistics so many times that you start to really, really feel the long tail and all these like kind of interesting challenges that just wouldn't have come up when you're doing kind of testing and development and small scale deployments. And so it's the real world, like you start to really get exposed to it. Scale is big. does that make world models and we're going to get into defining those in just a second but does that point you make there Boris make world models more or less useful when it comes to bringing AI out of our screens and into the world?
10:56I think it's incredibly useful I mean in fact we kind of go through you know similar thought process where you know we're creating general capabilities for specialized type of machinery so these are manipulation machines that manipulate whether it's like earth or aggregate material you know lumber farmland and these are all like kind of specialized manipulation machines. And you can get off the ground by using more traditional simulation approaches where you're modeling kind of earth physics and sensor models and so forth. At some point, it is unequivocally the right long-term answer to go to world models where you like your own data, kind of teach these systems in order to represent the right physics.
11:40It becomes more scalable, becomes easier to start doing things like reinforcement learning where you need computational efficiency. So it feels like it's still, you know, there's a lot of complexity to it and you do have to have the right data for particular domains. But it feels like that's the future of simulation when you fast forward to the number of years. Let's talk about world models. The way that I think about this, Jeff, is that LLM's model language and world models model reality. your company odyssey says that uh world models are causal multimodal systems that learn and predict and interact with the world over long horizons can you translate that into english for everyone out there who does not speak nerd think of this as like a learning to simulate what happens so you know some information about what you're seeing and what you're learning is potential futures so you think about uh think about like you would a sort of computational simulation rolling forward in time um at its simplest that's kind of the way the best way to think about it.
12:33Now, the question is, what are you simulating? The easiest form of data is visual state, and that's where we're seeing most progress. And certainly in the sort of physical AI category is most world models have an inherently visual element. We don't see ourselves stopping there. Audio is a very useful extension. The way I like to think about it is it's almost like mimicking human senses of sight and sound, right? It's how we perceive and interact with the world. And I think that's a good place to start in terms of state. I think a lot of people consider world models from the perspective of self-driving cars, which makes it appear to be kind of like a video game style interface.
13:06Is that the way we should think about like visualizing a world model or is it something a little bit more esoteric than that? And I'm thinking about the candy cane version versus the big boy edition. In one level, you can think of it as a video game. I mean, to be honestly, the industries we focus on primarily are games and robotics. In many respects, it gives us the sort of the complete, the sort of maximal coverage of a couple of industries that are big in their own right. But from a capability perspective, allows us to think about most of the things that matter to world models. So yeah, I honestly do think the same model will be used for general purpose robotics in terms of autonomous driving and will also sit underneath future video games.
13:41And I think that's kind of the key crux of this, where if you want to really, really solve this foundational technology, you have to solve generality. And as a result, that means that learning from games benefits, for example, construction diggers. But it doesn't help Ethan, because generality, I think, works quite a lot when you're in our gravity well, you know, at 1G, when you're up in 0G, there probably breaks down. Ethan, am I wrong here? No, yeah, it's something that we actually face day-to-day right now. I mean, we have our photorealistic sim, most people use MDIs, XM, and things that look like this.
14:15We use it internally, and it just, it doesn't get you there right now. And actually, it's kind of a question I have for Boris as well, where, you know, we're on that tail end of embodied AI using robot learning from expert examples. We can teleoperate using laser-based comms from Earth to space, And this is amazing within low Earth orbit. Once you move out farther and farther away, you're bound by the laws of physics. You have communication delays. You can't do it. And there's this ongoing debate right now within the field of how many expert examples, how much data can you actually throw at these models before it's diminishing returns?
14:47And it ties into what we're talking about a second ago about Waymo, about Tesla with autonomous driving. And how much actual data can you throw at these current models and heuristics and training pipelines that we have before you need to move into something new. And I think that's why it's so key that in the future, you have to have things like world models to collect some of this data and augment what you have in the real world. But we're still missing that key point. And so for certain industries, you know, things like construction, where you literally cannot move as much earth as a digger can, right?
15:19It's so important to have those capabilities or in space, when it costs, you know,$130 ,000 an hour to keep an astronaut alive to do something. You need that actual capability. Right. And so I think these are some of the key questions that the industry on the research side has to figure out moving forward. The challenge, it's an interesting challenge because like scaling was it's been surprising in the space how far they still go where you get improved performance if you have the right capacity. We saw this so Waymo and Thomas driving. You see this in research with, you know, grasping and kind of human noise.
15:53you see with LMs. And so it's kind of, you know, you think more and more of it should be clever architecture and so forth, but there is just like an element of, you know, scaling laws. Oftentimes it's getting the right data. And so you end up having to create adversarial scenarios that are very rare and kind of like push your learning rate faster because... Boris, can you explain what adversarial means in that context? I think a lot of folks listening don't fully get it. Of course, of course. Yeah. Sorry. This is where you are basically upsampling really rare scenarios that matter a ton, but you don't show up too often.
16:24And so in the case of autonomous driving, maybe that's like a toddler jumping out from behind a occluded car, and you're not going to drive around hundreds of millions of miles waiting for that to happen. You're going to try to create that using synthetic data or close course testing. We use mannequins or spicing in real data into a previous log. And so you're effectively manipulating your distribution so that you're creating examples that really matter and give you the most learning value out of it. And this is where you start to see hybrids of simulation and real world, but you have to get creative because what's challenging when you get into these safety-critical applications is that you're really optimizing for the one-in-five-million-mile case in the case of driving or a rare interaction with a dangerous heavy machine or these sort of like rare scenarios that Ethan might be dealing with in space.
17:20And so this is where you actually have like a bit more sensitivity. And the challenge of the world models has to overcome is that by definition, you're modeling kind of the distribution that you've learned. And you have a challenge now of how do you prove that that's actually representative of what you're going to encounter. And so this is one of the interesting kind of jumps where you go from an LLM where the consequence of being wrong in most cases is not that significant. and maybe you start to get into tougher long tail challenges with medicine and law, and you go to something like autonomous driving or these very, very consequential physical robotics, you're actually quite sensitive.
17:57And if you're not careful, you can get a really great representation of your very common situations, but actually still be pretty exposed when it comes to the rare cases. And so this ends up being one of the deepest challenges of a lot of robotics applications. We have these dangerous, physical, complex systems that need to be incredibly safe. In respect, one of the big challenges that we've seen is getting that, for example, you mentioned, what's a 1 in 5 million mile example? How much driving data is there with, I don't know, an elephant in the road? As big as Whamu is? Probably not many. And as a consequence, the question is, where do you get that information?
18:30I think this problem compounds as well with humanoids or anywhere where you're sort of needing some very specific embodiment, specific environment. And then, actually, Ethan, you point out, it gets worse as soon as you hit space. there's nothing inherent in the architectures of these models that is inherently limiting in how they're used i think the challenge is really understanding how to build them there's still a lot of algorithmic work going on coupled with where do you bring in data and how do you sort of bring them all together uh that sort of is the recipe that um that i i think is necessary so to your question of whether can this work in zero g uh yeah i think it absolutely can work in zero g it's just a question of is it are we bringing in that right data i think it'll cut right if we think back in terms of language model evolution, the early versions of GPT were kind of rubbish at biochemistry.
19:15Now they're pretty good. And there's a whole lot of things that in terms of maybe academic fields or very specialized domain knowledge, which will be brought into world models, even if they aren't today. I want to go a little bit more on world models because we're talking about them in a very specific context, but Odyssey does quite a lot of research. You guys recently put together something called Star Child. And also this is the best, I presume, acronym I've ever heard. But if you take the prioritized regret-driven optimization for a world-earning model, it becomes prowl. So tell us a little bit about the state-of-the-art for world models, and then we're going to apply that to our two other use cases.
19:47But what's the state-of-the-art, how fast is it progressing, and do we have enough compute? How fast is it progressing? Very quickly. Do we have enough compute? Never. It's a special state of the art. So Starsheld was the first multimodal world model. So this is basically not just visual state but also bringing in jointly generated audio. This is a bit different from what you might have seen in bio or avatar models for speech. This means as you're predicting frames forward, you're getting both pixels coupled with an audio wave at the same time. Which is a, turns out was a more difficult bruce problem than we expected.
20:20Which is probably why people hadn't done it. Why was it harder? Maybe the easiest way to describe it is if think about predicting one frame of video or a very short snippet, the actual audio information you're dealing with is very short. So it might be half a phoneme, so a phoneme being sort of a phonetic unit of sound. It makes it quite difficult to get sort of stable coherent generation, particularly as soon as you get speech. Now, this might seem like it matters more for sort of non-embodied AI, but I think it's quite an important signal. In many respects, models that can generate and model the world well are also very useful for understanding the world.
20:56And audio is a big part of what happens, right? Be it robot grasping, if something cracks or sort of crumples, that's a very useful feedback sound. If something happens sort of out of sight, but you hear an audible sound, that matters. So hence we see this as a core fundamental capability that has to be baked into world models. So that was Starchild. You mentioned Prowl. That was a piece of work we're quite proud of. historically world models have been an intrinsic part of reinforcement learning where you could think of the world model as a learned transition model or a learned simulator coupled with an RL agent and these are kind of two sides of the same coin.
21:33Now what was interesting historically is most people are focused on improving the agent in terms of decision making. Papers like CIMA from DeepMind is a good example and that's really where a lot of the research is focused. We looked at us and said well actually you kind of have garbage in garbage out problem so why don't you actually improve the world model. So what we did is basically send up a simulation environment using Minecraft and effectively trained an RL agent to provoke failure in the world model as it was exploring the world. Maybe it was failure to adhere to actions. Maybe the geometry failed.
22:05Maybe the physics failed. Lots of different reasons. And essentially it's rewarded to find failure modes. So as a result, you get a really nice framework that allows you to essentially boost the level of your performance in your world model given that simulation environment in an automated way. So you crack the flops, warp bottle goes up. How do you determine what is a failure state? Because if I'm thinking about, you know, construction machines, a failure state is if the bucket hits someone in the face and kills them, right? Bad, very bad, failure state. In Minecraft, there's not quite the same level of impact, which is slightly funny given what I just said, but it seems like a difficult environment to define failure state.
22:40Typically, you've seen like games that are often the starting ground of where a lot of AI development happens. This is a very long history through a lot of the early work I mean, DeepMind is a canonical example, also OpenAI. So I think it's the right sort of environment to explore these kind of concepts. There's no reason that you are limited to Minecraft, and we'll have some more work coming on that soon. So, for example, any ground-trough simulation environment can be used in this way to juice the performance of a fully learned world model. So that could be, for example, a space simulator, an Earth-moving simulator.
Read the full transcript
23:09There's no reason that can't be used in a closed-loop manner in the same way. Okay. So Boris, given all of that, I'm curious about how you guys have picked what AI technology to use inside of your machines, because you guys said in your public materials that you use large-scale machine learning models. Is that a one-to-one to world models, or is there a variation in what Bedrock uses to power its machines? The world models will be used for your kind of simulation and training purposes. Like the actual like ML model is effectively a large transformer based kind of end to end architecture. It's taking as an input sensor data, machine signals, goal of what you're trying to achieve, you know, a variety of other factors that kind of give you the pose.
23:51and then it's processing all of this with a particular structure and outputting a behavior for the machine as well as potential other behaviors like honking the signal that you're ready. And so that behavior can actually govern a complicated machine like an excavator that has six to nine degrees of freedom that the axiom architecture could power, a wheel loader, a bulldozer, a dump truck and other aspects. And so you're effectively learning from a variety of inputs. What's the core of it can be human demonstration where you have a lot of like examples of, you know, of like huge amounts of hours of actual real work.
24:30This is actually similar in spirit to how like the foundation for Waymo's training was, you know, the big shift that really caused it to break through was learning from hundreds of thousands and millions of miles of driving. They kind of captured the nuances of how you interact with downtown San Francisco. So the same subtleties exist in how do you interact with, you know, kind of complex different materials and, you know, get like inch precision on, you know, cuts and grades, do demolition, deal with different tools, different machine sizes, different types of machines. And so we're learning from all of this and effectively enabling these systems to be able to behave like an operator and work towards achieving a goal.
25:08And so that goal could be a large kind of cut to build a data center. It could be a trench to do kind of piping a foundation. Sure. And then, you know, eventually like this becomes the property that you start to see is that the foundations of what you learn in these machines, they actually carry over and give you a huge shortcut towards the next capability of the next machine. And we saw this at Waymo where going from San Francisco to Los Angeles, Phoenix, Austin, Atlanta, Miami, you start to get a higher and higher subsidy and less and less new data required until you're, you know, almost more of an operational and qualification problem.
25:46We saw the same thing going from car to truck, from surface street to freeway. And so if you structure things in the right way, you start to learn more and more of a foundation of what is it that enables large physical machines to manipulate the world around them. And that does more heavy lifting. And you need less and less incremental data that's specific to the new application. And it starts to get a similar property to how LOMs have generalized across language and other applications like grasping and so forth. And so it's quite exciting. And then the world models end up being a really powerful potential tool to simulate and do offline kind of like iterations, training, get more training data that doesn't just have to be physical.
26:27Because in all of these applications, physical testing is so expensive that what you really want to be doing is using the physical world to train your simulation. And in a simulator is actually where you develop and qualify your systems and deploy because that scales much better. And so you oftentimes will see one to a hundred to one to a thousand kind of price differences in a physical hour versus a simulated hour. So you want to push it into that direction as much as you can. So you start by collecting real world data, put that into your simulator to create a strong world model that you can do more testing, then take the results of that and then put them into the machine again and send it off on its own.
27:03Yeah. And the world model becomes like a tool to do all this. Like we've kind of experimented with the mix of elements. But yes, you do a lot of training offline where you're effectively training a large scale model to emulate the best behaviors of a human. And this is where it's interesting because you're not just optimizing just for safety. You actually can have all types of metrics that you actually care about. And the same way that a Waymo is 10 times safer than the human, you can actually become superhuman by upsampling the best behaviors and downsampling the worst. And this is where you become the superhuman operator that absorbs more and more scope over time.
27:34And just explain for me in idiot terms why you don't want to use a world model on a machine in the world and you want to use this transformer-based approach you mentioned earlier. Is it just a lower compute footprint? Is it just a better fit for the work? I would think of a world model as an environment in which you can realistically emulate the real world and simulate. When you're on a machine, you are in the physical world. So you're actually kind of like now acting on it. World model becomes an incredible tool for development in these sort of autonomy applications where you can climb the quality of your safety, your behavior, create scenarios you never encountered, test new versions of your software in scenarios that you would have to physically recreate.
28:23when you deploy physically your world is your simulator in some sense and so you basically just apply your final output in the final model you train. Now as Jeff was mentioning in the case of like a video game your world model could actually be the output because your end product is a virtual experience in the case of robotics your end product is a physical behavior and so you use world models more as a development tool. Maybe it's worth adding um there's sort of a fairly recent sort of approach that people have been using it basically using the world model as a backbone if you're familiar with the concept of a vla a vision language action model you can kind of think of this as deleting the language model and putting a world model in its back in its place as a backbone um and then uh effectively using that um at least the current research in literature suggests that effectively that is a more useful prior or more useful information um to to learning it but as boris says you still need to you still means you'll train that backbone and actually train it for the sort of the the behavior that you're looking to to solve with your robot i mean if i can jump in here like this is something that we use internally on the bla side of things and like boris i have a couple questions for you actually like how long horizon can your tasks actually get are they high scale perimitums right now are you going end to end very long horizon tasks across an entire construction site of multiple different things because this is something that we deal with number one we just don't have the quantity of data of any other domain.
29:49Right now, the amount of data that's been collected of high-scale, high-fidelity manipulation in zero-G dealing with the dynamic coupling problems of picking up something that is unknown mass and then trying to manipulate that and move it is probably very similar to the physics of picking up unknown masses of dirt and earth and moving this terrestrially. In fact, some of the people on the team wrote controls for Caterpillar and some of the heavy machinery there because of how similar, we call it the dump truck problem, actually. If you pick something up of sufficiently heavy mass and you try to move it, the controls of this entirely changes.
30:23Changes your entire controls, yeah. Exactly. And so this is why we have to rely so much on real world data because as much as you can have photorealistic sims and world models, it's just not enough fidelity. Number one, for safety, but number two, for you to have operationally whatever your primitives are done autonomously. And so for us, we're lucky where the business case closes. I can pay someone$200 ,000 a year to sit here and operate this robot. And because the cost of labor on orbit is so high, it closes for us. And this is another hot take I have. Teleoperation gets such a bad rap right now.
30:58Everyone's like, oh, it has to be fully autonomous, full autonomy, full autonomy. That's the only way robotics work. And in certain industries, this is just not true. People don't care who moves their cargo on orbit. They don't care if it's a magical rainbow unicorn, an astronaut or a robot, and they just need that task done. And so for us, we found this one place where we can deploy in orbit. We can have a human in the loop, which NASA loves because safety wise, deploying autonomous systems that are based off of simulated data and not collected data, as well as based on things like world models.
31:31It's just not there yet because you have humans floating around in a tin can up there. If you poke a hole in the wall, this is a very bad situation for everyone involved. And so, you know, for us, the way that we think about it is you deploy sequentially and full autonomy is this thing that's earned, not given. And so you deploy first the teleoperation with you collect all of your sensor data of the environment around you. You understand the physics of just moving in six degrees of freedom, your base before you even pick something up. And then as you pick something up, then you move forward. You start to understand what does this look like of moving an object that's similar mass to our mass of our robot.
32:07And from there, you can start to roll at high level primitives. So let's say I've collected, you know, with the VLA, for example, like Jeff was talking about, let's say I've collected, you know, 150, 200 examples of myself picking up an ISS cargo bag and moving it to, you know, the Columbus node or Columbus module. I can then at a high level tell this robot, hey, go pick up cargo bag in slot 3L and move it to Columbus, node three or whatever it might be. And we can do that high level primitive. And over time, as we collect more data of those edge cases and train and simulation of those edge cases and augment our actual model, we can then build out more and more longer horizon tasks that are not compound.
32:50Instead of moving a cargo bag from A to B, I can move the cargo bag A to B, unzip the cargo bag and unpack it. And after I've done that, I can start to implement experimentation and plugging things in and plugging things out all on one learned behavior, all on one actual task horizon. And over time, as we do this, we build that corpus and everything becomes more and more higher fidelity. And then you earn autonomy, but it's not necessarily something where you can just go in the computer, run thousands of hours and deploy. All right. Ethan, pause. Let Boris respond to that, then keep going. Otherwise, he's going to forget the first thing you asked him and then we're going to have to do it all over again.
33:26I love you, but let's give... That was like seven paragraphs of interesting things. We'll have to unpack it slowly. Super fascinating. I mean, this is an interesting challenge in like every kind of autonomy application where you can actually take advantage of potentially a split where you think of the long horizon planning as almost like a reasoning model that like lives above, you know, everything else that's happening. Our version of this is you might have autonomy for an excavator doing a task and then a bulldozer and a dump truck. And each of them internally has, you know, the model that kind of operates their behavior, but it's more local.
34:01Like they're doing this kind of strip of a cut. They're doing kind of loading tasks. The reasoning model becomes kind of the orchestration model becomes an opportunity to coordinate all of them. And you can take advantage of the fact that you have very different constraints in that model where the physics are a little less relevant. And you can learn from much broader data and break up this problem in a way that isn't as vulnerable to the sort of physical challenges you mentioned. And we have the exact same ones where, you know, you could be carrying, you know, two tons of earth in one bucket kind of stretched out, you know, 20 feet.
34:35And that creates like tip over risks, different vehicle, you know, vehicle dynamics. And so if you're able to do reasoning in a higher level model that then hands off to this much more kind of physically complex kind of reasoning that has to happen locally, you can potentially simulate that reasoning model a lot more aggressively. and then optimize the components that are actually much more vulnerable to the physics of space or the specific controls of an arm or whatever the challenges might be. And so it becomes like a decomposition of complexity where you kind of have a way to lean into the strengths of each approach.
35:19And it makes it more simulatable versus taking on the biggest pain points of every single one out altogether. I've been looking for this for the last two minutes. So here it is. If you're curious what the Columbus module on the ISS looks like, not that world's best resolution image, but there she is right there. It is literally a 10 can. It's like one of those eight ounce Cokes you get at the dentist's office you don't really want. Very small little squat thing. All right, Ethan, back to you. No, I mean, we think about it the same way. It's going to have this higher level thing, planning these individual tasks, and then you have to get much less high fidelity data of that, you know, one-to-one human operator.
35:53And then you can start to expand this out across. And then to kind of jump on, one of the biggest tasks for us is every about 60 to 90 days, depending on resupply schedule, we send up three and a half tons of cargo. And we have some of the most highly trained people, highly trained scientists just go move the bags in and out. It's kind of ridiculous that they do this. Out of the 16-day duty cycle, they'll spend 14 full days just moving things. And so that's why we look at things very similar to the construction industry of that dump truck problem. of I've just moved this bucket of two tons of dirt.
36:26How do I actually control this? Because that changes your entire vehicle dynamics. And if you think about it in space, you know, you move your arms forward, your body moves backwards. And so this problem becomes magnified tenfold. And the only ways that we found to solve this is to use things like adaptive sliding controllers and use humans in the loop, but then also learn from that with RL and VLAs. Adaptive, what was the mechanism you used to move things around? Like adaptive sliding controllers. So controllers that will actually then change themselves based off of what you're picking up, based off of estimates that you actually have and in different environments.
37:00Got it. So highly adaptable based on whatever the task is. Got it. Okay. Jeff, you want to weigh in here on building world models for zero G? Because I feel like everyone's talking a lot of smack about not having enough data. This is really tough. And I'm curious when we're going to get Moon 1 from you guys. It's going to solve the problem. Hold my beer. I was going to say, last six months. And none of this is intractable. I mean, for example, the hierarchical planning that you are sort of both talking about, this is pretty common. I mean, like almost every autonomy stack has got looks of some kind of flavor of this at every point in history in the last decade, frankly.
37:36So the way I kind of see this evolving is you'll always have some kind of like system decisions which will be dependent on the application that might be safety related or something like that. There's probably going to be some kind of controller in the loop. And the question then is where do you want to devote the horsepower of AI in terms of key decision making? And there can be different abstractions. Essentially as the core foundation models get better to maybe use a language parallel, LLMs are amazing provided you can fit things in context. And essentially you hit limits as soon as things drop out of context.
38:04So that's getting better. It's almost like this rising tide that improves the core reasoning and planning capabilities of these models. World models will be very, very similar. It's a relatively early technology comparison. Imagine the algorithm march over a number of months. So what I would expect is basically a lot of the complexities of that to sort of collapse over time, which is what we saw play out in autonomous driving, saw play out in LLMs, the same thing will happen. So the question I think is how do we make a space world model? I think it's a combination of two things. Either a closing loop through a simulation environment where you can kind of close the loop through some like physical interaction manipulation environment with gravity set to zero.
38:46Great. That's something that any world more company can work with. Similarly, from a data perspective, any observational data or visual state and what happens and how this works from a visual perspective, also great. So I think it's just simply a matter of that and bring that into the the data mix, but there's nothing inherent beyond that that is intrinsically limiting. So yeah, if you want to send some data our way. Yeah, I was just about to say that. We like Jeff. So let's say, Ethan is going to go up to space to the ISS with Voyager Technologies, a public space company. They're going to take their, Ethan, correct me here, is it Joy or Joyride your robot?
39:20So yeah, the robot is named Joy and the mission is Joyride. Okay, I got this mixed up in the head. Okay, by the way, good branding. uh so joy goes up into into the columbus module or whatever and bops around doing its testing and so forth if they took like you know telemetry data sensor data you know video data as much as they could from from a reasonably long mission and sent it over to odyssey and said can you guys try to take this and improve a zero g world model for us how far do you think a mission's worth of data from icarus would get odyssey and building that how long is a mission like how many hours Yeah, we're up there for a full year.
39:57So Q127, Q128. So we'd have thousands of hours to ship to you. I mean, I was going to say like a thousand hours is probably sufficient in most cases. I mean, it's certainly maybe sort of a parallel. So we have some of our web models in use of robot foundation model applications. We're not a robot foundation model developer. Rather, we basically provide this as a model to people who then build on it. In this case, there's often some domain misalignment in terms of data, but it's like order of magnitude sort of in the sort of hundreds of thousands to maybe a low millions of samples usually gets us a pretty long way.
40:31Might be a little bit more in terms of zero G, but there's also quite a lot of data that exists in the public internet of things moving in zero G. So we're not starting from scratch in that regard. So certainly very happy to see we add some zero G to the mix. Yeah, we'll ship you an SD card or something. but I think it's going to be a little bigger than an SD card. I mean, dear God, if that's all the data you send over, I think Jeff would laugh. He'd be like, what is this data for ants? That's pathetic. I don't think I've ever worked with less than a petabyte of data in my life. Okay. But okay.
41:02This actually goes back to what Ethan was talking about earlier on about, you know, space distance and, and temporal lag in communications. If you wanted to get a petabyte of data down, Ethan, is that even possible with current throughput between us and Leo lower Yeah. So there's a bunch of different communication schemes and they all have different latencies and different bandwidths. So traditionally, to buck to the ISS, we use an S-band relay. You go from Earth to Geo, Geo down to Leo back to Geo, Geo back to Earth. Oh, you actually do the... Oh, I didn't know about the last loop there. Interesting.
41:33Yeah. And so this historically has very low bandwidth. Like, for example, on that S-band, we'll get about a gig a day pushing it. And so this really sucks for us. But if you look at the new generation of infrastructure going up there, so I'll, you know, pick one amount of hat, you know, a Voyager with Starlab or VASP with Haven. These have laser-based comps. So you get about 100 milliseconds latency in communication, and you get about 25 gigabits per second. And so this is a massive increase from what currently is state-of-the-art. And what's really exciting about what we're doing is, while we won't have the full bandwidth while we're up there on the stay station, because it's like a 30-year-old Toyota Corolla at this point, what we do get is we get all the data actually back down because we're getting the robot back.
42:20So it's coming down on a Dragon capsule. And so we'll actually be able to physically bring it down and actually process the data afterwards. What's the backup if it doesn't? Yeah, yeah. And then we have the one gig a day over the course of the year. So that's the backup. Okay. Yeah, that's really the backup. So you're going to lose a lot of data. But, you know, let's just hope that that dragon capsule comes down and, you know, the astronauts don't die as well. Oh, well, now I feel like an asshole. I guess there would be humans in there too. I was just thinking about your one pathetic little SD card getting too hot and burning up.
42:53Boris, we're talking a lot about, you know, where the rubber meets the road here, or I guess where the robot meets the zero G, but you're definitely out there in the real world. I was really impressed that your company recently moved 65 ,000 cubic yards of dirt on one project alone. So are you now just like in data abundance as a company? Because I feel like if you've done that much actual earth moving after all the training you've talked about, all the data, all the sampling, all the modeling, you must just have a ton of fresh and useful data. It's accumulating. It's how everything's relative where I think there's always, particularly given how many types of machines and types of work are out there, there's always a frontier of like new diversity, you know, new variety and so forth.
43:33it's funny as Ethan was talking I'm like thinking you know the data problem like everybody has a version of it where you know we generate a terabyte per machine per day and so it's a thousand gigabytes and so it's like we're like yeah we're like you know not quite doable with Starlink and so we're still kind of shuffling physical you know data around and then there's you know kind of like work your way trying to hope that you know you can't compress your way into something that it could be purely through satellite. Yeah. So for us, we have a really helpful property that we are operating an existing heavy machinery that we can retrofit.
44:11And so we do not need to create a brand new machine that is obviously expensive, complex, and also not in existence, which makes data harder. The fact that we can go and take existing$600 ,000 Caterpillar machine and turn it into autonomously capable, whether it's for data collection or for autonomy testing, that opens up a lot more flexibility in how we get the data. And so we are able to have really rich partnerships with general contractors and subcontractors, both for data, but also to go in our actual deployments with them, like what you mentioned, where we were helping build a factory in Arkansas.
44:50in this case uh we've been on a number of data centers i've done a manufacturing facility uh a couple warehouses uh and we're not far off from going um those were supervised autonomy tests where we were doing autonomy but with a safety operator and um later this year we'll be going full drive over since so that means that these will be operating on real sites doing work um all day long with nobody in the cab um and in the initial use cases will be mass excavation and then moving on to more and more types of work. But what's exciting is that some of these projects, you know, we're now doing projects that will be many hundreds of thousands of cubic yards of earth.
45:30And these are subsets of projects that get into the many millions of cubic yards of earth that become the opportunity. And so when you think of some of these data centers, for example, they're getting built, they'll be getting built on many hundreds, if not thousands of acres, and can be, you know, two, five, even 10 million cubic yards of earth. And so these are the scale of projects that are actually happening. That's almost as much dirt as Peter Thiel dug up to build his bunker in New Zealand.
45:57No, jokes aside, 65 ,000, sorry, it's not fair to make you guys enjoy my humor. 65 ,000 cubic meters is exactly 26 Olympic sized swimming pools. So that's how much dirt you guys moved in case you wanted a more prosaic term. All right. So listen, all this is great. I'm excited about everything from world models to saving time construction sites, building things faster, going up and learning. It's great. It's so great. It's optimistic. It's very human positive. Why does this not form more of a conversation we have at the public level? And what do we need to do as people who are really bullish on having more intelligent machines in the world and better models behind them change the conversation?
46:38I'm so bummed to read the news and read what people are saying, read the polling about data centers and everyone being very concerned, and then coming here and having a blast and just... There's an optimism, right? Like on what's possible with a lot of this work. Maybe I'll kick off. Like it does worry me because when you see approval rating for AI, like below 20%, that puts it in danger of being a bipartisan issue that like actually, you know, like impacts the space in a way that is detrimental, not just to the space, but to like, you know, overall society. And so I think it's, there's not enough, you know, the communication around it has certainly not been great, but there's so many successes in, you know, the breakthroughs you can start having in medicine, the way this democratizes education, like the way that, you know, today in construction, there's like a astronomical shortage because it's been systematically underinvested.
47:37And now when the demand is skyrocketing because of infrastructure and data centers and non-train manufacturing, you start seeing housing prices go up and so forth. And you have a world that is like fundamentally can be expansionary. And everybody does have a legitimate fear of local job disruption, but almost always it's much easier to see the local direct impacts where something is lost than all of the astronomic positives that have happened in every single technological transformation that are actually harder to predict. Where you see the mobile economies that started to form after mobile phones, what the internet created, even if it hurt retail stores locally in some places.
48:19Right. And like the biggest example that I can remember that kind of hits on the physical transformations that we're talking about is the Industrial Revolution, where it has that sort of a feel today where when it happened, everybody is petrified that all the jobs would be lost or be complete unemployment. Everything would be absolutely horrible. And, you know, just a few people, they make money off of it. And what ended up happening is that even though there was like locally some really painful kind of transitions that might be geographically or in profession fixated, by the time that dust settled, the productivity skyrocketed by like order to magnitude, you know, plus the number of jobs actually increased, not decreased, and the average salary doubled and the rate of poverty significantly decreased in the entire country.
49:02And so you have these like really, really non-zero-sum impacts that I think we're not doing a good job of conveying externally. Jeff, there's the big sovereign AI pusher over in the UK. I know there's talk about building kind of a UK-specific AI model. I know you're going through some prime minister turnover, but it did seem that the Keir Starmer government was at least generally in favor of building out AI. So are y 'all doing a better job with the pitch here, the optimism, the forward-looking element of this, the we are going to have better medicines and better technologies and cheaper goods?
49:32Yeah, maybe the context where we're partly transatlantic company with us, half the team here in London and half the team in Palo Alto. I think AI is ultimately a global problem and a global question. And the reality is companies who are pursuing this, particularly at the foundational scale, you have to think about winning this globally. And I think it's a good thing. As Boris said, it's not a zero sum game. In my mind, it's more about focusing on how do we make this a positive sum game. And I think there are so many things where, like many technologies, in many respects, I think all of us here are building something that is not possible today and the world will be better at the outcome.
50:08And I think in many respects, that's really why, certainly what motivates me and I think certainly what motivates most folks in the AI field. So as to whether we're doing a better job, I'm not entirely convinced. I think every society is grappling with this in terms of what it means, how the world will work, how relationships between countries will work. It's a bit of an unknown, but I don't think that unknown is to be feared. I think in many respects, It's a thrill of discovery and that motivates people to do ridiculous things like go into space. Yeah, I think something to be celebrated. I wonder, Ethan, if we just got everyone to read 10 times as much science fiction, we could solve this branding problem.
50:45Because if you pick up a copy of Red Mars or just even something that's older, but just aspirational about going out into space, like or Rendezvous with Rama, just pick another classic. Like, to me, those got me so excited about the possibilities of going faster and further. Your company being a good, I would say, step towards our front porch as a species before we go off even further. How do we get folks to be a little bit more open-minded and optimistic, especially amongst the younger set that are currently dealing with the butt end of AI today, which is a slightly softer job market for recent college graduates?
51:16You know, it's two things for us. Number one, we plan the end of physical AI in the AI field. But then also we work in the space industry and there's nothing more expensive looking than the space industry. When a rocket goes up, it looks expensive. It looks like someone took a pile of cash and lit it on fire. And people are like, what does this do for me? If you're New Glenn, that's true. You know, and like the thing about like programs like New Glenn is there's so much of the positive end where NASA and space have very bad PR problem. They don't like to take credit and they should take a credit for a lot more things than we have day to day.
51:54And we really have to ground what we do on how it impacts people's day to day lives. And I think that's the thing that every time I've sat down and had a conversation of why it's so important to do what we do, then this is what lands with people. You know, like, for example, you know, your cell phone. The reason why we have miniaturized computers was because the Apollo program. We needed them to put on the rocket to take us to the moon. The reason why you have memory phone, that's NASA. It created Invisalign. It created from a NASA program. Invisalign? The things for your teeth? Invisalign. That was a spinoff of technology that was made by the NASA program and NASA funding.
52:30Things like your credit card. Every time you swipe it, that timing signal, that comes from a satellite. Your GPS, when you came into work, you would not have that today without a satellite, without the space industry. And I think the one for us is we have the most advanced lab in the world floating around us with people working on it every single day for the last 25 years. And they created this cancer therapeutic called Keytruda. It's the number one cancer therapeutic in the world right now. And so we did the research and development of this drug where you can do protein crystallization on orbit.
53:05And we found out the way that proteins crystallize in space, totally different than Earth, way more effective. We changed the way we manufacture this drug right now on Earth with a random positioning machine to try to replicate zero G because we can't. And that one drug made $29 billion last year in revenue for Merck, the pharmaceutical company, and has saved hundreds of millions of lives. There you go. That's the one you want. I think we need to drop the financial one, even though we're all capitalists here, and just go straight to the petted kittens and planted trees. Because I think that there's so much money going around right now that every time we bring it up, it doesn't land well.
53:39The joke that I made at the start about how Icarus hasn't raised a nine-figure Series B yet is a joke amongst us because we know that the markets and problems are so big that it's worthy of investment. But to an average person, they're like, my school fell down. Why are you giving Jeff$310 million? He's got nice hair, but does he really need all of that? Couldn't he get by with$35 million? You should see how GPU bill. Maybe we need to definancialize things a little bit. I don't know. I don't want to live in a world where we have the shot at such rapid progress and a brighter future, and we end up talking ourselves down, and dumbing ourselves down to the point in which we slow down.
54:17And I don't know if we're going to lose to China per se. I think it was Jeff talking about competition, but I don't even want that to be a possibility as a fan of democracy. So do we need a new leader in technology? Have we just failed to find the right person to carry this torch? I think in respect to the revenge of the nerds, right, is a good thing. In respect, I think a lot of what sort of built the technology industry is people being like, oh my God, I can build something ridiculous of my garage and it's really, really cool. It enables amazing new things. And I think it's often being sort of couched in large business interests, which are true.
54:54I mean, respect that's a good thing. It brings the capital behind the problem, but at the same time, I think that soul is valuable. And I think sometimes being a bit more open and human with the public in that regard can be useful. Yeah. Boris, do you think we should have partial AI lab nationalization. We're seeing the horseshoe theory in practice here. We got Bernie Sanders and Donald Trump both coming to the same conclusion. I don't think it's a particularly good idea, but I'm happy to be wrong. What's your take? That kind of scares me as well. I'm with you. There's not a, there's a lot of, not a huge amount of examples that's actually worked out well in other industries.
55:32At the very, very least, it massively slows down innovation if you now have that gate across. Now, it doesn't mean that they're, you know, like given the power of some of these models, you don't have, you know, some checks on how and, you know, when they, the sort of ways that they can be applied. There's export controls on GPUs. Like there's strategic reasons on why that might make sense. Right. But it kind of scares me. It's almost like saying we're going to government control software. Right. It's such a. Math in this case. Right, math. It's like, okay, well, there was proposals to control models in general.
56:15It's like, well, at what point does it become a large model? What if you're one parameter under? It's kind of silly in the big picture. And it's trying to kind of apply a hammer to something that is just operating on a very different dimension. So it worries me. But at the same time, I think it's probably not doing the industry. is doing the industry a disservice to bulldoze forward and put your head in the sand and not actually try to communicate and kind of create a bridge to government, to the rest of society where people do genuinely feel like they're not benefiting equally from it, even if a lot of their benefits are actually going to come in five years when like all these technologies mature.
56:54I think the only 100 % losing strategy when it comes to selling AI to the public is to be dismissive and rude to people who have concerns because I think what that does is just poisons the conversation. I'm not going to play it because I want to save a question for actually for you again, Boris. But I saw a video of Theo Vaughn, the podcaster who became very well known during the last election. And he was talking about data centers. And he was talking what I would say is a standard TikTok set of points about data centers. And everyone was just dunking on them. Instead of trying to politely educate him, help him out, be supportive, they were just calling him an idiot, calling him various slurs.
57:27And I was like, that's not at all how we're going to build the coalition here to keep this from becoming a bipartisan issue that's going to get crushed. So I'm hoping for some more positivity. All right. I want to do one more thing before we go and talk about the return of fable for fun as a final question. NVIDIA recently announced Cosmos 3, which is a quote, open frontier model for physical AI and the world's first omni modal unifying vision reasoning world action and generation. So Jeff, are they encroaching on your core domain here? Or is this something that's more like a test bed for other folks to learn from?
58:01Overall, I think it's positive for us. In many respects, they're another voice advocating for this as a category. NVIDIA historically has done this, where they use their research arm as a way to basically prove markets and validate markets. And essentially, as soon as the market is mature, they typically go and find the research tiers another way. And this has been played out quite a number of times. So yeah, in many respects, there are things that are comparable to our models. There are things that are different in terms of our models. There's some great researchers on it, but I'm quietly confident in our models.
58:34Yeah. Well, I was just thinking that Nematron, the models from NVIDIA, are a good example of open-weight AI from the United States, which we could see a lot more of, and I think we wouldn't be any poorer. Now, Boris, there was also something from NVIDIA called Halos for Robotics, which is full-stack open robotics safety system. Is this something that you can then bring to your robots, or is this something you've already figured out, and they're just helping other people avoid the, not the dump truck problem, but the swinging bucket problem? And we were constantly kind of like surveying and kind of trying out these tools.
59:07When you get into these incredibly complex systems, we've at least found that you have to more deeply vertically integrate and really control the full stack because you have to deeply understand the interface with the systems, the latencies, the physics and controls, the cycle counts, the performance of every model. you have to have the architecture that is guaranteed to fail safe. And so we're constantly looking at these types of tools as a support for various building blocks. But I think given that there's still so much novelty to these large kind of physical systems and things are just not standardized and how does a car work versus an excavator versus a humanoid machine, there's no magic bullet.
59:55you kind of tend to have to still have a pretty deep control all the way through the stack. I think it's an understatement. Every single robot financial model company, including autonomous driving that I've spoken to, has a uniquely different view in sample rate, frame rate, in terms of the action rate, in terms of the way they've parameterized trajectories, including resolution, the number of cameras. Do people like diffusion? Do people not like diffusion? It's like, what's safety critical? What's not safety critical? What's your safety case? It's almost like you can't miss their potato head.
1:00:22It just never really works. When you talk to them, they're all so confident that they're correct and everyone else is wrong. And I'm always like, guys, you can't all be, maybe you can all be right, but you can't all be wrong. I don't know what to tell these very smart people. This is natural. Like, I mean, you tend to see this, right, where everyone kind of fans out and you kind of, imagine you're kind of like throwing darts at a big dartboard and you're trying to get points. And everyone's sort of exploring. And over time, basically, things converge. I think autonomous driving is sort of a little bit ahead of the curve compared to some of the other fields, just because it's a bit more of a mature industry, a little bit more defined problem at some level.
1:00:53The objectives will converge over time. It's just a question of enough people getting enough information. You end up with cross-pollination between companies. This will start to consolidate. In methodology, I mean. Okay. And so do you think once we get the methodology kind of aligned or harmonized, do you think we'll be able to go faster or will that actually slow us down because we'll be exploring fewer odds and ends and nooks and crannies in the research game? It'll make my life easier. I won't get such a wide range of feature requests. Jeff, I'm not actually optimizing for your personal retirement.
1:01:25I'm not going to lie. Typically, people consolidate around things which work better, and typically something that works better in adjacent domains tends to work well in another. So maybe things which work well in autonomous driving will probably have, in many cases, will work well in maybe construction. Although maybe that's a little bit different given you also have got manipulation of your environment at the same time. So there are adjacencies, and typically where there are adjacencies, you accelerate. I'll take a slightly different lens on this one is that this focuses on the autonomy model and the core approach of development.
1:01:59What's deceptive is how much of the challenge of solving, for example, something like autonomous driving is actually not at all in the autonomy model. But how do you evaluate how you able to improve your system when you're in a super long tail, one in millions of mile case without regressing accidentally in 15 other places? How do you actually qualify it? And so there's actually so much surface area that in a lot of ways, people have gotten quite good at climbing quality of models when you have the right data and so forth. And the hardest problem eventually becomes how do you get the right signal that you actually can drive 500 developers to optimize over?
1:02:35Or how do you actually test it and make sure that it's qualified and ready for driverless when you have safety critical applications? It's actually harder in your case with bedrock because if you think about like modern equipment, it's often hydraulically driven, right? If you're doing a big digger. And you know what's not super precise? Hydraulic driven, large buckets full of dirt. And so earlier in the conversation, you mentioned something like precision and like getting within an inch or something. I've seen these machines. They're not always very well maintained. So you probably have to also take into account like you might tell it to do something and you're going to get a slightly different outcome and have to handle that as well in the real world.
1:03:11So that's exactly right. Because you now have a variance where you do not like, you know, Waymo has thousands of jaguars on the road. They're tuned to be like as similar as you possibly can be. You know, you have like 72 different types of attachments to an excavator. You have tons of different sizes from, you know, like in terms of tonnage and, you know, and variances as well. And so what you have to do is create safety cases that are actually resilient to this and work at everything from subsystem to kind of overall system level. And so, for example, you can have controls systems, imposed systems that estimate and, you know, all these models.
1:03:45And you can like empirically kind of like measure your errors and your variances and collect data incredibly quickly over a wide range of machines and have very precise tolerances. And be confident that you have a particular type of margin of error that's going to be your, you know, like some giant percentage of time you'll be within this sort of tolerance. And then part of this is knowing that you will have exceptions where, for example, let's say a sensor broke or a machine is just like fundamentally different and you fall outside their range. You have to have fail safes to where it's always preferable to fail safely than to try to do something that's outside of your space you've qualified.
1:04:23Ethan, it's funny how much you're the opposite problem to this because you're building one robot. You're controlling all the sensors, all the tools, all the pieces. You don't have to go out in the field with, you know, someone didn't maintain this digger, so it's a complete mess. How precise do you need to be when it comes to your machine interacting with the physical world? Are there still tolerances you can play with or is it like millimeter precise requirements? No, you still have to be within those centimeters, like especially because the environment that we're in is working hand in hand with humans.
1:04:52And so where you can have very clear workstates with something like large hardware that's on a construction site, make sure there are not many people around. So even if you do go into whatever failure state you're in, you can be safe. For us, if we go into the wrong failure state next to somebody and you hit them, it's not like there's a doctor up there. Everyone's trained in medical, but that can be a very bad situation. And so when we talk about the systems that we're interacting with, from the switches to the science experiments to even the cargo bags, these are made for human level fidelity.
1:05:24You're very good at grabbing very small things you have tabs and zippers that are sub centimeter size that we need to grab ourselves and if you take like robotic grip and you kind of break it down you get two things you get pinch you get power and so the way that we've made our end effectors are actually to replicate that so it's actually just two small fingers because of how precise we need to be where we're flipping switches uncertain experiments to turn them on and plugging in mil spec connectors um and for us like the the two things that make it hard is number one every single action in space the reaction is just so exaggerated on earth you have 1g dampening every single movement now for us we just don't have that you move your arms forward to the side your entire body starts rolling your cameras start rolling um and then when you pick up things sometimes we're picking up things that are more massive than our vehicle itself And so that shifts outside of your control matrix of your thrusters, that shifts your center of mass way outside of this.
1:06:25So traditionally, if you have your center of mass inside of this and you try to strafe left on the X axis and just move, you'll start pivoting in a circle if you've sufficiently moved your center of mass far forward enough. And so these are the types of things that we have to really care about to get our precision fine enough that you can do the tasks that matter. You know, I find it really funny that the reason why a lot of people want to build humanoid robots and they've told me this is because you can take a robot then and put it into a human environment and it can perform well. And they often do this in places where you have lots of room and could actually change the definition of work and create more of a purpose-built robot.
1:06:57You're going into a literally human defined space that's very small, very complicated, and you're not going humanoid. So I wonder if everyone else kind of adds it backwards because you're making, it looks like a robot, if I imagine like, you know, what are the dimensions on joy? Like a pizza box by like four pizza boxes? Yeah, that's actually like perfect. I've never thought of explaining it that way, but that's probably a great, great visual. You're welcome. But it's that size and it's supposed to interact with human things quite well. So Ethan, are people just, is figure at all making the wrong bet when it comes to what they're actually building to do the work?
1:07:34So I think it really depends on the environment and what you're trying to do. Like humanoids, the reason why these are compelling is economies of scale. If you can get something that does every single task on earth made for a human at 80 percent and you can just print these it makes sense on the financial side of things you'll never beat a purpose-built machine or robot for some task with a humanoid it just won't happen but that's not the point of humanoids the point is i can make a hundred thousand and deploy them everywhere and democratize labor and when you look at the earth and you say okay half of the earth's gdp is people doing things it's human labor, there's a very compelling argument there.
1:08:10When we look at space, we say, oh my God, I'm going to go look at these astronauts. I'm going to go talk to these astronauts. What do they actually do in this environment? Turns out legs suck. They don't help at all. They mess around your center mass. The only things you ever use your legs for is hooking into footholds so you can react against the vehicle itself so you're not floating and spinning in a circle. And so when you break down the environment, your first principles, a human has been evolved to work on Earth very, very well. but not space. We sent them there. It kills them. Your heart degrades.
1:08:42Your eyes go bad. Your bones get soft. Exactly. And so when we broke it down to first principles of what was most important, it takes shape as something very, very different. But the machine itself and the robot itself is still a general purpose platform. And so I think there's a very big difference between humanoids and general purpose platforms depending on the environments. You don't want a humanoid to go swim underwater you want something that's an underwater rov or similar class um the same way that you don't want a human to fly in the air but you might want a human to walk on earth so i think depending on your domain you get very different general purpose machines all right listen i could keep asking you guys questions literally all day long but it is getting a little bit late in london so we're gonna we're gonna wrap it up let me do this wrap towards a final a final quick round of questions here.
1:09:30So for each one of you, one, when does Fable come back? I'm curious. And then two, are we still approaching the GPT-3 moment for physical AI or have we already gone past it? And Jeff, we've heard the least from you lately. So we'll start with you. Ooh, when are we getting Fable back? Soon, I hope, would be my thought. I have no unique insight into when they might come back. I think that's probably out of my hands. You have to put a date on it. You can't waffle here. No, I don't know. this one okay uh pretend you're from texas two weeks okay good uh in terms of uh gpt3 moment for her physical ai i don't think we're there i think in my mind the gpt3 moment is when you see the spike of a sort of explosive rollout in demand um i think we're seeing promising signs but i don't think we're there i think we're on the trajectory all right boris same questions for you oh let's see when does it come back uh uh well and not in not in a form where they kind of like nerf the capability, but for real.
1:10:29And I'd say it sounds like it's a little bit longer because it got a little bit of a mess with, like, kind of, you know, government, but now they've got to, you know, to go with it. So, I don't know, I would say eight weeks is my guess. I don't know. Eight months. I guess we're first eight weeks. Yeah, me too. I don't know, I'm a little less optimistic. Sounds messy. You know, like Jeff, I actually think we're not there. I'm maybe a little bit more pessimistic on the GPT-3 moment anytime soon just because that benefited so much from almost infinite data in a common feature space of language. There's so much nuance and variability to the hardware, the sensors, the environments, the safety challenges, everything, that I think you will for longer see these more verticalized applications really focus and take things end to end.
1:11:13Even though the industries can be double-digit percentages of GDP, I think it'll be a while longer before you just have something that can genuinely go into brand new environments and be super capable like this. But likewise, I don't think it's impossible. I think we're on the way. It's just going to take longer. Okay. No, I really appreciate the context there. Ethan, you get the last couple of words here. So when does Fable come back? And the GPT-3 moment for robotics, future or past? Yeah. So when it comes to, you know, like Mythos and Fable, I don't think we're getting it anytime soon. I'd put it, you know, two and a half, three months.
1:11:45The sea getting worse. I'd put me in the fourth guess. You'd be like, five years. No, I mean, the government moves at the speed of government. I think once you classify something as a weapon and you put ITAR controls on it, this just makes everything so much more messy. You know, we've had our dealings with them. Just getting documents signed takes weeks. So I can imagine that this will take a very long time. And then, like, as far as GPT moment for physical AI, I think we're so early. I think it's a far, far time away. And I think like Forrest said, you're going to find very vertical niches and very, very constrained new sets right now.
1:12:23And then once you see that demand signal spike and once you see the capability spike is when it will explode. And so I think that's what makes it exciting for me is because we are on the, you know, the greenfield bleeding edge of it right now. And I think like two, three years time, we're going to see that massive explosion. Well, it's really optimistic because I hate doing work and I hate driving and I like to write. So if anyone can take care of all the physical stuff, I'll do all the typing and life will be good. All right, guys, we got to leave it there. This has been This Week in AI. My name is Alex.
1:12:51And today we are joined by my new besties, Boris from Bedrock, Jeff from Odyssey and Ethan from Icarus Robotics. Thank you all much. We'll have you back and we'll see everyone next week. Goodbye.
From the publisher
This Week In Startups is made possible by:
Paypal
Today’s show:
*When we think of “Physical AI,” most of us conjure images of humanoid robots and self-driving cars. But there’s a lot more to the category than just breakdancing Chinese bots and Waymos, from excavators autonomously digging tomorrows worksites, to a robot the size of four pizza boxes deploying to the ISS.
Guest host Alex Wilhelm sits down with three founders in the trenches of the Physical AI Space:
Boris Sofman of Bedrock Robotics, which builds autonomous construction equipment
Jeff Hawke of Odyssey, a frontier AI lab specializing in world models for robotics and video games
Ethan Barajas of Icarus Robotics, designers of the free-flying Joy robot that will join the ISS crew in 2027
Guests:
Boris Sofman on X: https://x.com/bsofman
Bedrock Robotics: https://bedrockrobotics.com/
Jeff Hawke on X: https://x.com/jeffrey_hawke
Odyssey: https://odyssey.ml/
Ethan Barajas on X: https://x.com/ethanbarajas11
Icarus Robotics: https://www.icarusrobotics.com/
Relevant Links
Bedrock Excavators Remove 65,000 Cubic Yards of Dirt: https://www.enr.com/articles/61982-bedrock-robotics-excavators-remove-65-000-cubic-yards-of-dirt-on-southwest-project
Odyssey $310M Series B article: https://techcrunch.com/2026/06/17/world-model-maker-odyssey-nabs-1-45b-valuation-backed-by-amazon-and-other-big-names/
Icarus “Joyride” Mission announcement: https://thedebrief.org/icarus-is-building-the-robotic-labor-force-for-space-voyager-technologies-is-sending-a-next-generation-zero-gravity-robot-on-a-joyride-to-space/
PROWL: Prioritized Regret-Driven Optimization for World Model Learning: https://arxiv.org/abs/2605.18803
NVIDIA Cosmos 3 report: https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf
ISS Columbus Laboratory Module: https://www.nasa.gov/international-space-station/columbus-laboratory-module/
NASA’s Astrobee flying robot: https://www.nasa.gov/astrobee/
Voyager Technologies: https://www.voyagerspace.com/
Vast Space: https://www.vastspace.com/
Blue Origin’s New Glenn: https://www.blueorigin.com/new-glenn
“Red Mars” by Kim Stanley Robinson: https://www.amazon.com/Red-Mars-Stanley-Robinson/dp/0553560735
“Rendezvous with Rama” by Arthur C. Clarke: https://www.amazon.com/Rendezvous-Rama-Arthur-Clarke/dp/0553287893
Theo Von “This Past Weekend” podcast: https://www.theovon.com/podcast
Anthropic statement on Fable/Mythos suspension: https://www.anthropic.com/news/fable-mythos-access
Timestamps:
0:00 The Year of Physical AI
1:28 Why everyone's bullish on world models
7:03 Physical AI is harder, but probably bigger
9:08 How to think about world models
16:00 When data becomes "adversarial"
25:00 How Bedrock powers autonomous excavators
30:48 Why teleoperation gets a bad rap
39:10 Meet Joy the space robot
41:33 Laser comms coming to next-gen stations
46:29 Why everyone hates data centers
57:33 When does Claude Fable come back?
Subscribe to the TWiST500 newsletter: https://ticker.thisweekinstartups.com
Check out the TWIST500: https://www.twist500.com
Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp
Follow Lon:
Follow Alex:
LinkedIn: https://www.linkedin.com/in/alexwilhelm
Follow Jason:
LinkedIn: https://www.linkedin.com/in/jasoncalacanis
Check out all our partner offers: https://partners.launch.co/
Great TWIST interviews: Will Guidara, Eoghan McCabe, Steve Huffman, Brian Chesky, Bob Moesta, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland
Check out Jason’s suite of newsletters: https://substack.com/@calacanis
Follow TWiST:
Twitter: https://twitter.com/TWiStartups
YouTube: https://www.youtube.com/thisweekin
Instagram: https://www.instagram.com/thisweekinstartups
TikTok: https://www.tiktok.com/@thisweekinstartups
Substack: https://twistartups.substack.com
