In short
Podcast Episode Notes: The 20-Year Journey to Fully Autonomous Cars with Dmitri Dolgov of Waymo
Podcast Overview
- Title: Cheeky Pint
- Host: John Collison
- Guest: Dmitri Dolgov, Co-CEO of Waymo
- Episode Focus: The evolution of self-driving cars and Waymo's journey to achieving fully autonomous rides.
Key Highlights
- Waymo's Scale: Waymo provides over 500,000 autonomous rides weekly across 11 cities.
- Foundational Background: Dmitri Dolgov joined Google's self-driving car initiative in 2009 and became co-CEO in 2021.
- Operational Insight: Discussion on the transition from scientific research to practical application and the operational aspects of Waymo's services.
Episode Breakdown Background & Early Life
- Dmitri's Journey:
- Grew up in Russia; family moved during the Soviet Union's collapse.
- Passion for physics and math influenced his educational choices.
- Notable mention of the Russian diaspora's impact on tech in the UK.
Technical Discussion
- Waymo's Architecture:
- Utilizes a combination of sensors: cameras, LiDAR, and radar for 360-degree environmental perception.
- Data processing done in real-time using AI encoders and decoders.
- AI plays a crucial role in decision-making for vehicle operations.
- Teacher and Critic Models:
- Waymo employs "Teacher" (off-board foundation models) and "Critic" models to enhance AI training.
- The integration of multiple models allows for predictive behavior and decision-making in complex environments.
- System Design Philosophy:
- Emphasis on a modular architecture rather than a purely end-to-end approach to enhance safety and efficiency.
- Importance of real-world data and feedback loops in training AI models effectively.
Evolution of Self-Driving Technology
- Iterative Progress:
- Technological advancements (AI, compute power) have been essential for scaling up.
- Early models required significant adjustments and improvements based on real-world testing.
- Driving Challenges:
- Addressing nuanced driving situations, especially in urban environments with high pedestrian interaction.
- Differentiation between freeway driving and complex urban drop-off scenarios.
Future of Autonomous Vehicles
- Scaling Operations:
- Transition from developmental to operational phases, focusing on data collection, validation, and expansion into new cities.
- Plans for international operations in cities like London and Tokyo.
- Hardware Innovations:
- Discussion on the development of a custom-built vehicle aimed at enhancing passenger experience.
- Introduction of a new vehicle platform with features designed for comfort and accessibility.
- Economic Considerations:
- Examination of ride-hailing economics, particularly in low-density areas like rural Alaska.
Cultural and Social Impacts
- Societal Changes:
- Predictions about how self-driving technology will reshape urban landscapes, parking needs, and social driving behaviors.
- The potential for higher throughput traffic systems and reduced congestion as autonomous vehicles become more prevalent.
- Public Perception and Acceptance:
- Addressing public concerns about safety and operational reliability.
- Importance of creating a positive rider experience to foster acceptance.
Conclusion
- Waymo's Vision: Aiming for seamless, autonomous transportation in major metropolitan areas with a focus on safety, efficiency, and user experience.
- Long-term Goals: Dmitri envisions a future where autonomous vehicles are ubiquitous, fundamentally changing how people interact with urban environments and transportation systems.
Additional Resources
- Waymo Research Article: [EMMA: End-to-End Multimodal Model for Autonomous Driving](https://waymo.com/research/emma/)
---
This markdown file captures the essential discussions and insights from the episode featuring Dmitri Dolgov, providing a clear and structured summary for readers interested in the evolution of self-driving technology and Waymo's pivotal role in it.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VODmitri's Journey from Russia to California
0:45 to 2:41
Dmitri shares his early life in Russia and his family's migration to the U.S.
“I was thinking about the next thing I wanted to do, and I really like that technical school in Russia.”
The Technical Foundations of Waymo
2:41 to 4:35
Dmitri describes the architecture and sensor modalities used in Waymo's self-driving cars.
“You know the classic engineering interview question of what happens when I type google.com and hit enter?”
Real-Time Processing and Cloud Capabilities
4:35 to 6:26
Discussion on how the Waymo driver processes data in real-time and its cloud features.
“Nowadays, we talk about it using AI terminology as encoders that take this data in.”
Debating Self-Driving Technologies
6:26 to 8:11
Dmitri discusses current debates in self-driving technology and their relevance.
“So I can think of end-to-end versus the more modular approach.”
20 Years of Evolution in Self-Driving Tech
8:11 to 10:12
Dmitri reflects on the technological evolution of self-driving cars over two decades.
“The simulator, of course, is what powers our synthetic generative environment that can run on the cloud for training and for evaluation in close level of the system.”
Building a Foundation for Autonomous Driving
10:12 to 14:01
Exploration of the foundational models and architecture behind Waymo's technology.
“Was it just a story of we needed 20 years of going down the wrong cul-de-sacs before we eventually arrived at the right approach?”
Driving Dynamics: Lessons from LLMs
14:01 to 16:41
Explore how autonomous driving models can learn from language models and the challenges they face.
“The easiest thing you can do is pick a model.”
The Role of Context in Driving
16:42 to 19:38
Understanding the multi-agent interactions that make driving complex and how they relate to language and context.
“And if all you did was kind of observing how other people drive when you trained the system.”
Nuanced Challenges of Drop-offs
19:39 to 22:28
Discuss the intricacies of drop-off and pick-up for self-driving cars and the importance of context.
“having N10 architecture that's augmented with that structure, allows you to kind of play in both of those worlds.”
Scaling Autonomous Driving Technology
22:29 to 24:24
Examine the advancements in self-driving technology and the transition from research to deployment.
“Freeways, for most of the time, Not much happens.”
Show all 26 chapters
Operating Domains in Self-Driving
24:25 to 27:08
Learn about how Waymo evaluates its technology across different environments and scenarios.
“Now, that said, there is a lot of work to do in specialization and in validation before we can deploy responsibly.”
Evolution of Waymo's Driver Technology
27:09 to 28:03
Discover the progression of Waymo's driver technology and its impact on deployment strategies.
“be within the operating domain or outside of it.”
Advancements in Autonomous Driving Generations
28:03 to 29:18
Explore the evolution from the fourth to fifth generation of autonomous driving systems and the role of AI.
“this is what's on the high basis today, we really wanted to take a huge bite out of that operating domain.”
The Future of Custom Autonomous Vehicles
30:17 to 33:19
Discuss the transition from conventional vehicles to specialized self-driving designs.
“But one is maybe everyone in this space has a very charismatic demo of a vehicle that is custom made for self-driving.”
The Sixth Generation of Waymo Vehicles
33:20 to 35:28
Understand the features and improvements of Waymo's sixth generation vehicle and sensor stack.
“But if you think about the value proposition, right?”
Sensor Technology in Self-Driving Cars
35:29 to 37:36
Examine the differences between LIDAR and radar in autonomous vehicles and their applications.
“It is very different from the fifth generation.”
Challenges and Excitement in Autonomous Driving
37:37 to 42:00
Learn about the challenges facing autonomous driving and the excitement of rapid technological advancements.
“What are LIDARs versus radars better at in a self-driving company?”
Emergent Behavior in Autonomous Vehicles
42:00 to 45:30
Learn about the surprising capabilities of autonomous vehicles and the technology behind them.
“I think it's seeing the capability and the scaling laws from this approach of starting with that cornerstone of the foundational model and then specializing to t-shirts and then distilling.”
Waymo's Current Operations and Expansion
45:30 to 47:45
Understand Waymo's operational metrics, including rides, revenue, and city expansions.
“Is this the perfect example to explain what we were talking about earlier?”
The Convergence of Driver Assistance and Autonomy
47:45 to 49:54
Explore the relationship between driver assistance systems and full autonomy in vehicles.
“It seems now clear that in 15 years, most miles that are driven will be autonomous.”
Operational Infrastructure Behind Waymo
49:54 to 53:44
Gain insights into the operational infrastructure supporting Waymo's self-driving cars.
“You have to tackle, if I think about the hardest parts of building a fully autonomous rider-only system, they are very different from what you do for a driver assist system.”
Future of Waymo and Autonomous Traffic
53:44 to 55:55
Discuss the future implications of widespread autonomous traffic and its impact on society.
“How well-behaved are the Waymo riding population in terms of not living a mess in the car?”
Impact of Autonomous Driving on Traffic and Urban Planning
56:01 to 57:08
Explore how autonomous driving could change traffic dynamics and urban landscapes.
“But presumably there'll be other second order effects, like we'll want higher throughput traffic lights.”
Waymo's Journey and Google's Commitment to Self-Driving Tech
57:09 to 58:44
Understand Waymo's development journey and Google's long-term vision for self-driving.
“Well, because, again, your car is just sitting there 90 % of the time, right?”
Challenges and Breakthroughs in Self-Driving Technology
58:45 to 1:01:00
Learn about the technological challenges faced and breakthroughs achieved in self-driving.
“Yeah, no, on the latter part, I just have to give credit and huge kudos and gratitude to Larry and Sergey and Alphabet Leadership Center company.”
Google's Culture of Innovation and Talent Development
1:01:01 to 1:02:15
Discover how Google's culture fosters innovation and enhances technical talent.
“It feels like Google really recognized your talents.”
Transcript
Automatic transcript. May contain errors.0:01Dmitri Dolgov:Dmitri Dolgov is co-CEO of Waymo. He joined Google's self-driving car project in 2009 as one of its first engineers and was repeatedly promoted until he took it over in 2021. Waymo is Google's most successful moonshot and now provides over 500 ,000 fully autonomous rides each week. Cheers, by the way. Yeah, cheers. You grew up in Russia, right? I grew up in Russia. Then I was actually a Soviet Union. Right, exactly. My dad is a physicist, so the Soviet Union started falling apart, and then he had a visiting position in Kyoto University for a year. We moved there as a family, and then he went to Berkeley, and I kind of tagged along.
0:44And then I graduated from high school. I was thinking about the next thing I wanted to do, and I really like that technical school in Russia. The Russians are serious about the physics. They are. They are. So I went back to Russia, and I got my bachelor's and master's there.
1:00Dmitri Dolgov:What year was this that you went back to Russia? 1994. Okay. So that was kind of almost peak Russian optimism in a sense where it was opening up. It was. It was. Yeah, yeah. No, I actually remember talking to my mom about it. And, of course, my parents grew up in the Soviet Union. They've seen it. They were born right before the war. And then they saw, you know, they lived through some really tough times. And I remember talking to my mom and saying, she, you know, in fact, I got my green card here in the U.S. before I went back, and she insisted that I do it. And I was actually, at the time, I wasn't thinking of coming back.
1:37But then I was pretty excited about where Russia is and the trajectory it's on. And, you know, being young and naive, I was like, there's no turning back.
1:47Dmitri Dolgov:And so why did you decide to come back? There's more of a play-by-play than that. Yeah, yeah, no, school, it was pretty clear. to me like I wanted to continue studying math and computer science. And while the undergrad and masters that I got in physics and applied math, that I think was still an incredibly strong foundational school of Russian math and science. Graduate school, it was very clear to me that the best way to do it was in the US. So I came back. I'm struck by the founders of the two most valuable UK companies are Russian math nerds who both went to the same school. Nikolai at Revolut and Alex Gurko at XTX.
2:37Dmitri Dolgov:But yeah, it's a strong diaspora. There is a company not far from here where one of the founders also has a similar pedigree. A company that we're called to. Exactly. You know the classic engineering interview question of what happens when I type google.com and hit enter? As you know, talk me through whatever you like. HTTP and DNS and BG. You can go down to whatever level of stack you want. Do you want to maybe just describe, when I take a ride in a Waymo today, what's happening at a technical level? Like, what is the architecture? Let me answer your question. It's happening in real time, but this is going to be only a part of the story because we're going to be talking about kind of the inference, the real-time inference part of it.
3:23And if we want to have a deeper, richer technical conversation, I think it would be interesting also to zoom out and talk about the entire ecosystem of what goes into building, evaluating, and deploying the Waymo driver. But when you're driving around or being driven around, we think about what we're building as a driver. Obviously, it's not a car. So it has a number of sensors that are positioned around the vehicle. We use three different sensing modalities. There's cameras, there's lighters or lasers, and there are radars. Those are the primary ones. There are also microphones, directional microphone arrays, but those are the primary three for sensing the world.
4:10They all have very nicely complementary physical properties. They all have 360-degree coverage around the vehicle, so the Waymo driver sees 360 all the time. So all of the data goes into a computer, you would expect. And the software that process, now it's all AI. I can see a specialized AI in the physical world. So it processes the sensor data. Nowadays, we talk about it using AI terminology as encoders that take this data in. And then there's the decoder, the action, the generative part, if you will, in the car. And the generative task there is to figure out how to drive. And that is, of course, connected through a specialized interface to the car where we can actuate the vehicle.
5:01And that's why you see the steering wheel turn and it drives you around.
5:05Dmitri Dolgov:Okay, so I get into my car. There's three main families of sensors, LiDAR, radar, and cameras. And then it is using that to first build a model of what's going on in the world, where are all the other cars and things like that. And then you say, make decisions and then actuate that with the car. That is the system that you're living in. And is all that inference done locally or presumably yes, nothing's in the cloud? Nothing real time. Nothing real time. And there are some things that can happen in the cloud, but they're not required. Got it. What's an example of a nice to have that happens in the cloud?
5:39You can imagine a situation where we do, you know, some of it is not directly related to the task of driving. But let's say after you leave the car, we want to check that, you know, the car is not dirty. You didn't leave anything there. If you did leave an item, well, if you left in a mess, then we want to send the car to one of our depots, get it cleaned up. If you left an item there, maybe on your phone, we want to detect that and then send it to our lost-in-phone and let you know. So that we do by asking a model that actually lives off-board as opposed to having to put it on the car, right?
6:21Because it's not a real-time task related to the driving. So that's one example of something that...
6:25Dmitri Dolgov:There are all these debates that go on on Twitter around self-driving. So I can think of end-to-end versus the more modular approach. There's cameras only versus array of sensors. And I can't tell, are these debates actually interesting to an expert in the field? Or do you think these are just settled matters and they're just grist for the algorithm? I understand where the questions are coming from. I do find that often the way they're posed and the way the debate happens is losing a lot of the nuance and a lot of detail that really matters. are to me the most interesting technical questions are in that level.
7:14Because the way we think about building the Waymo driver, it starts with a large off-board foundation model. I can imagine building a big model that understands how the physical world works and understands the important properties of what it means to drive, the social aspects of driving, and what it means to be a good driver as opposed to a bad one. So that's the foundation. Then we specialize it into what we call three main off-board teachers. There are still large, high-capacity off-board models. There's the Waymo driver, there is the simulator, and then there's the critic. And those then get distilled into smaller models that you can run inference on faster.
8:08So the WEMO driver becomes the backbone, the male backbone of what's in the car. The simulator, of course, is what powers our synthetic generative environment that can run on the cloud for training and for evaluation in close level of the system. And the critic... Sorry, does the simulator ever run locally? No. No, it doesn't. Yeah. However, what I think is interesting, in a way, the way the decoder works, the way the model works, If you think about the generative task in the simulator of creating those realistic worlds and how other people behave, how cars, pedestrians, cyclists, and the task that you have to solve on the car in real time, there is this fundamental shared capability of understanding how these objects relate to each other and predicting what they might do in the future if you are running on the car.
9:04and then generating, you know, some sampling, those probabilistic behaviors in the simulator. So it's a different model, but there is, you know, this is why the shared foundation model is able to power both. And similarly, if you think about the critic, like the job of the critic is to find interesting events and then, you know, be opinionated about what's good behavior and what's bad behavior. Similar fundamental understanding, right? If you're running, you know, inference on the car, you still have to, like, figure out which of the multiple hypotheses
9:33Dmitri Dolgov:of these future worlds you want to take action to steer towards. Okay, and these are all downstream of the same foundation model? Let's start with the foundation model. Then you specialize in fine-tune, still off-board model. Those are the teachers, and then you distill. Each one of the teachers kind of distill, trains its own student, the driver, the simulator, the critic. Yes. You started working on self-driving 20 years ago. As you think about the tech evolution, is this just a scaling law story where we had to be able to throw enough compute at us? Were there architectural approaches we needed to wait to have be invented?
10:16Dmitri Dolgov:Was it just a story of we needed 20 years of going down the wrong cul-de-sacs before we eventually arrived at the right approach? You know, could you, knowing what you know now, could you have a successful Waymo in market in 2015? Or was there some enabling technology? No. Technology breakthroughs that happened over the years were critically important. Primarily in AI, but also in other areas, like, you know, compute. You know, heavy compute than you know. Now, I wouldn't characterize it as like going, you know, a thousand different dead ends and then having to retract and then finding like the one right path.
10:56I would characterize it as iterative learning and evolution. And then, you know, transformers came around. But transformers, for example, are very general architecture, right? Powers of LLMs, powers, you know, our models. But how you apply them to that space, I think this is where... It doesn't just fall out of transformers. Exactly, right? And of course, people like to talk about architectures, but architecture is important. But really, a lot of it comes down primarily to your metrics, to your evaluation mechanisms, to all of the training recipes, and of course, new data. Yes.
11:30Dmitri Dolgov:LMs are good at text or tokens specifically, and obviously perform best at domains that have some kind of single corpus of text they can work on, like coding, where it's very helpful that everything was just kind of textual already. and part of the success has been creating textual representations for domains such that we can then put a lens against them. Can you describe how you encode the world that you're seeing? I mean, are you just building a 3D bitmap essentially? So this is where I think we can get a bit into the
12:14this question of what is the interface between the encoder and the decoder parts. And I think that touches also on the thing you flagged earlier where people like to debate end-to-end or not end-to-end. So the way... Let's talk a little bit about end-to-end and then get back to what is the interface between those two. So when you say end-to-end, what do we mean? mean that it is some large ML model. Typically, you don't build them monolithically. You have different parts and different subgraphs. But what's important is that you can propagate and backprop the gradient and the loss function all through the different layers.
12:57So every layer, you can learn the weights and the representations that matter for the final task. You don't force it through some narrow funnel between, let's say, the encoder and the decoder.
13:09Dmitri Dolgov:Yeah, I think of a simple view of N10 being, you know, pixels go in and car actions come out, which is a bit of an oversimplification. Yeah, that's exactly right. And this is kind of the basic vanilla version of it, right? If you think about what will it take to build the driver that's capable of fully autonomous operations, If you think about this entire ecosystem of the driver, the simulator, the critic, if that's all you do, pixels in, trajectories out, it becomes very difficult to do all of those three and achieve the high level of safety and performance that we require. And it becomes very difficult to kind of do it at scale.
13:57However, it's kind of a very easy way to get started. You collect some data, kind of like an analogy to the LLM world. The easiest thing you can do is pick a model. The easiest way to get started nowadays would be just take a VLM. It already has a language aligned camera encoder. And then it has a decoder that can predict, generate text. And you can fine tune it and say, hey, instead of text, generate trajectories. You know, very, very doable. In fact, a while ago, we published a paper called AMMA that did exactly that. And it will actually, in the nominal case, drive pretty darn well, which is mind-blowingly impressive.
14:45That is very funny, yeah.
14:47Dmitri Dolgov:And, I mean, there's some intuition. You're saying you can take an off-the-shelf model, which has nothing to do with driving to start with, and you'll get these good results. That's right. In the nominal case. I just want to be clear. it's orders of magnitude away from what you need. Yeah, you should not try it on the streets, but it works. But for example, it's like a talking horse. It's impressive that it's talking, you know? Exactly, exactly. And you can actually, the product that you wanted to build was maybe a driver assist system, not a fully autonomous system, then maybe that's all you need to do.
15:16And then for that, you don't need all this other machinery of the simulator and the critic because the number of nines is drastically lower. But this is interesting because there is some intuition behind why that works. If you think about the hard parts of driving, it's not unlike having a conversation, except if in the LLM world, you're modeling language or maybe modeling a dialogue in the space of sentences and words. what makes driving hard is also this kind of multi-agent social interactive part of it. If I do something, that's going to affect you, it's going to affect somebody else. And the history matters.
16:02It's not local and just geometric. Context matters. Semantics matters. But it's in a different, it's not in the language of words, it's in the language of kind of body language, if you will, right? And we see that empirically validated if you do this approach. Okay, so then let's say we build this thing. Just cameras, camera encoder, pixels go in, trajectory go out. The quality is sufficient to drive. In the normal case, it's not sufficient to deal with the long tail of the edge cases and hit the high bar of superhuman safety that we require. So then you start asking the question, what else do you need?
16:42Yes. And if all you did was kind of observing how other people drive when you trained the system. Maybe observing just passively how people drive and how they interact. Maybe also driving the car yourself and then using imitative learning to train it. Mind that that's not enough. You have to do something in closed loop. You have to do things like RLFT, which is also parallel to what we see last time. RLFT? RLFT, reinforcement learning based fine tuning.
17:16Dmitri Dolgov:Okay, yeah, yeah. So, similar to the reinforcement learning with human feedback in the LLM world, right? You want to do maybe a closed-loop proper, closed-loop driving, where you explore all kinds of different situations, and then you give it a reward signal to kind of keep it in distribution. For that, then, you need a realistic simulator, right? You also, if you want to have a good RL system, you need to have an opinion for the reward function. and this is where the critic comes in, right? If you have a purely end-to-end system, let's look at the simulator. Now, what do you do? You have to, you're then constrained to just go from pixels to trajectory, right?
17:57That's all you can run the system on, right? And it's a very high-dimensional space, so it's a hard problem to generate everything. But even if you solve that, it just becomes incredibly inefficient to run it in the full way of pixels to trajectories and simulation for training or for evaluation. So this is when intermediate representations come in. There are some intermediate representations in the world in this task, in the physical world, we know are correct. They're not sufficient, but they're not generality limiting. There's an object here, there's a concept of a road, there's signs, there's speed limits.
18:38So this is where augmenting that learned representation, those learned embeddings beddings from the encoder decoder with that more structured representation is what we do. And we find that this kind of gives us additional knobs to simulate in that space, just pixels to trajectories. It allows us to have additional safety validation layers in real time. And it also gives us additional mechanism to specify the reward function for evaluation of critic or for training. So this is, again, we've gone full circle. Is it N10? Yes, it is. But if you want to do it at scale for full autonomy, it's augmented with all of this other stuff.
19:24Dmitri Dolgov:That's very interesting on the simulating point. It's just very hard to simulate for an N10 model because it's easier to deal in intermediate representations rather than coming up with the pixel perfect view of the world. You need both. So having N10 architecture that's augmented with that structure, allows you to kind of play in both of those worlds. Yeah, yeah, yeah. What are you looking to do as a self-driving car? I mean, it sounds funny, but I think people maybe don't realize that there are many different things that you're looking to solve for, where you're looking to get the person to their destination, you're looking to get them there reasonably promptly, but also drive quite smoothly, and also have many lines of safety, and also not annoy other drivers and get honked at, and, you know, and, and, and.
20:10So what are some of the reward functions or kind of things you're optimizing for that maybe are not obvious to people? So safety is the primary focus, right? But of course, we also want to be a smooth driver for both people in the car and other actors. And we also want to be a predictable well-behaved one so that it can nicely fit into the whole social ecosystem of our roadways.
20:40Dmitri Dolgov:It seems like one of the issues that has quickly emerged with self-driving is the fact that people can't have nice things or, you know, not everyone is nice to the robots. And so, you know, whether you're, you know, driving through a dodgy area or getting blocked or, you know, maybe I'm not going to drop you off here. Maybe I'm going to go around the block and, you know, drop you somewhere better. but all of these, as you say, kind of other human issues. How do you go about solving this? A lot of the ones that you mentioned are just things that we need to work on and understanding, honestly, you said, that if we're not dropping you off exactly where you want it to be dropped off or we don't give you a good interface to tell us, that's on us.
21:29You just got to make it better.
21:32Dmitri Dolgov:It feels like the drop-off is actually a pretty nuanced part of the self-driving journey. Like the highway stuff and the 35-mile-an-hour roads, that is all nailed, but there's just a lot of nuance in the drop-off experience. I'd say they're all hard. You picked freeways and you picked drop-offs for different reasons. For drop-offs, you're absolutely right. There are a few things that are maybe not obvious. You just think about this problem. But it's understanding where you want to go and making it as convenient as possible for you and pickups from drop, it's not exactly symmetrical. But then I was also understanding the context of the station where you, you know, where do you stop?
22:13You don't want to block a driveway, you don't want to, you know, double park, although in some cases where if it's a quick one, maybe it's okay. So there's a lot of nuance that goes into doing that well so that it's a smooth, less frictionless experience for the rider. Yes. As well as other folks. Yeah. Freeways, for most of the time, Not much happens. They're very well structured because we designed them that way. But there is still that long tail of really complicated stuff that happens where the consequences of a bad event are much more severe. Speed is much higher. Everything is quadratic in speed.
23:00but we see a lot of stuff imagine grills falling off of freeways imagine people getting into accidents and kind of spinning out of control you see one of those flatbed trucks with just
23:11Dmitri Dolgov:like a bunch of stuff piled in it and you're driving behind it, I don't know, I always find it very nerve wracking, looks a bit I know, yeah, and we're like we've seen them leave a trail yes, yes, yeah, okay so it's a different set of problems, but it feels I feel like the general sentiment with Waymo is that the driving has mostly now been solved by you guys, and it's kind of a question of scaling up and maybe some super long tail stuff, really snowy conditions. Like, is that your sense internally, or is there actually much more nuance to it than that? I would say the, yeah, it's not like we're done with engineering.
Read the full transcript
23:48I would say that we've clearly moved past the stage of scientific research and kind of deep core technology development to this new phase of accelerated global scaling and deployment. We still have work to do, but I don't see today any limitations or any gaps in the core technology.
24:14Dmitri Dolgov:The driving is good enough now. Well, the core technology, I think, is good enough that I can't think of any aspect of driving that is not supported by the fundamental technology. Now, that said, there is a lot of work to do in specialization and in validation before we can deploy responsibly. We're not driving everywhere in the world. We are planning to start operating in London and in Tokyo this year. Do we have a driver that you're using today in San Francisco that we can just plop down in London and go? No. But what we're seeing is incredibly encouraging from the perspective of, is the core technology there?
25:00So now it's a matter of collecting the data, doing some specialization and validation. And you can see the signs are different in both of those places. People drive on the other side of the road. But that's actually not that hard for computers. And core technology generalizes really well, but you still work that you have to do. What generalizes least well? Increasingly, we're finding, especially now that we're able to kind of hook the Waymo AI to the AI in the digital world, the VLMs, and kind of inherit the general world knowledge from VLMs, we're seeing really strong results from like zero-shot or few-shot learning because of that general knowledge that we bring in.
25:39But there are a few things like, say, cold weather, cold winter weather, where it affects the entire stack. So it's not just the AI. We actually have to. The hardware, yeah. You need the hardware. You need to have the proper cleaning solution, heating elements in it, and then you think about things that are completely solvable by computers like motion control and slippery surfaces. So that takes a bunch of work. You don't get that for free from just pulling in some VLM decoder.
26:11Dmitri Dolgov:Was it the case? I mean, my impression, not knowing anything, is that in the early days, there was maybe a lot of San Francisco-specific work or Phoenix-specific work in the early markets, whether it be mapping or something else, and that you guys seem to either have solved that in generalizing it or just scaled up your ability to do the city-specific work. What enabled the rapid city expansion? We usually think about it, the capability of the Wynwood driver as well as deployment, not primarily and directly in that space of cities or zip codes. I think about the operating domain. And then next, freeways, cold weather, snow, rain, fog, density, etc., etc.
27:03And then that's what we are building, that's where we're evaluating, and then that maps to a city, like a particular city, be within the operating domain or outside of it. So, if we rewind history a little bit, our initial deployment where we started offering a fully autonomous commercial service for the first time was in 2020 in Chandler, Arizona. And that was on what we called the fourth generation of the Waymo Driver. This was, if you remember, the Pacifica minivans with different hardware, different software. There, we were super focused on doing the whole thing end-to-end. Learn how to build the driver, evaluate it, deploy regularly, operate it end-to-end 24-7 with customers, learn from the customers.
27:57And then we're very focused on that operating domain of mostly Chandler, which is a medium, low-complexity one. Then when we made the jump to the fifth generation of our system, this is what's on the high basis today, we really wanted to take a huge bite out of that operating domain. And we collected data all over the United States, all different states, different cities. When we chose to deploy in the hardest parts of San Francisco, hardest parts of Phoenix, we made a big jump on the hardware side and most importantly on the software, the AI side. and I would say that was the big discontinuous jump.
28:35And that's what you're seeing now after we've scaled up and iterated all of the aspects of building and deploying driver, this is now why you're seeing us go in parallel and scaling in the US.
28:50Dmitri Dolgov:So driver version 5 was just a much more generalizable stack than version 4? And what was it about it? Was it just that it had been trained on a much wider dataset? It was when we made this big bet on AI. Yeah. There was a lot more, you know, little AI models and ML models in the fourth generation. Got to meet a much bigger bet and jump to kind of AI as the backbone for the fifth generation. AI is the backbone as the core engine, as in you're saying that Gen 4 had lots of small little AI subsystems? Okay. Yeah, yeah. And that's been, so we've made that jump and we've been iterating and improving the model since then.
29:37Dmitri Dolgov:As we're seeing with Waymo rolling out widespread autonomy, it has second order changes on the entire system. In this case, traffic patterns or other drivers' behavior, or eventually how cities are laid out. And autonomous systems are coming in many domains. In commerce, soon agents are going to be transacting without human intervention. We're basically getting driverless commerce. And Stripe is building the economic infrastructure for AI. And as part of that, we're letting payments be initiated by humans or by agents. So if you want to sell to agents or if you want to let your agents spend money all around the web, check out Stripe's Agenda Commerce Suite.
30:15Dmitri Dolgov:Can we talk about hardware a second? So, lots of hardware questions. But one is maybe everyone in this space has a very charismatic demo of a vehicle that is custom made for self-driving. And so, you know, it's often the van with the, you know, no steering wheel, seats facing in both directions. You know, you guys have one. Tesla has the steering wheel-less cyber cab. You know, Cruise had the Cruise Origin. and yet we're still driving in Jaguars that have a steering wheel in the front and are pretty similar to consumer cars. And it's interesting to me because if we were talking about this 10 years ago, we might say, well, yeah, developing a custom car, that's relatively straightforward.
31:08Dmitri Dolgov:We know how to put a bunch of sensors on a new car, but the software will take a long time. And what's interesting is we've made huge progress in the software, but interestingly the cars are still derivatives of you know cars that people are driving and so i'm curious why you just think the custom hardware has not happened as of 2026 it's obviously it's a small improvement compared to you know waymo is the big improvement but it's just interesting that it still hasn't happened well let's say our sixth generation of the vehicle and the driver is our version of that oh no i know it is oh hi you know platform right so that is you know still has the, you know, we can talk about, you know, whether you want to have the seats pointed backwards or not.
31:49I actually think it looks nice in the demo, but practically speaking, maybe not the way to go. But that is, it is a custom design vehicle, and it is, we put a lot of thought into moving away from a car that's designed around the driver to a car that's designed around the passenger. And it's much more spacious, and it's happening. It's, you know, we're not It's not open to the public yet. But, you know, I took a ride in it the other day, fully autonomously, and that's coming this year. Yes.
32:22Dmitri Dolgov:How much better is it as a passenger experience? You'll tell me once you give it a try. I love it. Okay. So it's, yeah, it's all about the space and the convenience of, you know, ingress and egress and the screens and the interface of the passenger. So we put a lot of thought into every aspect of it. It has sliding doors. It's very easy to get in. It has a flat floor. It is, yeah, if you sit in the back, you can like fully stretch out. And there's so much space there. And it looks, you know, from the outside, you know, it looks fairly big. Yes. Right? But the actual footprint of that is barely, barely, barely larger than the I-PACE.
33:03So it's kind of amazing that, you know, you walk in and it feels like you're in the living room.
33:07Dmitri Dolgov:Yes. I guess my question is just, you know, Waymo does, you know, 25 million rides a year, run-ride-ish, with the Jaguar I-Pace. And it's interesting that so much scaling has happened with self-driving so far on the old, you know, retrofit. Maybe that's to be expected. But I think, well, it matches the high. I don't think it's a given. You're right. But if you think about the value proposition, right? Of course, there is the safety of it. You don't have to worry about it. There's also the privacy. Being in the car by yourself, maybe with other folks, but not having to share that space with another human, right?
33:53No, we have great products, yeah. But I guess this is why we're seeing such consistency in the car. It drives well, very predictable. And you can go beyond that, right? You can specialize even more to make the experience even more magical around the rider. But I guess it would have been disappointing if without the specialized car, and I think I would have been surprised if we leveled off at some other much lower level of customer. Because a car seems like more of an optimization improvement, but the core of the value proposition comes from those other factors.
34:28Dmitri Dolgov:Yes, yes. I guess just take risk on one thing at a time. we'll start by doing the software layer and then we'll build a specialized car or something like that. That's right. Yeah. Yeah. It's also, I mean, as you said, it's a big investment. Yes. You have to like, you de-risk the fundamentals. Yes. And throughout our history, we were very focused on setting the most, the biggest goal for the company to de-risk the most important questions. We talked about the third generation where we wanted to deploy something and go end to end. We talked about what was the goal with the fourth generation, sorry, the fifth generation, and then there's the sixth generation, right?
35:05So it was the sixth generation where it made sense to go and spend all this effort into the custom.
35:11Dmitri Dolgov:And the sixth generation is both the custom vehicle. Is it also a new generation of the driving stack? It is the new hardware. The sensors, the hardware, the self-draining hardware they're putting on the Ojai vehicle, is the sixth generation. It is very different from the fifth generation. It is simpler, it is more capable, it is much lower cost, it's a fraction of the cost is comparable to what you would get like a fancy 8S system nowadays, the driver assist system. The software is pretty much the same. So that's another, so when we talk about generalizability of the Wiimote driver, we talk about weather conditions, we talk about cities, but it also generalizes well to different vehicle platforms and different sensor configurations.
35:59Dmitri Dolgov:Okay, so Gen 6 is a new vehicle and a new sensor stack, but it's almost a tick-tock cycle happening here. It's a similar software. That's right. And then we're going to put the sixth-generation Waymo driver on other vehicle platforms like the Hyundai Ioniq that's coming later in the year. What is different about the sixth-generation hardware stack and how did you make it cheaper? So it still has the same three sensing modalities, but we've made significant optimizations in all three. So unification, simplification, and there's just the kind of just writing the... Is it a classic case of manufacturing scale?
36:43Well, scale hasn't fully come into place, but all of those, if you think about the supply chains, the industries, cameras is pretty mature. radars way many years ago used to be bulky complex very expensive you know when we were putting them on planes but then we started putting them on cars now you can get a decent automotive radar for tons of dollars there is a variant of the automotive radar it's called imaging radar gives you a richer so that is also has come down in cost drastically but it's a little bit behind your standard automotive radars. LIDARs are following the same very predictable, very well-known trend.
37:29So we're writing that, and we're also learning from the previous generation to just make improvements and simplifications and optimizations.
37:36Dmitri Dolgov:I have a very silly question. What are LIDARs versus radars better at in a self-driving company? LIDAR... Are they complementary? They're very complementary. It's all blasting.
37:54Effectively, like you're blasting photons out there and then they bounce off of something, they come back, you measure what comes back. The frequencies are very different. So laser gives you its very, very high resolution. So you can think of it as like a laser beam that goes out, spins around, it shoots out millions of these laser pulses per second. and then each one comes back and you're kind of sampling the 3D structure of the world with very high resolution. So LIDAR for very fine-grained mapping. That's right. Radar has much lower resolution, but because of the physics of it, it degrades much better in adverse weather conditions.
38:39So fog, snow, heavy rain. So it's not going to be occluded by particles between it and the target. So imagine driving in super dense fog. Yes. We're close to San Francisco, so probably don't have to think that hard. It can be really hard to see, so cameras degrade. Yes. Laser, depending on the size of the particulates, can degrade better or worse than camera. Radar is not well affected. So you can imagine driving on a freeway, then radar will give you really good returns for cars that are absolutely invisible in the camera space. That's interesting.
39:18Dmitri Dolgov:So does that mean there are some environments where you'll be relying significantly more on radar? But the performance is good enough? Well, it's a combination of the sensors, right? So we rely on, you know, each one is noisy, right? How the noise characteristics show up in different environments is different. But it is, I mean, it's not like we switch from one to another. It's not like we estimate what's happening with the world through cameras and through radars and through lightars, and then we compare. No, they're like, there's an encoder for camera, there's an encoder for lightars, and they all go into the system that gives you jointly the best view of what's happening in the world.
40:00So if it's a nice, bright, sunny day, cameras are very valuable. If it's pitch dark, or you have sun in your face, or you're blinded by the headlights from an oncoming car, then camera will degrade. There's still some noisy signal, but it will degrade. And LiDAR is completely unaffected.
40:19Dmitri Dolgov:Are there technical problems that are your white whale or you're still chasing or you are particularly interested in solving, even if they're kind of niche for the, you know, we really want to have, you know, driving when it's actually snowing nailed or steep hills in San Francisco or, you know, are there problems you've been very interested in historically or still are? I am super excited right now about the accelerating global expansion. More cities in the United States and going internationally. So being, I don't understand I'm not answering your question about the knowledge. I'll come back to that.
40:57But really that's the thing that I'm today most excited about. Just getting to a place where any major metropolitan area, you can fly into the airport and then take a Waymo and go anywhere you want to go. That is insanely exciting to me right now. So then technically, what I'm most excited about is all of the rapid progress in AI. and the world models, the foundational model work. And it is just such a massive boost to how much we can simplify the system, how much we can bring down the cost, and how we can scale globally. And there's some magic that happens that I don't think I would have anticipated a few years ago.
41:51So that I find from the technical perspective
41:53Dmitri Dolgov:just insanely thrilling. Yes. when you talk about kind of the progress in AI, what are the most fun parts of it for you these days? I think it's seeing the capability and the scaling laws from this approach of starting with that cornerstone of the foundational model and then specializing to t-shirts and then distilling. You get such big wins in performance across the board. I just need to use you you know invest something into you know the architecture or get a better data or training recipe yes and then yeah you invested that early stage and then it just has massive amplification ripple effects so that is in some ways is kind of magical and then you I guess then you see on the car and I've had some moments where you know car does something and you look at a log And I've been surprised.
42:58Like it does things that I didn't think it was capable of doing. So it's that.
43:07Dmitri Dolgov:When you see emergent behavior, that's kind of a proud moment. One example, yeah. When you build a system and then you think, you understand how it works and you understand fully the limits of its capability and performance, and then it does something kind of almost magical, it's accelerating. So one example I can give you, I think I've shared some videos of that publicly in some talks, was this example where the situation that happened in San Francisco, a fairly benign situation where at an intersection, our light is red, there's new cross traffic, a bus goes by, and it stops partially blocking.
43:51annoying. Our light turns green, so we start to go, we're nudging around the bus, and then you see a pedestrian being detected on the other side of the bus. And then your car responds appropriately, it slows down, goes a little bit wider, and then a pedestrian actually emerges from the bus and we go on our own way. So the first time I looked at that log, And what's going on here? I know we have pretty darn good sensors, and the software is very capable. But we don't see through stuff, right? That's not how cameras or lighters and radars work, right? I saw the pedestrian through the bus.
44:30Dmitri Dolgov:I saw the pedestrian on the other side of the bus. And it's not like you look at the windows, you're like, okay, radars shouldn't, it's a massive metal box. Look at the sensor data, and the radar shouldn't be able to go through it. Camera, you can't see in the camera because there's reflections and there's people on the bus, so it's not like you can see through the windows. So what is going on? Maybe it's noise or some coincidence. And the first time I saw it, I couldn't actually believe it. It's like, no, no, there's something. It doesn't sound right. So what actually turned out was happening is that our peripheral lighters bounce under the bus and there was just a little bit of very, very noisy reflection of the movement of the person's feet that was enough for the AI models that, hey, likely there's a pedestrian there and I'm going to, you know, I'll detect it as such.
45:24And moreover, there's enough data there to predict what they're going to do. Yes. It just kind of blew my mind.
45:31Dmitri Dolgov:Is this the perfect example to explain what we were talking about earlier? The value of one, fusion across a sensor suite, But then secondly, building, I mean, relatedly, building an intermediate representation of what's going on, where if you're just dealing with pixels, I mean, the person behind the bus does not exist in pixel space. And so you need to have some representation of the world that exists to be able to reason about the person behind the bus. I think it's an example where giving it kind of an, using that intermediate representation to boost the level of performance of all parts of the model is what's happening here.
46:18Just imagine solving this problem with a black box, purely open loop, imitative system. Hard to impossible. Is it impossible? No. In practice, what would it take to achieve that level of performance? Yes. Very, very difficult.
46:35Dmitri Dolgov:What metrics can you share on just where the business is at today in terms of rides, revenues, cars on the roads? We have about 3 ,000 cars on the roads. We're doing about half a million rides per week. That translates to about over 4 million fully autonomous miles. per week. We are operating in a fully autonomous mode in 11 cities in the US. In 10 of those, we have riders, public riders. What's the ghost city? The ghost city is Nashville. We just started there. So we just opened it up to riders in four new cities in one day. That was one of those little but super exciting moments where I thought back to the history, like how long did it take us from the first time we started fully autonomous rider-only operation to the first time we had external riders in four cities.
47:43Dmitri Dolgov:That's about eight years. And then the other week, we just launched four in one day. Yes, yes. It seems now clear that in 15 years, most miles that are driven will be autonomous. like there'll be some burn in period and there's lots of old cars in the road and some of that will be by level 4, level 5 systems expanding in new cities and that expansion continuing some of it will be you referenced existing driver assist systems and kind of getting up to level 2 and level 3 and existing systems across current car brands getting more and more capable What do you think that working your way up from the lower levels versus working your way expanding from existing products like Waymo?
48:35Dmitri Dolgov:What will that convergence look like? Because we're going to eat it from both sides. I don't believe we will. And I actually think this... That's a great answer. Cars will get smarter. There's going to be advances in driver-assisted systems. And if there is, at the same time, from level four autonomy, there is simplification and the sensors of today are not going to be the sensors of tomorrow. So they'll be much more integrated. They'll be simpler. There'll be much lower cost. So from that perspective, there is a path of convergence. And there's also a path of convergence from the product lines.
49:16There's been ride hailing and you can take a ride through the Waymo app today. Eventually, they'll be on your personal car. So that I see. And talk about the technology. And I see it just as fundamentally two different problems. There's driver assist systems. And then there is full autonomy. And I think it's deceptive to think of them as kind of incremental on one spectrum of complexity.
49:43Dmitri Dolgov:Okay, but you think one cannot work one's way up from driver assist systems to full self-driving? You think you have to start building a full self-driving system? You have to tackle, if I think about the hardest parts of building a fully autonomous rider-only system, they are very different from what you do for a driver assist system. And of course, some work in the space helps you, right? I don't want to say you can't make the jump, but it is a qualitative jump. Yes. When can I buy a Waymo so that I don't need to wait for it when I want to go? I can just like, when I'm ready, I can walk out the door and it's there.
50:26I'm not going to give you a date today, but you're not the first person to bring this up as a product request. Do we know it? Okay. I'll add it to the list.
50:37Dmitri Dolgov:Yes, you know, that waiting for the car, it should be nice just sitting in the garage there and you keep your stuff in it and everything. It's not the first time you've heard that request.
50:47Dmitri Dolgov:So it seems to me operationally very intensive and very hard. Like a self-driving car is actually not self-driving. It takes a village. You have all of the human operator ready to step in. And there was that thundering herd incident that you guys talked about in San Francisco that kind of highlighted that for people. And then there's just like keeping the cars clean and, you know, keeping everything running in that regard. And so, can you describe just what the operational infrastructure that sits behind Waymo looks like? Sure. And I will say that we are overall, you know, in all of those areas on a path of increasing efficiency and automation.
51:35Yep. So, the number of manual steps that one had to do five years ago to launch a Waymo versus where we are today is drastically different. But nowadays, if you look at one of our depots, it's like a fully automatically orchestrated dance of autonomous vehicles. So the way it looks, what it looks like today is cars will automatically go on there to pick up their riders, serve their trips. If for some reason they need to come back, maybe they're low on energy, maybe somebody left a mess in the car, they will automatically come to the depot. If it is, so cleaning today is a manual process. So it'll get flagged in the car.
52:34We have fleet management systems. Say, hey, car number 378 needs cleaning. And we'll actually, on the sensor dome, we're able to display icons. So we'll show you like a little emoji. And there's people whose job it is to clean the car. So they'll come and clean it up. If that cleaning is not required and it's just charging, We'll automatically pull into a charging stall. And we'll say, hey, I need charging. We don't yet have automated charging. In the future, you can imagine that being fully automated. But a person will come in and plug in a cable, and the car will charge, and then say, hey, now I'm ready to go.
53:12And it will get unplugged, and the car will pull out of its parking stall
53:16Dmitri Dolgov:and then go on its merry way. One of the new Porsches, I think it is, has inductive charging, just like your iPhone, where you just drive over the charging mat. I was amazed that that works at car scale, but presumably in the future they'll just be able to drive on with the charging mat. Or do you think just robotic plug-in will be easier? We'll see. We'll see. I don't know. I think there's some questions about efficiency and how that plays into the overall cost and which one will be most cost-beneficial remains to be seen, I think. How well-behaved are the Waymo riding population in terms of not living a mess in the car?
53:54We have wonderful riders. You have the most amazing customers in the world. Generally, I would say they are very good. I think there is something about, I talked about not having a person in the car. It's not somebody else's car. In some ways, you kind of want to preserve the, I think generally people want to preserve the nice aspects of it. It's a broken window thing.
54:20Dmitri Dolgov:It's so clean to begin with. I know. It's kind of like, I think that's the general trend that we see, right? And it's like, because there's not somebody else's space, you're in it, it feels like it's your own. So you don't want to mess up your own space. I don't want to speculate too much on the psychology of things. However, I will say that it varies. And you can imagine a college town on a Saturday night, and that's a different distribution. Yes, yes. Will I be able to get Waymo at any address that has USPS service in the US? Or will there be some head-tail dynamic where Ketchkin, Alaska is just never worth it?
55:04Eventually, it will. Absolutely. There's no doubt in my mind. I think it's just a matter of when and what modality would make the most commercial sense.
55:15Dmitri Dolgov:This is for rideshare versus privately owned. For right, it's not a technical problem. I mean, technology is solved. But then if you're in the middle of nowhere and there's just not enough density of the trips, does it make sense for the right hailing service that Waymo is running to have cars on standby? Yes, yes. Probably not, right? They can be deployed somewhere else and you probably don't want a horribly bad ETA. And this is where a personally-owned vehicle that is equipped with the Waymo driver is maybe how you will see it materialized. Relatedly, what will the second-order effects of, say, majority autonomous traffic be?
55:50Dmitri Dolgov:Like, it feels like a lot of things will work better where, as you say, you know, when someone merges into a lane very poorly and everyone all the way back has to, you know, slam on the brakes, that's kind of anti-social behavior. And so it feels like higher quality and more pro-social driving will just, I mean, basically reduce traffic a little bit, even for the same number of cars on the road. But presumably there'll be other second order effects, like we'll want higher throughput traffic lights. And yeah, how else will things change? So the first thing I think that you mentioned is that's a huge deal.
56:25I just need to think about traffic jams. What's that saying? The Navy SEALs? Slow is smooth and smooth is fast.
56:35Dmitri Dolgov:Traffic jams are like you accelerate abruptly then you come to a stop, and sometimes you have the traffic, like, what happened? Well, an old lady crossed the road three hours ago, and we still have the standing wave there, right? So if everybody was kind of a smooth, predictable driver and a consistent driver, and you would still have those traffic jams at the time off, but then the time constant to clean it out, I think, would be very different. But longer term, things like parking lots. Right now, if you look at what is our most interesting pieces of land allocated to, it's parking lots, it's garages.
57:18And why is that? Well, because, again, your car is just sitting there 90 % of the time, right? If more cars become fully autonomous, then there's no need of that, right? and then imagine, just imagine what you can do with your favorite city in the world if you don't have to spend that money, that huge fraction of it on, just keeping these chunks of metal sitting around.
57:42Dmitri Dolgov:Yeah, I don't think people often realize how big a deal parking minimums are for the layout of the urban landscape. The coffee shop here where I am would like to have outdoor seating, but can't because it would reclaim parking spots. Yeah, wouldn't it be wonderful? Yeah. I have a few more questions, But I'm curious to talk about Google's relationship with self-driving, where, again, it feels like right now, Waymo is, aside from everything else AI-related, kind of the most exciting thing happening at Google. But it was a very long journey to get here. I mean, I feel like you could say that Google almost started working on it too early because you were saying there's been a bunch of recent enabling technologies.
58:29Dmitri Dolgov:And so did it require Google starting when it did so early? Or could one have spun up this project in 2015, 2020? And then how did Google keep the faith when it always felt like it was perennially two years away? Yeah, no, on the latter part, I just have to give credit and huge kudos and gratitude to Larry and Sergey and Alphabet Leadership Center company. It is part of the culture and the DNA of the company is to have that vision and have the stamina and conviction to go the distance.
59:19So to the other part of the question, was it too early? I don't know. I think what we've been seeing, clearly all of the breakthroughs that we've seen over the years have changed how we're building the system. but the complexity of the problem is such that you need to go through these iterative cycles it's not still and we've seen many waves of technology breakthroughs in 2013 ImageNet came around okay that is the right time to start a BSL and transformers came around and VLMs and all of those are super powerful and you have applications and other spaces. In the digital world, they certainly have an impact on our AI and the physical world.
1:00:15There are no silver bullets. They drastically reshape that early part of the curve. It's always been the nature of this problem. It's very easy to get started. It's deceptively easy to get started. But it is super hard to go the full distance and get the number of knives that you have to get. and there's the standard engineering rule of thumb that every next nine takes 10x more. So I, yeah, maybe there is a more optimal path, but I don't see that there's some magical moment where the true complexity of the problem goes away, and then you can just take some off-the-shelf components and you're a business.
1:00:56If that were the case, then I think the industry would look very different today.
1:01:00Dmitri Dolgov:Yeah, yeah. Last question I have. You've been promoted a lot at Google. It feels like Google really recognized your talents. Just what do you think Google does? Like Google is famously one of the very best in the world at technical talent. And say, you know, the current AI wave more broadly happening, you know, is either stuff happening at Google or generally Google alumni. But just what have you observed firsthand from how Google does this so well? Yeah, I would say Google, that culture of Google of not accepting the status quo, having a big vision
1:01:47and investing in technical talent, the people who can go the distance and realize the vision, that is part of the culture. I think this is what you're seeing. And with the breakthroughs in AI, in the digital world and all of the early investments in Transformers and other fundamental technologies, quantum computing. And I guess we're not unlike those efforts as well. Dimitri, thank you. Thank you.
From the publisher
Waymo is now doing 500,000 rides a week across 11 cities. Co-CEO Dmitri Dolgov came to the pub to discuss how the team went from scientific research to global scale. He gives a masterclass on the sensor stack (and why you still need Lidar), how they use "Teacher" and "Critic" models to train the AI, and why he believes cars that require human supervision will never naturally evolve into robotaxis. They also cover the new custom-built vehicle that feels like a living room, the economics of ride-hailing in rural Alaska, and the "Russian math nerd" diaspora that seems to run the UK tech scene.
Timestamps
(00:00:22) Russia
(00:02:51) Waymo architecture
(00:09:59) Why now?
(00:19:46) Driving nuance
(00:29:37) Stripe Agentic Commerce Suite
(00:30:17) Hardware
(00:40:20) Emergent behavior
(00:46:36) Scaling
(00:57:56) Google
Article:
EMMA: End-to-End Multimodal Model for Autonomous Driving – Waymo Research: https://waymo.com/research/emma/




