In short
How Waymo built its fully autonomous driving stack and scaled from research to deployment, including sensor fusion (cameras, LiDAR, radar), AI “foundation model + teachers (driver, simulator, critic) + distillation,” why driver-assist won’t naturally become full autonomy, and what’s required to expand globally (e.g., cold weather hardware).
Key claims
Waymo has nearly half a million fully autonomous rides per week and operates in 11 US cities; core driving technology is “good enough,” with remaining work focused on validation, specialization, and responsible deployment.
Guests
Dmitri Dolgov, Waymo co-CEO; joined Google’s self-driving project in 2009 as one of the first engineers, promoted repeatedly, and became leader in 2021.
Notable examples
A San Francisco intersection case where the system detected a pedestrian “through” a bus; Dolgov said it was actually LiDAR reflections under the bus plus intermediate world representations enabling prediction. He also described depot operations (cleaning/charging orchestration) and discussed upcoming London/Tokyo expansion and Gen 6 custom vehicle/sensor simplification.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VODmitri Dolgov's Early Life
0:49 to 1:48
Explore Dmitri's upbringing in Russia and his academic journey.
“From the sensor stack and why LiDAR still matters, to the role of simulation and critic models in training the AI.”
Transition to the U.S. and Education
1:48 to 3:24
Dmitri shares his motivations for moving to the U.S. for further studies.
“So the Soviet Union started falling apart.”
The Technical Underpinnings of Waymo
3:24 to 4:52
Dmitri discusses the technical architecture of Waymo's self-driving system.
“that I think was still an incredibly strong kind of foundational, you know, school of Russian math and science.”
How Waymo's Sensors Work Together
4:52 to 6:55
Learn about the various sensors used in Waymo vehicles and their functions.
“and talk about kind of the entire ecosystem of what goes into building, evaluating, and deploying the Waymo Driver.”
Inference and Cloud Integration
6:55 to 8:13
Dmitri explains the role of local inference versus cloud tasks in Waymo's operation.
“What's an example of a nice to have that happens in the cloud?”
Foundation Models and Their Evolution
8:13 to 9:29
An overview of how foundational models are developed for self-driving.
“I do find that kind of often the way they're posed and the way the debate happens is losing a lot of the nuance and a lot of detail that really matters.”
Iterative Learning in Self-Driving Tech
9:29 to 11:43
Dmitri reflects on the iterative learning process in developing self-driving technology.
“So the WEMO driver becomes the backbone, the male backbone of what's in the car.”
The Future of Self-Driving and AI Integration
11:43 to 14:01
Discussing the future of self-driving technology and AI's role in it.
“You know, could you, knowing what you know now, could you have a successful Waymo in market in 2015 or was there some enabling technology?”
Understanding ML Models in Autonomous Driving
14:01 to 15:40
Learn how ML models are structured and their application in building autonomous drivers.
“Typically, you don't build them monolithically.”
Challenges in Autonomous Driver Development
15:41 to 17:44
Explore the complexities of building a driver for fully autonomous operations and the safety considerations involved.
“And you can fine-tune it and say, hey, instead of text, generate trajectories.”
Show all 33 chapters
The Role of Simulator and Learning Techniques
17:45 to 21:05
Understand the importance of simulators and reinforcement learning in training autonomous systems.
“In the normal case, it's not sufficient to deal with the long tail of the edge cases and hit the high bar of superhuman safety that we require.”
Navigating Human and Road Interactions
21:06 to 24:59
Discuss the various dynamics of self-driving cars interacting with human drivers and road conditions.
“What are you looking to do as a self-driving car?”
Progress and Challenges in Self-Driving Technology
25:00 to 26:40
Learn about the advancements in core technology and the ongoing challenges faced in scaling and deployment.
“Like, is that your sense internally, or is there actually much more nuance to it than that?”
Adapting Self-Driving Tech to Different Environments
26:41 to 28:00
Examine how self-driving technology adapts to different geographic and environmental challenges.
“results from like zero shot or few shot learning because of that general knowledge that we bring in.”
Waymo's Operating Domain and Initial Deployment
28:00 to 29:19
Explore how Waymo defines its operating domain and its journey in autonomous service deployment.
“We usually think about it, the capability of the windward driver as well as deployment, not primarily and directly in that space of cities or zip codes.”
Advancements in Waymo's Driver Generations
29:20 to 30:59
Learn about the significant improvements from Waymo's fourth to fifth generation systems.
“Then when we made the jump to the fifth generation of our system, This is what's on my basis today.”
Custom Hardware in Self-Driving Vehicles
31:00 to 33:02
Discover the reasons behind the persistence of retrofitted vehicles in self-driving technology.
“has a very charismatic demo of a vehicle that is custom-made for self-driving.”
The Future of Waymo's Sixth Generation Vehicle
33:03 to 35:56
Understand the innovations and features of Waymo's sixth generation vehicle and driving stack.
“How much better is it as a passenger experience?”
Cost Reduction in Sensor Technology
35:57 to 37:55
Examine how Waymo is achieving cost reductions in sensor technology while maintaining performance.
“The sensors, the hardware, the soldering hardware they're putting on the Ojai vehicle is the sixth generation.”
The Role of LIDAR and Radar in Self-Driving
37:56 to 41:20
Delve into the complementary roles of LIDAR and radar in enhancing self-driving capabilities.
“imaging radar and gives you a richer something.”
Excitement for Global Expansion and AI Progress
41:21 to 42:00
Hear about the thrilling prospects of Waymo's expansion and advancements in AI technology.
“More cities in the United States and going internationally.”
The Magic of Foundational Models
42:00 to 43:19
Learn about the transformative impact of foundational models in AI development.
“And the world models, the foundational model work.”
Emergent Behavior in Self-Driving Cars
43:20 to 45:40
Discover how unexpected behaviors in AI systems can lead to excitement and breakthroughs.
“So that is, in some ways, is kind of magical.”
Real-World Challenges and Metrics
45:40 to 47:39
Explore the current operational metrics and the challenges of scaling Waymo.
“So what actually turned out was happening is that our peripheral lighters bounce under the bus.”
The Future of Autonomous Vehicles
47:40 to 49:49
Understand the future trajectories of autonomous driving technology and market expansion.
“We are operating in a fully autonomous mode in 11 cities in the US.”
Operational Infrastructure Behind Waymo
49:50 to 52:00
Gain insight into the operational aspects and efficiency of Waymo's services.
“And there's also a path of convergence from the product lines.”
User Experience and Future Accessibility
52:00 to 55:40
Discuss the user experience factors and future accessibility of Waymo's services.
“I will say that we are overall, you know, in all of those areas on a path of increasing efficiency and automation.”
Impacts of Autonomous Traffic
55:40 to 56:00
Examine the broader implications of a shift to predominantly autonomous driving.
“I think it's just a matter of when and what modality would make the most commercial sense.”
The Future of Ride-Hailing with Autonomous Vehicles
56:00 to 56:41
Explore the implications of Waymo's autonomous technology on transportation density and efficiency.
“But then if you're in the middle of nowhere and there's just not enough density of the trips, does it make sense for the right hailing service that Waymo is running to have cars on standby?”
Impact of Autonomous Driving on Traffic Dynamics
56:41 to 57:50
Understand how autonomous driving could reshape traffic behavior and urban planning.
“And so it feels like higher quality and more pro-social driving will just, I mean, basically reduce traffic a little bit, even for the same number of cars on the road.”
Waymo's Journey and Google's Vision
57:50 to 59:18
Learn about Waymo's development journey and the strategic vision behind Google's investment in automation.
“to, you know, it's parking lots, it's garages.”
Challenges in Building Autonomous Systems
59:18 to 1:01:12
Delve into the complexities and technology cycles behind developing reliable autonomous systems.
“And then how did Google keep the faith when it always felt like it was perennially two years away?”
Google's Culture of Innovation and Talent Development
1:01:12 to 1:02:35
Discover how Google's culture fosters innovation and nurtures talent in the tech industry.
“There's the standard engineering rule of thumb that every next nine takes 10x more.”
Transcript
Automatic transcript. May contain errors.0:00When you're driving around or being driven around, say, you know, we think about what we're building as a driver. I can imagine building a big model that understands how the physical world works and understands the important properties of what it means to drive, the social aspects of driving, and what it means to be a good driver as opposed to a bad one. I would say that we've clearly moved past the stage of scientific research and deep core technology development to this new phase of accelerated global scaling and deployment.
0:39Dmitri Dolgov:Waymo is now doing nearly half a million fully autonomous rides a week across multiple cities, a shift from long-term research to real-world scale. In this episode, originally aired on the Cheeky Pint podcast, Waymo co-CEO Dmitry Delgov joins John Collison to break down how they built the system behind it. From the sensor stack and why LiDAR still matters, to the role of simulation and critic models in training the AI. They also get into why driver assist won't naturally evolve into full autonomy. what it takes to scale globally, and how the product itself is changing from custom-built vehicles to entirely new economies of ride-hailing.
1:23Dmitri Dolgov:Dimitri Dalgov is co-CEO of Waymo. He joined Google's self-driving car project in 2009 as one of its first engineers and was repeatedly promoted until he took it over in 2021. Waymo is Google's most successful moonshot and now provides over 500 ,000 fully autonomous rides each week. Cheers, by the way. Yeah, cheers. You grew up in Russia, right? I grew up in Russia. Yeah. Then I was actually Soviet Union. Right, right, exactly. My dad is a physicist. So the Soviet Union started falling apart. And then he had a visiting position in Kyoto University for a year. We moved there as a family. And then he went to Berkeley.
2:03And I kind of tagged along. And then I graduated from high school. I was thinking about the next thing I wanted to do. and I really like that technical school in Russia. The Russians are serious about this. They are, they are. So I went back to Russia and I got my bachelor's and master's there.
2:20Dmitri Dolgov:What year was this that you went back to Russia? 1994. Okay. So that was kind of almost peak Russian optimism in a sense where it was opening up. It was, it was. Yeah, yeah, no, I actually remember talking to my mom about it. And of course, my parents grew up in the Soviet Union. They've seen it. they were born right before the war and then they saw, they lived through some really tough times and I remember talking to my mom and saying in fact I got my green card here in the US before I went back and she insisted that I do it and I was actually at the time I wasn't thinking of coming back but then I was pretty excited about where Russia is and trajectory it's on and being young and naive There's no turning back.
3:08And so why did you decide to come back? There's more of a play back. Yeah, yeah, no, school, it was pretty clear to me. Like I wanted to continue studying math and computer science. And while the undergrad and master's that I got in physics and applied math, that I think was still an incredibly strong kind of foundational, you know, school of Russian math and science. graduate school, it was very clear to me that the best way to do it was in the US. So I came back.
3:40Dmitri Dolgov:I'm struck by the founders of the two most valuable UK companies are Russian math nerds who both went to the same school. Nikolai Ash Revolut and Alex Gerko at XTX. But yeah, it's a strong diaspora. There is a company not far from here where one of the founders also has a similar pedigree. A company that we're close to. Exactly. You know the classic engineering interview question of what happens when I type google.com and hit enter as talk me through whatever you like, HTTP and DNS and BG, you can go down to whatever level of stack you want. Do you want to maybe just describe when I take a ride in a Waymo today, what's happening at a technical level?
4:32Dmitri Dolgov:Like, what is the architecture? Let me answer your question. I was happening in real time, but this is going to be only a part of the story because we're going to be talking about kind of the inference, the real-time inference part of it. And if we want to have a deeper, richer technical conversation, I think it would be interesting also to zoom out and talk about kind of the entire ecosystem of what goes into building, evaluating, and deploying the Waymo Driver. But when you're driving around or being driven around, we think about what we're building as a driver. Obviously, it's not a car. So it has a number of sensors that are positioned around the vehicle.
5:13We use three different sensing modalities. There's cameras, there's lighters or lasers, and there are radars. Those are the primary ones. there are also microphones, directional microphone arrays, but those are the primary three for sensing the world. They all have very nicely complementary physical properties. They all have 360-degree coverage around the vehicle, so the Waymo driver sees 360 all the time. So all of the data goes into a computer, you would expect. And they're the software that process, now it's all AI. I can see a specialized AI in the physical world. So it processes the sensor data.
5:56Nowadays, you know, talk about it in the, you know, using AI terminology as, you know, encoders that, you know, take this data in. And then there's the kind of the decoder, the action, you know, the generative part, if you will, in the car. And the generative task there is to, you know, figure out how to drive, right? And that is, of course, connected through kind of a specialized interface to the car where we can actuate the vehicle and that's why you see the steering wheel turn and it drives you around.
6:25Dmitri Dolgov:Okay, so I get into my car, there's three main families of sensors, LiDAR, radar, and cameras, and then it is using that to first build a model of what's going on in the world, where are all the other cars and things like that, and then you say, make decisions and then actuate that with the car. That is the system that you're living in. And is all that inference done locally or presumably yes, nothing's in the cloud? Nothing real-time? Nothing real time in the cloud. And there are some things that can happen in the cloud, but they're not required. Got it. What's an example of a nice to have that happens in the cloud?
6:59You can imagine a situation where we do, some of it is not directly related to the task of driving, but let's say after you leave the car, we want to check that the car is not dirty, you didn't leave anything there. If you did leave an item, well, if you left in a mess, then, you know, I want to send the car to one of our depots, get it cleaned up. If you left an item there, you know, your phone, or we want to detect that. And then, you know, send it to our list on phone and let you know. Right? So that, you know, we do with the kind of by asking a model that actually lives off board as opposed to having to put it on the car, right?
7:41Because it's not a real-time task related to, you know, the driving. So that's one example of something that...
7:46Dmitri Dolgov:There are all these debates that go on on Twitter around self-driving. So I can think of, you know, end-to-end versus the more kind of modular approach. There's cameras only versus array of sensors. And I can't tell, are these debates actually interesting to an expert in the field? Or do you think these are just settled matters and they're just grist for the algorithm? I understand where the questions are coming from. I do find that kind of often the way they're posed and the way the debate happens is losing a lot of the nuance and a lot of detail that really matters. are to me the most interesting technical questions are in that level.
8:35Because the way we think about building the Waymo driver, it starts with a large off-board foundation model. I can imagine building a big model that understands how the physical world works and understands the important properties of what it means to drive, the social aspects of driving, and what it means to be a good driver as opposed to a bad one. So that's the foundation. Then we specialize it into what we call three main off-board teachers. There are still large, high-capacity off-board models. There's the Waymo driver, there is the simulator, and then there's the critic. And those then get distilled into smaller models that you can run inference on faster.
9:29So the WEMO driver becomes the backbone, the male backbone of what's in the car. The simulator, of course, is what powers our synthetic generative environment that can run on the cloud for training and for evaluation in close of the system. And the critic, they value it. Sorry, does the simulator ever run locally? No. No, it doesn't. Yeah. However, what I think is interesting, in a way, the way the decoder works, the way the model works, if you think about the generative task in the simulator, of kind of creating those realistic worlds and how other people behave, how cars, pedestrians, cyclists, and the task that you have to solve on the car in real time, there is this fundamental shared capability of understanding how these objects relate to each other and predicting what they might do in the future if you are running on the car and then generating some sampling those probabilistic behaviors in the simulator.
10:29So it's a different model, but this is why the shared foundation model is able to power both. And similarly, if you think about the critic, the job of the critic is to find interesting events and then be opinionated about what's good behavior and what's bad behavior. Similar fundamental understanding, if you're running inference on the car, you still have to figure out which of the multiple hypotheses of these future worlds you want to... take action to steer towards.
10:59Dmitri Dolgov:Okay, and these are all downstream of the same foundation model? That's right. So start with the foundation model. Then you specialize in fine-tune, still off-board model. Those are the teachers, and then you distill. Each one of the teachers kind of distill trains its own student. Yes. The driver, the simulator, the critic. You started working on self-driving 20 years ago. As you think about the tech evolution, is this just a scaling laws story where we had to be able to throw enough compute at us? Were there architectural approaches we needed to wait to have be invented? Was it just a story of we needed 20 years of going down the wrong cul-de-sacs before we eventually arrived at the right approach?
11:44Dmitri Dolgov:You know, could you, knowing what you know now, could you have a successful Waymo in market in 2015 or was there some enabling technology? No. Technology breakthroughs that happened over the years were critically important, primarily in AI, but also in other areas, like compute, heavy compute. Now, I wouldn't characterize it as like going, you know, a thousand different dead ends and then having to retract and then finding like the one right path. I would characterize it as iterative learning and evolution. and then, you know, transformers came around, but transformers, for example, are very general architecture, right?
12:26Powers of LLMs, powers, you know, our models. But how you apply them to that space, I think this is where... You didn't just fall out of transformers. Exactly, right? And of course, you know, people like to talk about architectures, but architecture is important, but really a lot of it comes down primarily to your metrics, to your evaluation mechanisms, to, you know, all of the training recipes, and of course, new data. Yes.
12:50Dmitri Dolgov:LLMs are good at text or tokens specifically, and obviously perform best at domains that have some kind of single corpus of text they can work on, like coding, where it's very helpful that everything was just kind of textual already. And part of the success has been creating textual representations for domains so that we can then put LLMs against them. Can you describe how you encode the world that you're seeing? I mean, are you just building a 3D model, like a 3D bitmap essentially, or? So this is where I think we can get a bit into
13:35the question of what is the interface between the encoder and the decoder parts. And I think that touches also on the, you know, thing you flagged earlier, where people like to, you know, debate end-to-end or not end-to-end. And so the way, let's talk a little bit about end-to-end and then get back to like what is the interface between those two, right? So when you say end-to-end, what do we mean? We mean that it is some large ML model. Typically, you don't build them monolithically. You have, you know, different parts and different subgraphs. But what's important is that you can propagate and backprop the, you know, gradient and the the loss function all through the different layers.
14:18So every layer you can learn the weights and the representations that matter for the final task. You don't force it through some narrow funnel between, let's say, the encoder and the decoder.
14:29Dmitri Dolgov:Yeah, I think of a simple view of NTN being, you know, pixels go in and car actions come out, which is maybe a bit of an oversimplification. Yeah, that's exactly right. And this is kind of the basic vanilla version of it, right? if you think about
14:49what will it take to build the driver that's capable of fully autonomous operations if you think about this entire ecosystem of the driver the simulator the critic if that's all you do pixels in, trajectories out it becomes very difficult to do all of those three and achieve the high level of safety and performance that we require and it becomes very difficult to kind of do it at scale.
15:18However, it's kind of a very easy way to get started. You collect some data, kind of like an analogy to the LLM world. The easiest thing you can do is pick a model. The easiest way to get started nowadays would be just take a VLM. It already has a language-aligned camera encoder. and then it has a decoder that can predict, generate text. And you can fine-tune it and say, hey, instead of text, generate trajectories. Very, very doable. In fact, a while ago we published a paper called AMA that did exactly that. And it will actually, in the nominal case, drive pretty darn well, which is mind-blowingly impressive.
16:06That is very funny, yeah.
16:08Dmitri Dolgov:And I mean, there's some intuition. You're saying you can take an off-the-shelf model, which has nothing to do with driving to start with, and you'll get these good results. That's right. In the normal case. I just want to be clear. It's orders of magnitude away from what you need. Yeah, you should not try it on the street, but it works. It's like a talking horse. It's impressive that it's talking. Exactly, exactly. And you can actually, if the product that you wanted to build was maybe a driver-assist system, not a fully autonomous system, then maybe that's all you need to do. And then for that, you don't need all this other machinery of the simulator and the critic because the number of nines is drastically lower.
16:43But this is interesting because there is some intuition behind why that works. If you think about the hard parts of driving, it's not unlike having a conversation. except if in the LLM world, having, you know, you're modeling language or maybe modeling a dialogue in the space of sentences and words. What makes driving hard is also this kind of multi-agent social interactive part of it. And if I do something that's going to affect you, it's going to affect somebody else. And the history matters. It's not local and just geometric. Context matters. Semantics matters. But it's in a different, it's not in the language of words, it's in the language of kind of body language, if you know what, right?
17:34And we see that empirically validated if you do this approach. Okay, so then let's say we build this thing, just cameras, camera encoder, pixels go in, trajectory go out. The quality is sufficient to drive. In the normal case, it's not sufficient to deal with the long tail of the edge cases and hit the high bar of superhuman safety that we require. So then you start asking the question, what else do you need? And if all you did was kind of observing how other people drive when you trained the system, maybe observing just passively how people drive and how they interact, maybe also driving the car yourself and then using imitative learning to train it, that's not enough.
18:23You have to do something in closed loop. You have to do things like RLFT, which is also parallel to what we see last time. RLFT? RLFT, reinforcement learning based fine tuning. Okay, yeah, yeah. So similar to the reinforcement learning with human feedback in the LLM world, right? You want to do maybe closed loop, proper closed loop driving where you explore all kinds of different situations and then you give it a reward signal to kind of keep it in distribution. For that, then, you need a realistic simulator. Right? You also, you know, if you want to have a good RL system, you need to have an opinion for the reward function.
19:07This is where the critic comes in. Right? If you have a purely end-to-end system, let's look at the simulator. Now, what do you do? You have to, you're then constrained to just go from pixels to traject, right? That's all, you know, you can run the system on, right? and it's a very high dimensional space so it's a hard problem to generate everything. But even if you solve that, it just becomes incredibly inefficient to run it in the full way of pixels to trajectories and simulation for training or for evaluation. So this is when intermediate representations come in. There are some intermediate representations in the world in this task, in the physical world we know are correct.
19:50They're not sufficient, but they're not generality limiting. There's an object here, there's a concept of a road, there's signs, there's speed limits. So this is where augmenting that learned representation, those learned embeddings from the encoder-decoder with that more structured representation is what we do. And we find that this kind of gives us additional knobs to simulate in that space, just pixels to trajectories. trajectories. It allows us to have additional safety validation layers in real time. And it also allows us, you know, it gives us additional mechanisms to specify the reward function, you know, for evaluation of critic or, you know, for training.
20:34So this is again, like we've gone kind of full circle of it. Is it end-to-end? Yes, it is. Yes. But if you want to do it at scale for full autonomy, it's augmented with all of this other stuff.
20:45Dmitri Dolgov:That's very interesting on the simulating point. It's just very hard to simulate for an end-to-end model because it's easier to deal in end-to-end, or it's easier to deal in intermediate representations rather than coming up with the pixel perfect view of the world. You need both. Yeah. So, you know, having end-to-end architecture that's augmented with that structure allows you to kind of play in both of those worlds. Yeah, yeah, yeah. What are you looking to do as a self-driving car? I mean, it sounds funny, but I think people maybe don't realize that there are many different things that you're looking to solve for where you're looking to get the person to their destination, you're looking to get them there reasonably promptly, but also drive quite smoothly and also have many lines of safety and also not annoy other drivers and get honked at and, you know, and, and, and.
21:31Dmitri Dolgov:So what are some of the reward functions or kind of things you're optimizing for that maybe are not obvious to people? So safety is the primary focus, right? But of course, we also want to be a smooth driver. so that for both people in the car and other actors. And I also want to be a predictable well-behaving so that it can nicely fit into the whole social ecosystem of our roadways. It seems like one of the issues that has quickly emerged with self-driving is the fact that people can't have nice things or not everyone is nice to the robots. And so, you know, whether you're, you know, driving through a dodgy area or getting blocked or, you know, maybe I'm not going to drop you off here.
22:24Dmitri Dolgov:Maybe I'm going to go around the block and, you know, drop you somewhere better. But all of these, as you say, kind of other human issues, how do you go about solving this? A lot of the ones that you mentioned are just things that, you know, we need to work on. And understanding, honestly, you know, said that if we're not dropping you off, we're exactly where you want it to be dropped off. or we don't give you a good interface to tell us, that's on us. We just got to make it better. It feels like the drop-off is actually a pretty nuanced part of the self-driving journey. Like the highway stuff and the 35-mile-an-hour roads, that is all nailed, but there's just a lot of nuance in the drop-off experience.
23:06I'd say they're all hard. You picked freeways and you picked drop-offs. For different reasons. For drop-ups, you're absolutely right. There are a few things that are maybe not obvious. You just think about this problem. But it's understanding where you want to go and making it as convenient as possible for you. And pick-ups from drop-ups, it's not exactly symmetrical. But then it was also understanding the context of the situation where do you stop? You don't want to block a driveway, you don't want to double park. Although in some cases where if it's a quick one, maybe it's okay. So there's a lot of nuance that goes into doing that well, so that it's a smooth, frictionless experience for the rider, as well as other folks.
23:50Freeways, for most of the time, not much happens. They're very well structured because we designed them that way. But there is still that long tail of really complicated stuff that happens where the consequences of, you know, a bad event are much more severe, right? The speed is much higher. Everything is, you know, quadratic in speed. So, but we see a lot of stuff there. You imagine grills falling off of freeways. Imagine people getting into accidents and kind of spinning out of control. Can you see one of those flatbed trucks
24:31Dmitri Dolgov:with just like a bunch of stuff piled in it and you're driving behind it? I don't know. I always find it very nerve-wracking. It looks a bit... I know. Yeah. And we're like, we've seen them, you know, leave a trail. Yes. Yes. Yeah. Okay. So it's a different set of problems. But I feel like the general sentiment with Waymo is that the driving has mostly now been solved by you guys. And it's kind of a question of scaling up and maybe some super long tail stuff, really snowy conditions. Like, is that your sense internally, or is there actually much more nuance to it than that? I would say the, yeah, it's not like, you know, we're done with engineering.
Read the full transcript
25:08I would say that we've clearly moved past the stage of scientific research and kind of deep core technology development to this new phase of accelerated global scaling and deployment. Yes. We still have work to do, but I don't see today any limitations or any gaps in the core technology. The driving is good enough now. Well, the core technology, I think, is good enough that I can't think of any aspect of driving that is not supported by the fundamental technology. Now, that said, there is a lot of work to do in specialization and in validation before we can deploy responsibly, right? We're not driving everywhere in the world.
26:00We are planning to start operating in London and in Tokyo this year. Do we have a driver that you're using today in San Francisco that we can just plop down in London and go? No, right? But what we're seeing is incredibly encouraging from the perspective of, like, is the core technology there? So now it's a matter of collecting the data, doing some specialization and validation. And you can see the signs are different. In both of those places, people drive on the other side of the road. But that's actually not that hard for computers. And core technology generalizes really well, but you still work that you have to do.
26:37What generalizes least well? Increasingly we're finding, especially now that we're able to kind of hook the Waymo AI to the AI in the digital world, the VLMs, and kind of inherit the general world knowledge from VLMs, we're seeing really strong results from like zero shot or few shot learning because of that general knowledge that we bring in. But there are a few things like, say, cold weather, cold winter weather, where it affects the entire stack. So it's not just the AI, but you actually have to. Hardware, yeah. You need the hardware, you need to have the proper cleaning solution, heating elements in it, and then you think about things that are completely solvable by computers, like motion control and slippery surfaces, right?
27:25So that takes a bunch of work. You don't get that for free from just pulling it. some, you know, VLM decoder.
27:32Dmitri Dolgov:Was it the case, I mean, my impression, not knowing anything, is that in the early days, there was maybe a lot of San Francisco-specific work or Phoenix-specific work in the early markets, whether it be mapping or something else, and that you guys seem to either have solved that in generalizing it, or just scaled up your ability to do the city-specific work. Quash enabled the rapid city expansion? We usually think about it, the capability of the windward driver as well as deployment, not primarily and directly in that space of cities or zip codes. I think about the operating domain. And then that's just the freeways, cold weather, snow, rain, fog, density, et cetera, et cetera.
28:24And then that, that's what we are building, that's where we're evaluating. And then that maps to a city, like a particular city, be within the operating domain or outside of it. So where, if we rewind history a little bit, our initial deployment in where we started offering a fully autonomous commercial service for the first time was in 2020 in Chandler, Arizona. And that was on what we called the fourth generation of the Waymo driver. This was the, if you remember, the Pacifica minivans with different hardware, different software. There, we were super focused on doing the whole thing end-to-end.
29:08Learn how to build the driver, evaluate it, deploy regularly, operate it end-to-end 24-7 with customers, learn from the customers. And they were very focused on that operating domain of mostly Chandler, which is a medium, low-complexity one. Then when we made the jump to the fifth generation of our system, This is what's on my basis today. We really wanted to take a huge bite out of that operating domain. And we collected data all over the United States, all different states, different cities. When we chose to deploy in the hardest parts of San Francisco, hardest parts of Phoenix, we made a big jump on the hardware side and most importantly on the software, the AI side.
29:51And I would say that was the big discontinuous jump. And that's what you're seeing now after we've scaled up and iterated on all of the aspects of building and deploying driver. This is now why you're seeing us go in parallel and scaling in the US.
30:11Dmitri Dolgov:So driver version 5 was just a much more generalizable stack than version 4? And what was it about us? Was it just that it had been trained on a much wider data set? It was when we made this big bet on AI. I think there was a lot more, you know, little AI models and ML models in the fourth generation. Got to make a much bigger bet and jump to kind of AI as the backbone for the fifth generation. AI is the backbone as the core engine, as in you're saying that Gen 4 had lots of small little AI subsystems? Yeah. And that's been, so we kind of made that jump and we've been iterating and improving the model since then.
30:55Can we talk about hardware a second? So, lots of hardware questions,
30:59Dmitri Dolgov:but one is maybe everyone in this space has a very charismatic demo of a vehicle that is custom-made for self-driving. And so, you know, it's often the van with the, you know, no steering wheel, seats facing in both directions. You know, you guys have one, Tesla has the steering wheel-less cyber cab. You know, Cruise had the Cruise Origin. And yet, we're still driving in Jaguars that have a steering wheel in the front and are pretty similar to consumer cars. And it's interesting to me, because, you know, if we were talking about this 10 years ago, we might say, well, yeah, developing a custom car, like, that's relatively straightforward.
31:49Dmitri Dolgov:We know how to put a bunch of sensors on a new car. but the software will take a long time. And what's interesting is we've made huge progress in the software, but interestingly, the cars are still derivatives of, you know, cars that people are driving. And so I'm curious why you just think the custom hardware has not happened as of 2026. It's obviously, it's a small improvement compared to, you know, Waymo is the big improvement, but it's just interesting that it still hasn't happened. Well, let's say our sixth generation of the vehicle and the driver is our version of that. It is the O-Hike platform, right?
32:24So that is, you know, it still has the, you know, we can talk about, you know, whether you want to have the seats quite backwards or not. I actually, you know, think it looks nice in a demo, but practically speaking, maybe not the way to go. But that is, it is a custom designed vehicle. And it is, we put a lot of thought into, you know, moving away from a car that's designed around the driver to a car that's designed around passenger. And it's much more spacious, But it's happening. It's not open to the public yet. But I took a ride in it the other day, fully autonomously, and that's coming this year.
33:02Yes. How much better is it as a passenger experience? You'll tell me once you give it a try. I love it. Okay. So it's all about the space and the convenience of ingress and egress and the screens and the interface of the passenger. So we put a lot of thought into every aspect of it. It has sliding doors. It's very easy to get in. It has a flat floor. It is, yeah, if you sit in the back, you can like fully stretch out. There's so much space there. And it looks, you know, from the outside, you know, it looks fairly big. But the actual footprint of that is barely, barely, barely larger than the I-PACE.
33:43So it's kind of amazing that, you know, you walk in and it feels like you're in the living room.
33:47Dmitri Dolgov:Yes. I guess my question is just, you know, Waymo does 25 million rides a year, run-ride-ish, with the Jaguar I-Pace. And it's interesting that so much scaling has happened with self-driving so far on the old retrofit. Maybe that's to be expected. But I think, well, it matches the high. I don't think it's a given. You're right. But if you think about the value proposition, of course there is the safety of it, you don't have to worry about it. There's also the privacy, being in the car by yourself, maybe with other folks, but not having to share that space with another human, right? No, Wayne has great products, yeah.
34:35But I guess this is why we're seeing such consistency in the car, it drives well, very predictable, and you can go beyond that, right? and you specialize even more to make the experience even more magical around the rider. But I guess it would have been disappointing if without the specialized car, and I think I would have been surprised if we leveled off at some other much lower level of customer adoption. Because a car seems like more of an optimization improvement, but the core of the value proposition comes from those other factors.
35:08Dmitri Dolgov:Yes, yes. I guess to just take risk on one thing at a time, we'll start by doing the software layer and then we'll build a specialized car or something like that. That's right. That's right. Yeah. Yeah. Yeah. It's also, I mean, as you said, it's a big investment. Yes. So you have to like, you de-risk the fundamentals. Yes. Um, and you know, throughout our history, we were very focused on setting the most, you know, the biggest goal for the company to de-risk the most important questions, right? We talked about, you know, the third generation where, you know, we wanted to deploy something and go end to end.
35:40We talked about the, what was the goal with the fourth generation and then, oh, sorry, the fifth generation, and then there's the sixth generation, right? So it was the sixth generation where it made sense to go and spend all this effort into the custom.
35:51Dmitri Dolgov:And sixth generation is both a custom vehicle. Is it also a new generation of the driving stack? Yeah, it is the new hardware. Yep. The sensors, the hardware, the soldering hardware they're putting on the Ojai vehicle is the sixth generation. It is very different from the fifth generation. It is simpler. It is more capable. It is much lower cost. It's a fraction of the cost that's comparable to what you would get like a fancy 8S system. Nowadays, the driver assist system. The software is pretty much the same. So when we talk about generalizability of the Wiimote driver, we talk about weather conditions, we talk about cities, but it also generalizes well to different vehicle platforms and different sensor configurations.
36:39Dmitri Dolgov:Okay, so Gen 6 is a new vehicle and a new sensor stack, but it's almost a tick-tock cycle happening here. It's a similar software. That's right. And then we're going to put the sixth-generation Waymo driver on other vehicle platforms, like the Hyundai Ioniq that's coming later in the year. What is different about the sixth-generation hardware stack, and how did you make it cheaper? So it still has the same three sensing modalities, but we've made significant optimizations in all three. Yeah. So unification, simplification, and there's just, you know, the kind of just writing the... Yeah, is it a classic case of, you know, manufacturing scale?
37:23Well, scale hasn't fully come in place, but all of those, if you think about the supply chains, the industries, cameras is pretty mature. radars way many years ago used to be bulky, complex, very expensive when we were putting them on planes but then we started putting them on cars. Now you can get a decent automotive radar for tens of dollars. There is a variant of the automotive radar it's called imaging radar and gives you a richer something. So that is also has come down and caused drastically but it's a little bit behind your standard automotive radars. LIDARs are following the same very predictable, very well-known trend.
38:09So we're writing that and we're also learning from the previous generation to just make improvements and simplifications and optimizations.
38:16Dmitri Dolgov:I have a very silly question. What are LIDARs versus radars better at in a self-driving company? LIDAR... Are they complementary? They're very complementary. It's all blasting. you know, uh, uh, effectively, like, you know, blasting, you know, photons out there and then, uh, they bounce off of something, they come back, you know, you measure what comes back. The frequencies are very different. So laser, uh, gives you, it's, uh, very, very high resolution. So you can, you know, think of it as like a laser beam that goes out, you know, spins around, it, you know, shoots out millions of these laser pulses, you know, per second.
38:59And then each one comes back and you're kind of sampling the 3D structure of the world with very high resolution. So LIDAR for very fine-grained mapping. That's right. Radar has much lower resolution, but because of the physics of it, it degrades much better in adverse weather conditions. So fog, snow, heavy rain. So it's not going to be occluded by verticals between it and the target. So imagine driving in super dense fog. Yes. We're close to San Francisco, so we probably don't have to think that hard. It can be really hard to see. So cameras degrade. Yes. Laser, depending on kind of the size of the particulates, can degrade better or worse than camera.
39:46Radar is not well affected. So you can imagine driving on a freeway, then radar will give you really good returns for cars that are absolutely invisible in the camera space.
39:56Dmitri Dolgov:That's interesting. So does that mean there are some environments where you'll be relying significantly more on radar? Well, it's a combination of the sensors, right? So we rely on, you know, each one is noisy, right? How the noise characteristics show up in different environments is different, but it is, I mean, it's not like we switch from one to another. It's not like we know we estimate you know what's happening with the world through cameras and through radars and through lighters and then we compare no they're like there's an encoder for camera there's an encoder for lighter there's a corner and they all go into the you know the system uh that gives you jointly the best view of what's happening uh in the world so if you were you know if it's a nice bright sunny day cameras are you know very valuable if you know it's pitch dark or you have like sun in your face or you're blinded by the headlights from you know oncoming car then camera will degrade.
40:52There's still some, you know, noisy signal, but it will degrade. And radar, LiDAR is completely unaffected.
40:59Dmitri Dolgov:Are there technical problems that are your white whale, or you're just, you're still chasing, or you are particularly interested in solving, even if they're kind of niche for the, you know, we just, we really want to have, you know, driving when it's actually snowing nailed, or steep hills in San Francisco, or, you know, are there problems you've been very interested in historically, or still are? I'm super excited right now about the accelerating global expansion. More cities in the United States and going internationally. So being, I understand I'm not answering your question about the knowledge.
41:37I'll come back to that. But really, that's the thing that I'm today most excited about. Just getting to a place where any major metropolitan area, you can fly into the airport and then take away more and go anywhere you want to go. Like that is insanely exciting to me right now. So then, you know, technically what I'm most excited about is all of the rapid progress in AI. And the world models, the foundational model work. and it is just such a massive boost to how much we can simplify the system, how much we can bring down the cost and how we can scale globally. And there's some magic that happens that I don't think I would have anticipated a few years ago.
42:31So that I find from the technical perspective just insanely thrilling.
42:36Dmitri Dolgov:Yes. When you talk about kind of the progress in AI, what are the most fun parts of it for you these days? I think it's seeing the capability and the scaling laws from this approach of starting, you know, with that cornerstone of the foundational model and then specializing to teachers and then, you know, distilling. It just, you get such big wins in performance across the board. I just, you know, you invest something into the architecture or get better data or training recipes. And then, you know, you invest in that early stage and then it just has massive amplification and ripple effects. So that is, in some ways, is kind of magical.
43:25And then you, I guess, then you see it on the car. and I've had some moments where, you know, a card does something and you look at a log and I've been surprised. Like it does things that I didn't think it was capable of doing. So it's that, you know, it's... When you see emergent behavior, that's kind of a proud moment. One example, yeah. You know, it's, you know, when you build a system and then, you know, you think you understand, you know, how it works and you understand fully, you know, the limits of its capability and performance, and then it does something, you know, kind of almost magical, it's exhilarating.
44:06So one example I can give you, I think I've shared some videos of that publicly in some talks, was this example where the situation that happened in San Francisco, a fairly benign situation, where at an intersection, our light is red, There's new cross traffic. A bus goes by and, you know, it stops partially blocking. Our light turns green. So we start to go. We're nudging around the bus. And then you see a pedestrian being detected on the other side of the bus. And then your car responds appropriately. It slows down, goes a little bit wider. And, you know, then a pedestrian actually emerges from the bus and, you know, we go on our own way.
44:54So the first time I looked at that log, what's going on here? I know we have pretty darn good sensors, and the software is very capable. We don't see through stuff, right? That's not how cameras or lighters and radars work, right? It saw the pedestrian through the bus. It saw the pedestrian on the other side of the bus. And it's not like, you know, you look at the windows, you're like, okay, you know, radars shouldn't do this massive metal box. Yeah. Yeah, you know, look at the sensor data. Yes. And like it just shouldn't, radar shouldn't be able to go through it, right? You know, camera, like you can't see in the camera because, you know, there's reflections and there's people on the bus.
45:33So it's not like you can see through the windows. Right. So like, what is going on? Maybe it's, you know, noise or some coincidence. And I, you know, first time I saw it, I couldn't actually believe it. I was like, no, no, there's something. It doesn't spell right. So what actually turned out was happening is that our peripheral lighters bounce under the bus. And there was just a little bit of very, very noisy reflection of the movement of the person's feet. That was enough for the AI models. Hey, likely there's a pedestrian there. And I'm going to, you know, I'll detect it as such. And moreover, there's enough data there to predict what they're going to do.
46:09It just kind of blew my mind.
46:11Dmitri Dolgov:Is this the perfect example to explain what we were talking about earlier? The value of one, fusion across a sensor suite, but then secondly, building, I mean, relatedly, building an intermediate representation of what's going on, where if you're just dealing with pixels, I mean, the person behind the bus does not exist in pixel space, and so you need to have some representation of the world that exists to be able to reason about the person behind the bus? I think it's an example where giving it using that intermediate representation to boost the level of performance of all parts of the model is what's happening here.
46:58Just imagine solving this problem with a black box, purely open loop, imitative system. be, is it, you know, impossible? No. In practice, what would it take to achieve that level of performance? Yes. Very, very difficult.
47:15Dmitri Dolgov:What metrics can you share on just where the business is at today in terms of rides, revenues, cars on the roads? We have about 3 ,000 cars on the roads. We're doing about half a million rides per week. that transits about over 4 million fully autonomous miles per week. We are operating in a fully autonomous mode in 11 cities in the US. In 10 of those, we have riders. What's the ghost city? The ghost city is Nashville. We just started there. So we just opened it up to riders in four new cities in one day. That was one of those little but super exciting moments where I thought back to the history, like how long did it take us from the first time we started fully autonomous rider-only operation to the first time we had external riders in four cities.
48:23It was about eight years. And then the other week, we just launched four in one day. Yes, yes.
48:29Dmitri Dolgov:It seems now clear that in... 15 years, most miles that are driven will be autonomous. Like there'll be some burn in Puritan, there's lots of old cars in the road, I think it'll actually take a little while. And some of that will be by level four, level five systems expanding in new cities and that expansion continuing. Some of it will be, you referenced existing driver assist systems and kind of getting up to level two and level three and existing systems across current car brands getting more and more capable. What do you think that working your way up from the lower levels versus working your way expanding from existing products like Waymo, what will that convergence look like?
49:17Dmitri Dolgov:Because we're going to eat it from both sides. I don't believe we will. And I actually think this... That's a great answer. Cars will get smarter. There's going to be advances in driver-assist systems. And there is, at the same time, from level four autonomy, there is simplification. And the sensors of today are not going to be the sensors of tomorrow. So they'll be much more integrated. They'll be simpler. There'll be much lower cost. So from that perspective, there is a path of convergence. And there's also a path of convergence from the product lines. Right hailing and what, you know, you can take, you know, a ride through the Waymo app today.
50:02Eventually, they'll be on your personal car. So that I see. And talk about the technology. And I see it just as fundamentally two different problems. There's driver assist systems. And then there is full autonomy. And I think it's deceptive to think of them as kind of incremental, you know, on one spectrum of complexity.
50:23Dmitri Dolgov:Okay, but you think one cannot work one's way up from driver assist systems to full self-driving? You think you have to start building a full self-driving system? I think you have to tackle, if I think about the hardest parts of building a fully autonomous rider-only system, they are very different from what you do for a driver assist system. And of course, some work in this space helps you. I don't want to say you can't make the jump, but it is a qualitative jump. Yes. When can I buy a Waymo so that I don't need to wait for it when I want to go? I can just like, when I'm ready, I can walk out the door and it's there.
51:07I'm not going to give you a date today, but you're not the first person to bring this up as a product request. Duly noted. Okay.
51:16Dmitri Dolgov:I'll add it to the list. Just, you know, that waiting for the car, it should be nice just sitting in the garage there and keep your stuff in it and everything. It's not the first time you've heard that request.
51:27Dmitri Dolgov:So how, it seems to me operationally very intensive and very hard. Like a self-driving car is actually not self-driving. It takes a village. You have all of the human operator ready to step in. And, you know, there was that thundering herd incident that you guys talked about in San Francisco that kind of highlighted that for people. And then there's just like keeping the cars clean and, you know, keeping everything running in that regard. And so can you describe just what the operational infrastructure that sits behind Waymo looks like? Sure. I will say that we are overall, you know, in all of those areas on a path of increasing efficiency and automation.
52:15Yep. So the number of manual steps that one had to do five years ago to launch a Waymo versus where we are today is drastically different. But nowadays, if you look at one of our depots, it's like a fully automatically orchestrated dance of autonomous vehicles. So the way it looks, what it looks like today is cars will automatically go on there to pick up their riders, serve their trips. If for some reason they need to come back, maybe they're low on energy, maybe somebody left a mess in the car, they will automatically come to the depot. right if it is so cleaning today is a manual process right so i'll get flagged in the car you know we have fleet management systems say hey you know car you know number you know 378 needs cleaning and we'll actually uh on the sensor dome we're able to you know display icon so we'll you know show you like a little you know emoji yeah yeah and you know there's you know people whose job it is to clean cars they'll come you know clean up if that's you know cleaning is not required and it's just charging, we'll automatically pull into a charging stall and we'll say, hey, I need charging.
53:43We don't yet have automated charging. In the future you can imagine that being fully automated, but a person will come in and plug in a cable and the car will charge and say, hey, now I'm ready to go. And it will get unplugged and the car will pull out of its parking stall and then go
53:58Dmitri Dolgov:on its merry way. One of the new Porsches, I think it is, has inductive charging, just like your iPhone where you just drive over the charging mat. I was amazed that that works at car scale, but presumably in the future they'll just be able to drive on the charging mat. Or do you think just robotic plug-in will be easier? We'll see. We'll see. I don't know. I think there's some questions about, you know, efficiency and, you know, how that plays into the overall cost and which one will be, you know, most cost beneficial remains to be seen, I think. How well behaved are the Waymo riding population in terms of not living a mess in the car?
54:34We have wonderful riders. We have the most amazing customers in the world. Generally, I would say they are very good. I think there is something about, I talked about not having a person in the car that's not somebody else's car. In some ways, you kind of want to preserve the, I think generally people want to preserve the nice aspects of it.
54:59Dmitri Dolgov:It's a broken window thing where it's so clean to begin with. I know, yeah. It's kind of like, you know, I think that that's the general trend that we see. Right. And it's like, because there's not somebody else's space, you know, you're in it. It feels like it's your own. So you don't like want to mess up, you know, your own space. I think I mean, I don't want to speculate too much on the psychology of things. However, I will say that it varies. And you can imagine, you know, a college town on a Saturday night. And yeah, that's a different distribution. Yes. Yes. Will I be able to get Waymo at any address that has USPS service in the US?
55:36Or will there be some head-tail dynamic where Ketchkin Alaska is just never worth it? Eventually it will, absolutely. There's no doubt in my mind. I think it's just a matter of when and what modality would make the most commercial sense. Is it for you? Right share versus privately owned? For right, it's not a technical problem. I mean, technology is solved. But then if you're in the middle of nowhere and there's just not enough density of the trips, does it make sense for the right hailing service that Waymo is running to have cars on standby? Probably not, right? They can be deployed somewhere else and you probably don't want a horribly bad ETA.
56:17And this is where a personally-owned vehicle that is equipped with the Waymo driver is maybe how you will see it materialized.
56:25Dmitri Dolgov:Relatedly, what will the second order effects of, say, majority autonomous traffic be? Like, it feels like a lot of things will work better where, as you say, you know, when someone merges into a lane very poorly and everyone all the way back has to, you know, slam on the brakes. That's kind of anti-social behavior. And so it feels like higher quality and more pro-social driving will just, I mean, basically reduce traffic a little bit, even for the same number of cars on the road. but presumably there'll be other second order effects like we'll want higher throughput traffic lights and yeah, how else will things change?
56:58So the first thing I think, you know, that you mentioned is, I think that's a huge deal. I just need to think about traffic jams. Yep. What's that saying? The Navy SEALs? Slow is smooth and smooth is fast. That's what like traffic jams are. You accelerate abruptly, then you come to a stop and sometimes you have the traffic like what happened? Well, you know, an old lady crossed the road three hours ago and we still have the standing wave. Right. So if everybody, you know, was a kind of a smooth, predictable driver and a consistent driver and you would still have those traffic jams at the time of, but then the time constant to clean it out, I think would be very different.
57:43But longer term, and you know, things like parking lots, right? Right now, if you look at, you know, what is our most interesting, you know, pieces of land allocated to, you know, it's parking lots, it's garages. And why is that? Well, because again, you know, your car is just sitting there 90 % of the time, right? If, you know, more cars become fully autonomous, then there's no, you know, right. And like, then imagine, just imagine what you can do with, you know, your favorite city in the world if you don't have to spend that money, that huge fraction of it on, you know, just keeping your, these chunks of metal sitting around.
58:22Dmitri Dolgov:Yeah, I don't think people often realize how big a deal parking minimums are for the layout of the urban landscape. The coffee shop here where I am would like to have outdoor seating, but can't because it would reclaim parking spots. Yeah, wouldn't that be wonderful? Yeah. I have a few more questions, but I'm curious to talk about Google's relationship with self-driving, where, again, And it feels like right now, Waymo is, aside from everything else AI-related, kind of the most exciting thing happening at Google. But it was a very long journey to get here. I mean, I feel like you could say that Google almost started working on it too early, because you were saying there's been a bunch of recent enabling technologies.
59:09Dmitri Dolgov:And so did it require Google starting when it did so early? or could one have spun up this project in 2015, 2020? And then how did Google keep the faith when it always felt like it was perennially two years away? Yeah, no, on the latter part, I just have to give credit and huge kudos and gratitude to Larry and Sergey and Alphabet Leadership Center company. Uh, it's part of the culture and the DNA of the company is to have that vision and have the stamina and conviction to go the distance.
1:00:00So to the other part of the question, you know, was it too early? I don't know. I think what we've been seeing, you know, clearly all of the breakthroughs that we've seen over the years have changed, you know, how we're building the system. But the complexity of the problem is such that like you need to go through these eater of cycles, right? It's not, you know, still, and we've seen many waves of technology, right? There's, you know, breakthroughs in, you know, 2013, ImageNet came around. Okay, like, that is the right time to start BSL driving company. And, you know, transformers came around, and, you know, VLMs, and that, and all of those are super powerful, and you have applications in other spaces.
1:00:50In the digital world, they certainly have an impact on, you know, our AI, in the physical world, there are no silver bullets. They drastically reshape that early part of the curve. It's always been the nature of this problem. It's very easy to get started. It's deceptively easy to get started. But it is super hard to go the full distance and get the number of nines, right? There's the standard engineering rule of thumb that every next nine takes 10x more. So I, yeah, maybe there is a more optimal path, but I don't see there's some magical moment where the true complexity of the problem goes away and then you can just take some off-the-shelf components and you're a business.
1:01:36If that were the case, then I think the industry would look very different today.
1:01:41Dmitri Dolgov:Last question I have. You've been promoted a lot at Google. It feels like Google really recognized your talents. Just what do you think Google does? Like Google is famously one of the very best in the world. at technical talent and say, you know, the current AI wave more broadly happening, you know, is either stuff happening at Google or generally Google alumni. But just what have you observed firsthand from how Google does this so well? Yeah, I would say Google, you know, that culture of Google of not accepting the status quo, having a big vision
1:02:28and investing in technical talent and the people who can go the distance and realize the vision, that is part of the culture. I think this is what you're seeing. And with the breakthroughs in AI in the digital world and all of the early investments in transformers and other fundamental technologies, quantum computing, And I guess we're not unlike those efforts as well.
1:02:54Dmitri Dolgov:To be sure. Thank you. Yeah. Thanks for listening to this episode of the A16Z podcast. If you liked this episode, be sure to like, comment, subscribe, leave us a rating or review, and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. Follow us on X at A16Z and subscribe to our sub stack at a16z.substack.com. Thanks again for listening, and I'll see you in the next episode.
1:03:48Dmitri Dolgov:but A16Z does not guarantee its accuracy.
From the publisher
Waymo is now delivering hundreds of thousands of fully autonomous rides each week — but getting there required more than better models. It meant building a complete system for training, evaluating, and deploying a driver in the real world.
In this episode — originally aired on the Cheeky Pint podcast — Waymo Co-CEO Dmitri Dolgov joins John Collison to break down how self-driving actually works today: from sensor fusion across LiDAR, radar, and cameras, to simulation, “critic” models, and the role of AI in decision-making.
They also explore why full autonomy is fundamentally different from driver-assist, what it takes to scale globally, and how recent advances in AI are reshaping the path forward.
Resources:
Follow Dmitri Dolgov on X - https://x.com/dmitri_dolgov
Follow John Collison on X - https://x.com/collision
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
