In short
```markdown
The TWIML AI Podcast - Episode #725
Waymo's Foundation Model for Autonomous Driving with Drago Anguelov
Episode Overview In this episode, host Sam Charrington speaks with Drago Anguelov, head of AI foundations at Waymo. The discussion dives deep into how Waymo is using foundation models and large-scale machine learning to enhance autonomous driving technologies. Key topics include multimodal sensor integration, safety measures, and the evolution of Waymo’s research stack.
Key Takeaways
Waymo's Research and Development
- Background of Drago Anguelov:
- Over 20 years of machine learning experience.
- Led the Waymo research team since 2018.
- Notable previous work includes significant contributions to image recognition and 3D vision at Alphabet.
- Waymo's Journey:
- From initiating operations in Phoenix to providing over 200,000 autonomous rides weekly across four major markets: Phoenix, San Francisco, LA, and Austin.
- Emphasis on user experience and safety, reporting that Waymo is statistically safer than human drivers.
Foundation Models and AI Integration
- Foundation Model Development:
- Waymo is developing a custom “Waymo Foundation Model” that incorporates advances in vision-language models and generative AI.
- The model aims to improve perception, planning, and simulation for self-driving cars.
- Multimodal Sensor Data:
- Integration of various sensors, including LiDAR, radar, and cameras, to enhance 3D spatial awareness and long-term memory processing.
- Challenges discussed include hallucination prevention and generalization across different driving environments.
Safety and Validation Framework
- Safety Metrics:
- Waymo's safety dashboard shows that their vehicles are over 80% safer than human drivers in terms of severe incidents.
- Continuous updates to safety data and rigorous validation practices are emphasized for public confidence.
- Validation Techniques:
- A comprehensive validation framework is in place to ensure rigorous testing under diverse conditions.
- Use of simulation to model various driving scenarios and reinforce learning through real-world data.
Future Challenges and Research Directions
- End-to-End Learning vs. Modular Approaches:
- Discussion on the balance between fully integrated systems and modular architectures for control and perception.
- The need for robust testing and validation mechanisms to ensure safety in modular systems.
- Emerging Technologies:
- Exploration of technologies like 3D Gaussian splatting and diffusion models for creating realistic simulations and enhancing machine learning models.
- Future challenges announced for 2025, including tasks for end-to-end driving models and agent behavior modeling.
Concluding Thoughts Drago Anguelov expresses excitement about the future of autonomous driving technology, particularly as they explore the potential of foundation models. The conversation highlights Waymo's commitment to safety, innovation, and collaboration with the academic community to advance autonomous driving capabilities.
Additional Information
- For complete show notes, visit [TWIML AI Podcast Episode #725](https://twimlai.com/go/725).
- Previous conversations with Drago Anguelov available for context on the evolution of Waymo’s technology.
```
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00There is some unique properties of the problems we're solving. For example, our stack relies on really powerful spatial awareness of everything going on around you, but then you need to understand the world in 3D even better. Another thing you want to incorporate is very long memory over a scene. You need to reason often based on the history of several seconds or more. That can be a lot of frames. You need to think how to prevent hallucinations. So all of this, you know, are still areas that are very much ripe for further work, even though we have these vision language models.
0:46all right everyone welcome to another episode of the twiML ai podcast i am of course your host sam charrington today i'm joined by dragomir angulov drago is vp and head of research at waymo before we get going be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Drago, welcome back to the podcast. Thank you. Pleasure to be back. I'm really looking forward to digging into our conversation. I went back and checked. It's been four years and a month since we last spoke. It was February 2021, and I have got to imagine a ton has been happening on your end.
1:23So I'm really looking forward to catching up on all that. We'll be especially digging into some of the work that you've been doing to incorporate foundation models and that whole idea into the self-driving experience at Waymo. But before we get going, I'd love to have you jump in and share a little bit about your background for folks who weren't listening four years ago. So, well, a lot has happened in the last four years. I'm Drago. I've been at Waymo for almost seven years now. I have been leading the Waymo research team since summer of 2018, and we focus on pushing the state-of-the-art in autonomous driving systems with machine learning or AI, or both.
2:11And I have been a machine learning researcher for over 20 years. I spent eight years at Alphabet working on image understanding with deep neural networks. We published some of the early architectures for image classification or object detection, and we won the ImageNet challenges. That was 11 years ago, 2014. And I used to work on Street View as well, on 3D pose estimation and 3D vision for Street View. As the vehicles used to drive around, we would reconstruct the world in 3D, align all the different trajectories of all the platforms. We would not just have cars. We had snowmobiles, trikes, bicycles.
3:01Sometimes we would put a trike on a boat. We've done all kinds of things at the time. And so you had to have accurate posts so that when you navigate for all these photos and the pancake in 3D needs to look good. Essentially, I led a small team that was enabling this for a few years. And then around 2015, I got into autonomous driving. Initially as head of 3D perception at Zoox, another reasonably prominent autonomous driving company. And I did this for two and a half years. And then I had the opportunity to start and form this Waymo research team that I have been leading and evolving together with the team since then.
3:52And I think soon it will be seven years. That's amazing. That's amazing. I think when we spoke, you guys were early days in the Phoenix market, trying to get some on the ground test experience for the vehicles. Talk a little bit about how far you've come since then. So I think 21, we had the Waymo 1 service in Phoenix. I think it probably was still with the older Pacifica vehicles. And since then, a lot has happened. And just recently, maybe two weeks ago, we announced that we are now giving over 200 ,000 trips to customers in four major markets and 200 ,000 trips a week. So this is fully autonomous, fully autonomous to paying customers.
4:48Right. At a reasonable scale at this point, it is. a significant part of people's lives in these cities and an option that they can clearly use anytime they would like. You can get the app. You can just hail away more in San Francisco, Phoenix, and L.A. directly. And now in Austin, in a partnership with Uber, they need to go to the Uber app to hail a vehicle. and often or reasonably often it would be way more. But in these four cities, we operate at reasonable scale in large territories too. So San Francisco, we have the full city and even a little bit of, I believe, Delhi City. So around 55 square miles in Phoenix.
5:40That's our oldest area where we provided the service and is the largest area, I think, generally in the Western world. I'm not sure in China how it goes, but Phoenix is over 300 square miles. And there we even provide rides to the airport directly. And then we have the other two cities. They're around, again, LA is maybe 90 square miles and Austin, where we're relatively recently around 37, I believe. When we were building it, we were not sure how enthusiastically it would be received and how much people would love being in an autonomous vehicle. I can tell you that at least for myself, and I'm a biased user, of course, because I contributed a lot of, well, systems of features that either are part of the vehicle or help support the vehicle.
6:36I think it's an amazing experience. I think a lot of people also really appreciate the safe driving, the comfort, the fact that, you know, there is significant amount of privacy or there by yourself and safety. As part of the experience, you can play your own music. I think generally people that try it, a lot of them are hooked pretty quickly. And so it's validated the vision that autonomous vehicles can be compelling. They can be safe. The company recently published some updated safety numbers. Can you talk a little bit about that? Yeah, so we share safety studies and also safety numbers of our service as we operate in all these cities.
7:17And we update them regularly. There's a hub that the safety dashboard on the Waymo website where one can keep checking in. And so the most recent update was at 50 million miles. And there we share statistics about our performance as we drove people for the first 50 million autonomous driven miles. Actually, every week we drive more than 1 million miles using our fleet. We drive customers. And so in this 50 million miles, there's different groups of incidents ranked in terms of severity. And so the most severe we track is, for example, incidents in which airbag was deployed. These are fairly serious high-impact incidents.
8:09And we are over 80 % safer than normally human drivers would perform. In our estimate, right, in the cities where we drive, in the conditions we drive, when we adjust for all of this, We find that we are, I think, 83 % fewer crashes that result in airbag deployment than if humans drove these miles. And that's a non-trivial amount. So if you're given the miles, that's over 60 crashes that we could potentially eliminate it by having Waymos drive this. And the second category of incidents is ones where there was some kind of injury, right? And that one, we're over, again, 80%, 81 % safer than an estimate of what humans drivers would, given our domain where we drive.
9:08And the third class is ones that are actually police reporting. So more minor incident, but still often you can report to police. That one, we're about, I think, 68 or so, 64 % more,
9:26how to say, fewer than our estimate of what humans would do. So these are significant improvements. And also, separately, we've had studies done by insurance companies, third party, that also confirmed that in terms of like property damage or you know injury claims we are about 80 percent less than what humans would do for comparable amount of driving that Swiss Ray is a company so uh this is another separate benchmark not just ourselves benchmarking us but we will continue releasing this but it's it's I mean that's what a lot of us are in the space, not just for the experience. I mean, I think we're always, many, many people are the joint autonomous vehicle domain are motivated by the great safety benefits and many other benefits that these vehicles can bring to the communities that they serve.
10:25And so, but this one's the main one. There's tons of people that, I mean, safety comes first. Our mission that Waymo is the world's most trusted driver, right? And they underpin everything we do. Is it still the case that the vehicles are to some degree remotely operated either on an intervention basis or otherwise? And do you publish stats around that and the number of incidents per mile or something along those lines? I think in terms of can a vehicle, can an operator inspect the vehicle? I think that is possible, right? But it happens rarely. And I think the fact that we've scaled to such amounts of vehicles and such scope, you're not going to do it.
11:11There's people watching over every vehicle, right, all the time. That's just not a good product. You know, what did you eliminate? You moved the driver to sit behind the workstation. That doesn't work. And so this is mostly in very rare cases. and also the vehicles, just to be clear, they're not teleoperated. So it's not like someone drives them remotely in these cases. At most, occasionally a person can confirm something the vehicle wants to do and usually the vehicle already needs to be stopped for this to happen. Yeah, it seemed like having a real-time loop for that kind of confirmation would be very difficult.
11:50Yes, and you do not want it. But then that means that you actually need to understand and deal with these issues, at least to understand to stop for real, right? Which is, you know, again, very rare. That's also not even a good experience. You don't want to do it. So while it's possible, it's very rare. So tell us a little bit about how, from a research perspective, you've been tracking and thinking about incorporating all of the ideas around foundation models, LLMs, VLMs that have sprouted up since the last time we spoke? So, yes, generally, we look at the state of the art in machine learning and AI, our team, and it has tried to adapt the advances to our domain.
12:40And usually, it has involved some creative adaptation. So there is some unique properties of the problems we're solving. and traditionally we had to do adjustments or fairly creative adjustments of the techniques we see right now every two years or so my experience in this space and I'm in in it since 2015 actually in December maybe early 2016 is there's big technological leaps and a lot of them have to do with machine learning or AI and I think generally in this space you want to keep keep reinventing your stack. So as these advances occur, you digest them, you want to be at the forefront and take the benefits.
13:29And I think what's happened in the last two years, if anything, there is even bigger jump than before with these multimodal large language models with generative AI technology that also spans multimedia, obviously images, video, understanding, audio, handling, all of these things, it's been a bit of a revolution. And, you know, we are keenly aware of the advances and we have been exploring them for the benefit of our driver, right? And so we've engaged in understanding and leveraging the technologies such that vision language models, right? Such as diffusion prediction for generative outputs that can include videos, but they can also include like generating driving scenes for testing we have done.
14:32We leverage things like 3D reconstructive advances. There's a lot of very interesting new technologies are relatively speaking. Like Gaussian splatting and that kind of thing. Consistence plotting before that nerve, right? These kinds of technologies are also very interesting. And we've been working on them. But generally, right, recently there has been these trends to scale the models and, you know, train them on a lot of data, scaled architectures, over time, ad reinforcement learning, right, to target even better reasoning capabilities. All of these things, I think that's the general trend in the space.
15:18You can see it with all the large models. We have been scaling our models for a while, and we've had actually transform architectures for a while. If you see our publications, we had a few even externally published around 22, 21 with transformers even before that for certain applications in perception. And so we have been on this trend, but I think what has changed more is just the understanding how much more scale we can leverage. And so now we have gone and scaled the models and evolved them beyond what normally we used to do before. And we have seen clear benefits from doing it. Now, just one thing I want to describe is why is our domain different than the vanilla vision language model?
16:16And by the way, speaking of vision language model, we have, for example, experimented with them. We have even published a paper called EMMA last year where we took a state-of-the-art multimodal large language model, what Gemini people like to call theirs and tried to adapt it for driving tasks. And adapted it reasonably successfully. We got some very nice results on the driving tasks. And this is the idea of converting trajectories into tokens and passing them through the VLM? So the way it works mostly is, okay, so what does Gemini bring to you? It brings to you world knowledge. It's pre-trained on the internet and it's a large model.
17:01So it understands a lot of concepts already that you don't have to teach it directly with your own labels or data. And then you fine-tune it on your tasks. So now you bring your own data and your own tasks, and you adapt it to do well given all this knowledge it brought to the tasks that you fine-tune it to do. And, of course, you need to adjust a bunch of recipes. You need to make sure that the tasks are well-represented in the Gemini inputs and that all the sensors are properly scaled and the right data sets are created and so on. but then when you fine-tune it, right, it does a reasonable job actually on these tasks but there are still limitations.
17:38For example, PowerStack relies or likes to rely on really powerful spatial awareness of everything going on around you. It's helpful for safety but then you need to understand the world in 3D even better, right? And we have additional sensors like LiDAR and RAIDA that are really superb for contributing to 3D spatial awareness that you want to incorporate, which a vanilla Gemini model does not do. Another thing you want to incorporate is very long memory over a scene. You need to reason often, you know, based on the history of several seconds or more. That can be a lot of frames. You need to incorporate this memory, somehow maintain it, somehow structure it.
18:19You need to think how to prevent hallucinations, right? So the model can predict things, but every once in a while, if you go to domains where it has not seen things, it can be wrong. So you need to think, okay, what are the way we have actually a lot of experience, you know, launching machine learning, powerful machine learning in every part of our stack, but doing it in a way with the stack design that is built to also mitigate all these issues of falling off the data manifold, hallucinating, and so on. So all of this, you know, are still areas that are very much ripe for further work, even though we have these vision language models.
18:57But we want that power. We want to bring external knowledge, right? Internet has a lot of knowledge when you're pre-training. There's a lot of value in having these kinds of pre-trained components. And we also can benefit from some of the architectures and scales. What are some of the specific tasks that you've explored using VLMs for? I mean, for example, they obviously can be explored for understanding all kinds of perception outputs, right? So they understand semantics very, very well. They're trained such that you show it an image, you can ask a variety of questions. And a lot of them are really powerful on all kinds of even fairly obscure details.
19:44That traditionally is a property that you want to benefit from. because otherwise in the old days, maybe five to 10 years ago, you would use humans to label all these concepts. Now that model, for example, would understand a lot of them out of the box. That's very powerful. All right. So the idea being the model sees the scene and as opposed to the old world of bounding boxes around humans, balls, scooters, that kind of thing, the model just knows what those things are and that they're important. Yeah, and without us having to label all these concepts. Right? So you can directly bootstrap yourself.
20:27The other power of all these models is the more scaled they are, the larger they are, the better they generalize. So if you see some new variant of the concepts you for some reason didn't label that you need to handle. Before, you would have to label a lot of examples for the small model to really capture what those things like. And then every new, particularly different looking thing, you'd need to label more examples. Here for perception, mostly I'm thinking here as I'm describing. Now the large models are just like, I mean, you can take Gemini or a lot of these other external vision language models.
21:05They mostly would understand what it is already. And it's a big model and it generalizes again. When it sees a trench, it can see a trench in a whole bunch of conditions and know it's a trench. For the small models, this is Harvard. It requires more bespoke engineering to ensure. And so that's a nice benefit, right? Now, the scaling part is appealing, but there's also this interesting limit. There's this question, how far can you push scaling? And how much of this knowledge can you bring in our models such that now we can get data from new cities, from new countries, from different software platforms You can move the cameras.
21:46You can change the lighters on the vehicle over time, right? Which we do. We have a six-generation driver that, you know, keeps evolving the sensor stack, which is then something to think about. So when you have all these properties, the question is, can we have a model that just generalizes across all of these things? We want to scale to dozens of cities. We want to do that much bespoke work. And the best way to scale a model is let's try to scale it in the data center first. So remove constraints of real-time response, which needed to be for safe driving. Remove the constraints even potentially to only see events as they unfold.
22:26You can see potentially how in the future. Remove constraints on how much you can scale the model. Now, how much understanding can you get out of this model? How much can you generalize? And so we have been pushing to understand this question, And I think we have been aiming to build what we call this Waymo Foundation model, which is not just a vanilla VLM adaptation, but we want to bring in all these capabilities I recently talked about that expand on the standard vision language model and build it large, build it in the data center and see how much can you just understand all the data that we see and we collect.
23:05And the better this does, the more straightforward it is then. This model then acts as a teacher to the models in the car. You can distill it down. And of course, you still need to do this with a thoughtful design, ensuring that the onboard stack fits all the safety constraints, all the latency constraints and so on. But then this thing is just the source of knowledge that you can always ask and both to mine data, to understand data in your mind. And there's a very interesting question, how far you can push it. But the more you push it, the more you accelerate the scaling or seamless scaling of the service, right?
23:48So we want to explore that. I think that's a very exciting direction and it's in tune with the time today, right? I think it's the age of these types of models. And it tackles one of the historically challenging aspects of autonomous vehicles, and that is that generalizability. Like the vehicle, you've got lots of data captured about a particular road during the day in clear weather, but the vehicle needs to operate in that same road at night in bad weather. historically you'd either need to capture that data manually or use some kind of data augmentation scheme to train the model so that it can adapt to those variations.
24:40But if you're able to achieve sufficient generalization with a scaled up model, that eliminates, I would think, a lot of the challenges of getting coverage in the markets you're in and and getting into new markets. So the reality is we have good generalization and coverage in the existing markets already, right? And we operate on them clearly and have passed a certain bar. But new market, and the more different the new market is, or the more different the new market and sensor stack is, the more it pushes your ability. And then, you know, you can do it a lot faster if you have this kind of capability than the process today, which would, of course, involve, you would go explore the market, collect data, label some data, understand anything that may be different, make sure you're confident that you're safe, right?
25:36Like there is a standard procedure how we ensure safety in the market, I think, which is important, but we could accelerate it. How long do you typically need to be testing and collecting data in a new market in order to launch in that market? I think it's been a very much evolving process, meaning the first time, say, we did Phoenix, we did East Chandler first, which is a bit of a suburban area of Phoenix. And suburban areas have one set of challenges, for example, unprotected left turns on the freeway, or not a freeway, but a reasonably high speed road, 45 miles an hour, and then people speed.
26:18So it can be a lot higher. And it's really important how you can go in front of traffic going at such speed. And there's a few other challenges there, but it took a while to sort out, right, and make sure our system is robust. Then you go to San Francisco and you see, oh, it's totally different. There's all these crowds and totally different behavior. And the streets, big part of the city, are a lot narrower, right? And it's just different dense urban scenario. So there it took a while, too. and then we go to Los Angeles which somewhat combines the properties of the suburban Phoenix and by then we had moved more to downtown Phoenix and so on so we had a bit more of a span everything from a bit more suburban to urban to dense urban and you have San Francisco so when you combine those two LA actually was fairly straightforward right so as we cover it's about a bit of covering the design domain, the more you cover it, right, through, and in our path we've been doing this, the easier it is for the next city.
27:23They generalize. So every new part you collect, it helps you with the next part. So it's an accelerating process too. Now, separately, what do we work on a lot today? Freeways. Freeways offers a whole set of new challenges compared to the ones we've seen in, say, the more suburban areas or dense urban, there's different things to be concerned about, right? And you need to validate, for example, on freeway, you need to be robust at 70 miles per hour or 65 miles per hour to any kinds of failures that can occur. And also you need to worry about, you know, stopping at any position on the freeway that's inherently unsafe, right?
28:09And sometimes there's not even shoulder for certain parts of freeway. So you need to think what do you do there. And that's a different bespoke work you need to think about. And another one, of course, would be snow ice. That brings yet another dimension. There's a few of these. So I would say I can't tell you an exact estimate because it really varies. But we have a process, a robust process, to ensure that we gain sufficient confidence in the market before we launch. In the markets that you're in now, do the vehicles operate on freeways or are they only on surface roads? So for most customers today, we do not provide freeway, fully driverless on freeway.
28:52But we are testing on freeway for a while, actually, without driver in some capacity. So it's still an area which we are scaling and working on. And so we, of course, would very much like to provide this to our customers, but it's still an ongoing process. For example, take Phoenix. We're clearly motivated to the freeway because Phoenix, for example, is 300 square miles. Going from one end of that area to the other, you really want to take the freeway. Otherwise, your service would take a while compared to what is possible. Well, you want to provide good service. So it's clearly an area of high importance.
29:36So kind of going back to the foundation model work, when you think about this challenge of scaling up the foundation model in the data center, one of the challenges that you mentioned is incorporating all of the different sensor data that you might want to incorporate that aren't necessarily native to VLMs and to images. How do you approach that problem? So it's a very interesting problem, actually. It's a bit open-ended how to do it best, right? And I think maybe one way to think about it is also audio is an example of this too, right? So initially when you do a model, people try images first, and then, you know, later audio became a domain tool.
30:28So you want to build a multimodal model. Now, how to actually build it and how to insert data from a new sensor into a model that was trained potentially or pieces of the model that were trained without it, it's actually an interesting research problem and there are many approaches. So I can't tell you how exactly we do it. We're experimenting with several and learning along the way. But we find that it's one of the key questions in our domain how to do that. is building a foundation model on top of existing VLMs. Like the part of the motivation for that is taking advantage of the world knowledge that is embedded in that.
Read the full transcript
31:03But the alternative might be incorporating, I'm sorry, training from the ground up a foundation model that understands these different modalities and, you know, is able to do the things that you want to do. So you can train what exactly that would look like either. Well, exactly. There's this interesting question of, okay, so let's say you train something fully from scratch yourself. But then you need to worry, well, how do I bring all this knowledge that people already usually put into this from the internet? Right? So you still want to somehow then maybe need to get that data? There's still some kind of fusion of modalities that needs to come together.
31:50It's just a different way of looking at the problem. Right. And so there is another thing, though, we can do, which is we, and it's also expensive, of course, to train the model fully from scratch, right? So you need to think that's one aspect of things. The other aspect is, okay, how can we learn and build a model from our domain, right? Using data from our domain directly separate from the fact, okay, one is you leverage, hopefully, world knowledge from models one way or the other, right? You need to think how. Maybe you can even use them to label data for you or you can take pieces of them.
32:25There's many things possible. Now the question is, how do I create, how do I train model in our domain to understand our sensors, right? And there's this interesting question. You can learn in our domain to predict the future also, right? So how do you pre-train large language models? You predict the next token in language. What can you do in our domain? You can do things like you can predict what the sensors give you in the future. So imagine if you go to tokenize your sensors, you can predict the next sensor token and the next sensor token. And if a model is able to predict accurately how sensors look in the future, whether it's LIDAR or camera or radar, they need to learn a lot about the environment because it will actually capture all the knowledge in here and to predict how things evolve so you can imagine how you will see them.
33:20And that's very powerful pre-training. This is equivalent to this kind of pre-training next word prediction. The issue is only, of course, a word can be tokenized in maybe 100 ,000 tokens or something to that effect, and then you're predicting a distribution over 100 ,000 discrete tokens, that's fairly straightforward. Now, our sensors, they're higher resolution. They're different multimodal. Now the question is, well, for those sensors, what's the vocabulary, right? How do you practically predict it? And that opens a whole bunch of interesting questions, right? And it's a question, do you need autoregressive models?
33:58Do you need diffusion-type predictors? How do you set it up or a mixture of those? What is effective? There is work in the space already in academia trying to build this, and in the industry already trying to build what it's called world model. You're dreaming the futures in the world that you are modeling, right, from your own sensors. And that can be useful to also train large models, right? But how to best build it is also a very fascinating question. And of course, the richer futures you dream, the larger models you need. because now you need to capture a lot of the nuance beyond a certain level of, oh, can I scale some model to just predict what I should do and maybe some high-level properties of that is versus predict everything in super high pixel, lighter scale level fidelity, right?
34:52You need potentially larger models to do that well. And so there's trade-offs, like how much of the compute now you're going to spend to pre-train these very large models, how much of that actually results in driving improvements, say, for you in the end? And how does the math work out? Is it compelling enough? This is also a very interesting question. But for good or bad, some of these world models are needed for the simulator. So you need to think how to build them regardless, right? We want to... So in our domain, it's not just about building the driver. There's two main problems. One is build a driver, right?
35:29And I talked to this also about this in my GTC talk. So the other problem, which I know is easier, is to validate it and ensure that you are confident this model or driver performs well. In the vast majority of cases we need to handle, right? Like the environments in which vehicles drive span all kinds of diversity and season. As we already talked about, different operational domains. There is all kinds of human behaviors. Some of them very rarely seen that you need to deal with. in front of you. You need to interact with humans. So all of this, of course, is great to test in a simulator and is a great tool to test at scale.
36:12But now the question is, how do you build this simulator? And now you want to use machine learning to build it, right? Before we dive into validation, I want to dig into this idea of predicting future sensor values. When I think about that, I think that there's a causal relationship between the vehicle trajectory or the vehicle behavior and then what the sensors pick up as the vehicle progresses through a scene. And I'm trying to kind of wrap my head around what it means to predict future sensor values as opposed to predicting the thing that is the root cause of that, the vehicle trajectory. How do you think about the relationships between future sensor prediction and the vehicle itself and what you're using as a control input or something.
37:05So I just contrast the two a little bit and see if we're thinking of the same setup. So in our Emma paper, for example, right, you have a vision language model. And one of the tasks we can teach you to do, even though we tried several, like you can predict the road graph that's in front of you, you can predict the 3D boxes that you see. But one of the interesting properties, the main one maybe you predict is what's the driving trajectory you're likely to follow. and you can of course sample several and the start of that trajectory actually then you can turn in the controls the rest is more speculative kind of showing you what it's likely to be like over time right and of course the longer you predict the trajectory by the way in our domain it's the more uncertainty it gets whether you're actually taking it because it also depends on certain assumption of the world kind of like next word prediction or next token prediction.
37:56But like, you know, I can think I'm driving this trajectory for a second, but if the vehicle in front of me swerves, I'll adjust it. So it's all relative to what the others do a little bit as well. That said, it's a very good output and the start of it, you always would drive, right? So it's related to planning quite strongly. But the trajectory, how do you train this model, right? In our domain, you would actually observe either our drivers or potentially how others drive. And you saw what they did in the future. And you collected it. And now you're teaching the model to predict just those tokens, driving tokens for a few steps in the future.
38:36That's a signal, right? That's the most topical signal for how you should drive. You watch how people drive, you see their tokens, and you learn to predict them. Now, when you think of sensors,
38:50it's essentially the environment that goes with how you drive. Instead of predicting just how you drove, now, for example, maybe a next level. Imagine is observe others and predict how they drive next to you. A lot of the models do this already, right? So not just predicting yourself. We have a model called MotionLM published maybe two years ago, and it's in even all the work, showing how to turn motion into conversation. And you can train the models with that level of abstraction. So you say, oh, I'm not just teaching our model to predict how I drove. I'm going to teach it also to predict the other people's motion tokens around us.
39:29And so you predict jointly how a whole group of them now behave. So that's a richer signal. You're teaching the model more things from everyone, example. And they're highly topical. Now you're teaching it, oh, not only how I behaved, but how everyone else reacts to me. And then maybe how I would react to them. You're modeling what we call the joint distribution of at least some group of agents around this. So that's a richer signal. Now imagine you go to sensors. That's maybe very high bandwidth, detailed variant of everything that happened before. You not only see the environment not only changes because you, say, did certain driving paths.
40:16It also changes because others reacted to you and changed their behavior relative to you. That's also reflected in the sensors. And beyond that, you need to think how everything actually changes appearance as you and how objects may look from the other side if you surround them and how reflections may work and how certain LiDAR intensity relates to the RGB pixels. Because now you're dreaming the future of the environment. the full environment in its full fidelity. So there's an argument that it's like a more grounded prediction and that it's reflecting the actual environment that the sensors are intended to see.
40:58It's the most high bandwidth prediction you have. This is all the information you actually had about the future, which is your full sensors. And that's the most you could try to predict. But then the issue is, are you overdoing it? So, for example, humans, I don't spend my time thinking how the tree looks from the other side. when they drive, right? Right, right, right. Or like, you know, as they go, how would the branches occlude each other from a different viewpoint? Or how the shadow will evolve on that pedestrian? I don't necessarily have to think about it. So while there's a lot of knowledge inherent that the model can capture, not all of it is relevant to driving.
41:41Some of it is. And so now there's this interesting trade-off. Do you really want to go all the way there? how much benefit you're going to get. Or do you do something? Clearly, some version of fidelity is relevant. So like predicting what people did a few steps in the future is the minimum you can do and the most topical. And then you broaden and it becomes increasingly less relevant. Like, you know, you want to model how the clouds move in the sky. Maybe it's relevant to it raining in an hour, but I'm not sure it will help your driving suit. So there's this trade-off. There's a ton of signal in these sensors, but you could choose to model it.
42:19It's helpful in some cases for a simulator to model it, but it doesn't always help driving. So we're discovering what is the right risk. Remind me or maybe talk a little bit about the way your approach to kind of end-to-end training versus subsystems that maybe map to traditional approaches to robotics or vehicles. vehicles like you know are there separate you know control perception etc systems or are you moving you know towards or away from end-to-end modeling I vaguely remember us talking about this a few years ago as one of the kind of differentiators of your approach but I don't remember which side of that divide if you want to call it we're on the practical side we will take the thing that works best, right?
43:18I think that's what I would say. Now, there is end-to-end in general is a bit of an overloaded topic. I know it's been made into a big issue by a lot of people. It's compelling in some embodiment, right? But let's untangle what end-to-end means, right? So end-to-end is, if you think, there's kind of mild definition and a strong definition. So end-to-end means typically, right, the strong definition is, hey, you're going to pass some gradients from, you know, sensors, from controls, what you want something to do through a whole system, maybe to the sensors. And you'll, along the way, learn features that you may not necessarily be able to inspect yourself or that have semantic meaning, they just emerge from the training process that may help you drive, right?
44:19End-to-end, in general, is a training strategy. It's orthogonal from whether you have modules in the stack, right? You can still have modules and train things end-to-end, right? If the modules are connected properly and you can pass gradients between them, you can, right? and you can still train many tasks end-to-end. So it doesn't have to be sensors to control, right? You can even potentially consolidate large pieces of the stack. So every piece of your stack became more end-to-end because initially, say, you had bespoke 10 different models handling certain tasks, and now you consolidated them.
45:00That part became more end-to-end, all right? Now, so it's a complex topic. Now, I would say that what we've learned over time is that you want few large components if possible. It simplifies the development, right? It allows you to essentially scale those components more and optimize them as opposed to have a bunch of little custom things, each doing different stuff with different data generation. So you want few consolidated components. And now the question is how few? And of course, the other part, you know, is if you have a few, you still want to ideally backpropagate through them. But what's the, how few and can you just make do with sensors to controls the extreme?
45:49And the sensors to controls, I think, is what people are really excited about in academia. And it's very powerful in an academic setting. And it's a compelling idea. Just collect enough data. You've got some black box in the middle that's sufficiently, you know, has a sufficient, you know, scale and number of parameters. and you've got some desired output and you let the middle part figure it out. And our EMA paper does this too, right? So we do train it end-to-end clearly, right? And predicts nice trajectories and generally does a good job with a fairly straightforward recipe. So that's appealing.
46:24But the big questions there is can you simulate and prove that this system does not hallucinate, does not... Provability, controllability, introspection, testability. How do you validate the system? Now I need a simulator that's always from sensors to controls. And in our domain, there's quite a lot of sensors, and they're not that easy to simulate. The technology for this has been evolving. And then you're really tied to this. You need to scale it. Imagine driving over a million miles a day in the simulator. and now having to simulate everything you could possibly see in these sensors. And usually most companies have at least a couple sensors, different ones.
47:13And you need to, at a million miles with over a dozen cameras, with a few LIDAR, for example, or radar, right? And it needs to be really realistic. Otherwise, it doesn't count. How do you do it, right? That is the challenge. How do you prove to yourself now that all these cases can be validated just with this setup? That's hard, right? And so there's push and pull also the other big questions. Let's say you train a huge model. And now you want to change it. Well, what do you do? The only thing you can do is you change the data to retrain the model. Now you fixed it, but maybe it broke something else.
47:53How do you, it's really hard to fix something really quickly if all you have is a model that goes into it. So there are all these considerations. So I would say that there is consensus, reasonable consensus in our space that you need a few large components, but there is no consensus yet if it should be one component. And I think a lot of it comes from the stability. And so if you want to say, in our case, when we're doing actually fully autonomous driving, it's top of mind. You design your stack in a way that you're sure you can test. So that's a constraint. So we need to, anything we do, we're doing with the mindset that, hey, we actually know how to test this at scale to actually deploy our driver.
48:35When you're doing driver assist, it's less demanding. You don't have to test it to this extent. So maybe there you can get by with a fully end-to-end model, as some people have. But in autonomous driving, when you have to release it there to drive, you know, hundreds of thousands of trips and millions of miles a week, that's a high bar, right? Between that and ability to fix things, these are top of mind when you build a stack, when it's fully end-to-end. So you need to have an ad. I think so these are the trade-offs and we are we're doing our best thing I think I think it's appealing to do end-to-end generally it learns good features when you can do it just prove to yourself you can validate.
49:21Does the more modular approach that you've taken give you the controllability testability introspection you know quote-unquote for free or is are those all active research topics as well like how to um how to ensure a degree of um i guess explainability was the word that was briefly escaping me explainability and controllability you you want if you have a model and it's some some big vlm even or whatever like we did with emma You want to have a way to detect and steer it when it's doing wrong things and it's hallucinating. So you need a harness around it. There's two kinds of harness you can have.
50:12You can build a harness around it that operates in the system you deployed, that makes sure with, you know, there's ways to do this. We have experience. You can make sure to catch this model when it hallucinates things that it should not do and control it on top. The other part is you can catch it in the simulator, the other part we discussed, right? But then you need to build a fully realistic simulator. And in both cases, you need to understand the world a lot deeper and more symbolically than just having sensors to control. I would say that. And you need to somehow get the confidence that you've run sufficient simulation samples in order to get the coverage you need.
50:57Oh, that's a whole science, actually. It's extremely hard. And I think there's this other interesting aspect of the simulator. So, for example, as you collect more data and make your model larger, you can have it to predict better trajectories for you. So it's very efficient in some way, data efficient too. If you just want to improve over time and you get more data, you can keep adding data and the model will improve in certain ways, at least over time, in general, right? when you want to test it you cannot do that like you actually need to think think of all the cases that could go wrong you need to create some ways to validate all of those they can require all kinds of hardware failures sensors going wrong rare cases you name it you need to somehow handle all of this in your release cycle that does not work as well like you just take data and you just feed it in the simulator and the simulator tests you better.
51:55It's also true, but it's not sufficient. So you need to be very thoughtful on the testing side. It's not purely a data-driven exercise. You need experts thinking what could go wrong and working on mitigating that. And it scales less straightforward than something where you can just put data in and driver would improve with more and better data suitably picked. That is true. but your ability to test it is not similarly scalable. How do you characterize the test coverage or the testedness of a given model? Do you need to get to approvably, you know, some standard of, you know, proved valid or is there more of a statistical determination of testedness?
52:47so like i think you need to be very comprehensive and that's why it's hard to do just it's not just some data driven model that you just do throw data in it somehow tests you but you're also not trying to get to like nasa levels of mathematical proof of uh you know complete controllability or are you i guess i mean there's also this question how do you even prove it right because uh if people do very unreasonable things some of them you cannot even mitigate right like you're stopping and someone tailgates you at full speed what are you gonna do you can't even get out right uh so it's all relative to what what's reasonable to expect but i guess if we have published our safety framework we're one of the fewer companies that have done it and uh that details the whole set of tests we subject our driver to when we want to release it And it's a very broad set.
53:46And it includes both kinds of hardware testing and failure of injection and, you know, kinds of standard analysis techniques. And it includes simulator and it includes driving in closed courses and testing with drivers and all kinds of things, right? Like we, this is actually difficult to develop. It's very comprehensive. it takes a lot of effort, but we must show that we can iterate over time and release improvements to our stack while going through this validation cycle. And I think that's one of our great things that we've been able to accomplish is to have that, right? That's it, it's bespoke, it's diverse.
54:32Many things are combined, but if you ask me, what are we testing for? You need to test for a lot of things. One obvious one is you should feel you can be safer than the human driver. And, right, that's one. Second one is you should be confident you don't get stuck all over the place either. The third one is, okay, you don't want to, you know, you want to be thoughtful and mindful around, say, ambulances and police vehicles. It goes on and on and on. The amount of things you need to be able to do is broad. You need to, in construction, for example, or if directed by an official with some gesture, you need to try to understand it.
55:14It's a broad set of cases. You need to check the chicken hand. And going back to the question about data only and end-to-end versus, you know, we didn't explicitly talk about like a rules-based, like how do you... incorporate that long list of scenarios into the model and uh you know just throwing it into throwing more examples into a data set isn't good enough like what how do you do it and there's different ways right i think generally we have a lot of experts right as well into what is good driving motion planning experts. That's one. Second is, there's generally this question of how do you define a reward function for driving?
56:09So what's good driving even? How do you define it? How do you ensure your stack satisfies it? It's a very interesting question. I mean, so far we discussed maybe imitating some drivers. That's one way to define it, but that's not the only way. You can say you should not do a whole bunch of things, so you can read this rulebook rulebook and ensure that certain things are true and optimize for it, right? But I think it's a difficult endeavor to define it and then make sure the model take that into account. It's an evolving space. We have reasonable approaches. I think generally, as you can see, something's easier to define than others.
56:52For example, don't hit things is somewhat straightforward. I'm thinking of like I was just yesterday, I was on a road and there was some work happening to a transformer on the side of the road. So they combined the two lane highway into one lane and they had signal people on each side with either the slow sign or the stops, you know, the sign that spins from slow to stop. Like, and I'm thinking about in the context of this conversation, like, is that something that you just get enough data about those situations and the car figures out what to do? or do you have some rules set somewhere that, you know, kind of defines for the car, like when this is happening, like this is what's going on and this is how you're going to behave or does it follow more readily from what the car in front of you is doing than I might think?
57:38It's complicated because a lot of those things as well normally shouldn't cross, say, the lane boundaries. But when they put the cones, you actually should, but then separate from the cones and the sign, the human actually gestures you to do something, right? That's what makes it very complicated. Yeah. Yeah. So we combine different approaches to make sure we can handle this as well as we can. It seems like an example of where, of at least a failing of a fully end-to-end process. you need to collect a lot of data about a lot of, you know, what are probably, you know, unique, individually unique situations that have this pattern in order for the vehicle to, on its own, do the right thing all the time.
58:28And it's limiting. One thing that, yeah, so if you did not see some situation, right, the model may just nail it or it may not, right? And that's a challenge. Now we need to know, does it nail it often enough where you say it's good enough or does it fail often enough where now you need to think of other things to do? That's part of the, you know, convincing yourself, validating the driver's talk. You mentioned earlier, we talked about NERFs and 3D Gaussian splatting and diffusion models. You mentioned those as technologies that you were looking into. And a lot of that comes up in the validation part of the equation.
59:13Can you talk about some of the ways you're using those technologies? So I can say the following. I mean, so far I was trying to tell you, hey, validation is difficult to scale. It's complicated. You need to handle all kinds of individualities and some of them need experts. It's true. But the best way, one of the best ways to scale validation is still to use machine learning to build the simulator better, right? And because simulator is one big part of your validation story. And now you need realistic scenarios and interesting scenarios to be played in this simulator. And traditionally, maybe 10 years ago when I started, the state-of-the-art in simulation was computer graphics methods.
59:59So you collect a whole bunch of assets, houses and trees and, you know, whatever, and bollards. And then when you drove it, you understand those assets, where they are and how they look. And then you just replace it with the assets. You place them there, ideally automatically. And then you have a world. You have a 3D world. Now you can drive in that 3D world and you can use computer graphic technology to make it look realistic. But even then, there was a sim to real issue where the models wouldn't react to the graphical versions of the things in the same way that they'd react to the real world versions of the things.
1:00:35It can easily be. And this is still true in a lot of robotics. Now, it's true that they're correlated, right? They look similar-ish, but they're not the same. And so why? Because, I mean, honestly, computer graphics, again, is certain approximation and who said the properties of the, you know, what should the sun be like and what is the diffusion aspects of the environment and is there fog? Like you need to set a whole bunch of knobs properly. What is the camera, you know, detailed? Is it grimy or not? And you need to somehow need to stop all of this. It's really hard in a computer graphics environment.
1:01:16There's a lot of knobs, even though it's nominally powerful. Right. And so then there is this gap. And also, how good are you at taking what you draw and place all these assets? Did you have all the assets? Did you not have them? If you place different assets, if you set the conditions differently, is that, again, the gap is there. So you're not guaranteed that you will be able to well reconstruct anything you drove. You can drive at night, in the fog, in the rain, and things may look different. Now, the dream from, you know, scalable simulator is to be able to just reconstruct it from your own sensors.
1:01:56We have enough sensors, they're extremely rich, they're, you know, cameras and 3D sensors. you should be able to go a long way into just building yourself the simulator from what you drove and so the the dream is right and that's something we work to enable is any situation i drove with my car i can build a simulation environment from i may have to populate it with additional traffic suitably that's where generative ai comes in i may need to evolve it in response to the things i do because my decisions impact the decisions of others they respond I need to model that correctly. Otherwise, you're not going to have realistic outcomes.
1:02:34If the agent drove like what it drove when you weren't in front of it, it will plow into you just because you did it nebly. You can't just replay what the agents did nebly and get all the signal you wanted. So you need intelligent agents too. So between reconstructing the world and being able to drive, say, this intersection in San Francisco, I want to be confident I can drive and being able to populate it with agents that behave like it is reasonable for, you know, pedestrians and vehicles to behave. That's a really interesting, challenging machine learning problem. And so I'm inspired by it.
1:03:08It's fascinating. I think some of these technologies are 3D Gaussian splats, diffusion modeling. 3D Gaussian splats and NERF are the reconstructive side of things. So you can reconstruct really well a representation in the environment you draw, but you can't, those technologies don't necessarily allow you to augment it or dream it or turn day into night or place new agents so easily as it would be now, for example, something like a diffusion model. So that's more of a generative type of model. You can dream new parts of the environment or agent behaviors, a complex one where the appearance matters or futures.
1:03:45But, you know, again, do you dream freeform or do you want to condition yourself and land yourself at some specific intersection in San Francisco? How do you do the two together? There's some interesting questions here that I think the field is now making progress on. So it's exciting. And we are alongside the field as well, making some, I think, exciting progress. One of the ways that you make some of this progress is by offering challenges to the broader community to help you solve some of the things you're focusing on in a given year. And you recently announced those challenges for 2025. Can you talk a little bit about those?
1:04:26Thank you for mentioning this. We've been engaged with the academic community since 2018. We started working on this when I joined Waymo. So in 2019, we released a dataset at the time, the largest, most comprehensive we had in terms of, you know, actually having high quality Wiimuth sensors and quite a few scenarios compared to what was reasonable for the time. And we kept upgrading it every year, adding more and more either tasks or potentially additional examples of interesting driving or things that need to be understood. and in parallel with dataset work, we organize challenges on tasks, which we believe are exciting and there's headroom to make more progress in terms of interesting approaches and solutions.
1:05:18So this is our sixth year. We actually have not announced them yet. We probably will announce them by the time this podcast airs. They're launching, supposed to launch in end of March and every year we do four challenges so we pick certain tasks we've done over 15 so far over six years this is our sixth year we have some very exciting ones so one of the interesting ones that we're aiming to launch that relates to our conversation today a bit is end-to-end driving from just camera inputs. We are sharing some very interesting scenarios among the ones Waymo has seen, and we want to see how people can take some of these large models and get them to generalize on reasonably rare conditions.
1:06:14So that's exciting. We also have a simulated agents challenge. We're unique in running such a challenge. it's a very interesting problem where you're building agent models that populate the simulator and we have ways to validate whether they're good models and we have improved the metrics a bit this year we have one on traffic generation again something we discussed so let's say I give you the road graph of an intersection I want you to potentially populate it with realistic traffic that you know The relevant positions, velocities, behaviors are reasonable and not ad hoc, but we have ways to measure how realistic a traffic scenario is and we want to see what people can do.
1:07:07And there is interaction modeling as well. That's a challenge we're bringing back with a bit improved metrics from, I think, maybe four years ago is when we were on it, when we last talked. So we're bringing that one back. and here I am back on your podcast also four years later. That's an exciting one. Modeling how agents react vis-a-vis one another well is an interesting challenge and I think a lot has changed so I'm excited to see how the new crop of methods that people try with doing it. Well, we will include a link to those in the show notes as well as a link to the past conversation if anyone wants to go back to that one.
1:07:47but thanks for taking some time to chat with us about what you are up to there and how you're approaching this new world of foundation models. Very cool stuff. Thank you so much for having me on. It's always a pleasure. And yeah, well, I look forward to you trying us and telling me how it goes as well when you're in the city with a Waymo driver. For sure, for sure. Thanks so much, Drago. Thank you. Bye. Thank you.
From the publisher
Today, we're joined by Drago Anguelov, head of AI foundations at Waymo, for a deep dive into the role of foundation models in autonomous driving. Drago shares how Waymo is leveraging large-scale machine learning, including vision-language models and generative AI techniques to improve perception, planning, and simulation for its self-driving vehicles. The conversation explores the evolution of Waymo’s research stack, their custom “Waymo Foundation Model,” and how they’re incorporating multimodal sensor data like lidar, radar, and camera into advanced AI systems. Drago also discusses how Waymo ensures safety at scale with rigorous validation frameworks, predictive world models, and realistic simulation environments. Finally, we touch on the challenges of generalization across cities, freeway driving, end-to-end learning vs. modular architectures, and the future of AV testing through ML-powered simulation.
The complete show notes for this episode can be found at https://twimlai.com/go/725.




