In short
Episode Notes: Eye On A.I. – #182 Dragomir Anguelov: The Role of AI and Machine Learning in Waymo's Self-Driving Cars
Podcast Overview Podcast Title: Eye On A.I. Host: Craig S. Smith Episode: #182 Guest: Dragomir Anguelov, Vice President, Head of Research at Waymo Episode Release Date: (Date not provided in the transcript) Episode Length: (Duration not provided in the transcript)
The episode dives into the advancements in artificial intelligence and machine learning that are pivotal to the development of Waymo's self-driving cars. Craig Smith engages with Dragomir Anguelov to explore the technology behind autonomous vehicles, the challenges faced in their deployment, and the future of self-driving technology.
Key Points Discussed
Introduction to Waymo and Dragomir Anguelov
- Dragomir Anguelov discusses his role as the head of research at Waymo, emphasizing the company's goal to make autonomous driving safe and scalable.
- Waymo's initial focus has been on ride-hailing services, but the aim is to leverage their technology for various mobility applications.
Machine Learning Advances
- The conversation includes the evolution of machine learning technologies applied to perception, behavior prediction, and planning in self-driving systems.
- Anguelov highlights the foundational work in large neural networks for image understanding, which is critical in autonomous driving.
Technical Stack of Autonomous Driving
- Dragomir describes the technology stack of Waymo, which integrates multiple sensors (cameras, LIDAR, radar) for environment perception.
- The onboard system processes data from these sensors to create a comprehensive model of the vehicle's surroundings, enabling safe navigation and decision-making.
The Role of Data
- Waymo has developed the Waymo Open Dataset to provide researchers with access to high-quality autonomous vehicle data, encouraging advancements in AI research.
- Anguelov emphasizes the importance of having extensive datasets to train machine learning models effectively.
Challenges and Scaling
- The episode delves into the complexities of scaling autonomous driving technologies across different urban environments, such as San Francisco and Phoenix.
- Anguelov mentions how the understanding of rare event scenarios (corner cases) informs model training and safety evaluations.
Expansion and Future Outlook
- Dragomir shares insights on Waymo's expansion plans, including future deployments in Los Angeles and Austin.
- He expresses optimism about the timeline for widespread robo-taxi services, projecting significant advancements within the next decade.
Safety and Trust
- Safety is underscored as a paramount concern, with Anguelov discussing Waymo's rigorous safety methodologies and the importance of building public trust.
- He contrasts Waymo's approach with other self-driving companies, particularly in terms of safety protocols and system robustness.
Convergence of Technologies
- The discussion touches on the integration of generative AI, large language models, and multimodal approaches within the context of robotics and autonomous driving.
- Anguelov reflects on how the convergence of these technologies is accelerating advancements in the field.
Key Takeaways
- AI and Machine Learning: The backbone of Waymo's autonomous vehicles, with significant improvements ongoing in machine learning applications.
- Safety Protocols: A multi-faceted approach to safety is crucial for public trust and successful deployment.
- Data Importance: Large datasets are vital for training effective machine learning models and improving the system's overall robustness.
- Future Growth: Expansion of Waymo's services is expected to grow rapidly, with advancements in technology supporting this trajectory.
Conclusion The episode concludes with a call to action for listeners to stay informed about the advancements in AI and robotics as they have the potential to revolutionize transportation and various sectors of everyday life.
Additional Information
- Waymo Open Dataset Challenge: Anguelov discusses the upcoming 2024 Waymo Open Dataset Challenge, emphasizing its role in advancing AI research.
- Public Engagement: Waymo encourages the public to experience its services in operational cities, highlighting the tangible advancements in autonomous driving technology.
For more insights and to listen to the full episode, visit [Eye On AI](https://eyeonai.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00In Waymo's charter, we would talk that the Waymo driver, we eventually want to enable many different kinds of use cases and form factors. I think our first kind of flagship application is right-hailing. But we're building a general stack, right? I mean, we discussed that maybe we can transfer a lot of the models or types of models that we build for autonomous vehicles to other mobility robots. But even before that, right, we aspire to build a stack that learns from data and generalizes across these platforms and enables various, I mean, autonomous driving applications, preferably with as similar of a system as possible or maybe the same system.
0:43So it will take a few years to continue this trend until people really feel it all over the place. And to me, mostly, I don't count how many cities we will be in. I want to build the stack that makes it economically great and scalable to do a city N plus one. And so there's a ton of headroom still possible with the technologies that are being developed. And so I hope we can facilitate it. I believe it's not far out, right? Hi, my name is Craig Smith, and this is Eye on AI. In this episode, I talk to Dragomir Angolov, head of research at the autonomous vehicle company Waymo. Drago discusses how Waymo's research team has leveraged machine learning across perception, behavior prediction, and planning to enable 24-7 fully autonomous ride-hailing services in San Francisco and Phoenix, with LA coming soon.
1:40He talked about the Waymo Open Dataset, including the 2024 Waymo Open Dataset Challenge, which kicks off this month, and its impact on advancing AI research beyond autonomous vehicles. Finally, Drago gives his views on when we can all expect to be summoning robotaxis on our phones wherever we live. I hope you find the conversation as exciting as I did. At home, at work, we all know one person who's password challenged. Sticky note reminders, emailing passwords, reusing passwords, using the word password as their password. Because data breaches affect everyone, you need 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone.
2:36Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach costs millions of dollars. 1Password secures every sign-in to save you time and money. 1Password lets you switch between iPhone, Android, Mac, and PC with convenient features like Autofill for quick sign-ups. All you have to remember is the one strong account password that protects everything else. Your logins, your credit cards, secure notes, or the office Wi-Fi password. 1Password generates as many strong, unique passwords as you need and securely stores them in an encrypted vault that only you have access to.
3:22I use 1Password, and you should too. 1Password's award-winning password manager is trusted by millions of users and over 100 ,000 businesses from IBM to Slack. It beat out 40 other options to become Wirecutter's top pick for password managers. Plus, regular third-party audits and the industry's largest bug bounty keep 1Password at the forefront of security. Right now, my listeners get a two-week free trial at 1Password.com slash IonAI. That's IonAI all run together, E-Y-E-O-N-A-I. That's two weeks free at 1Password.com slash IonAI. 1Password.com slash IonAI I was a tech lead of 3D vision and pose estimation for Street View and got a lot of exposure to all kinds of platforms that the team was collecting data with, including boats, snowmobiles, backpacks, and of course cars.
4:31And afterwards, I got back a bit closer to my machine learning AI roots. I worked on a team and eventually ended up leading a team that developed, even in those early years, 2013 to 2015, large neural networks for understanding images. And so in those times, we had a very exciting set of developments and architectures that we invented. And we won the ImageNet challenges in 2014. That used to be a very popular competition on object understanding and detection and classification. So we won those two challenges in 2014 with the team that I led and some collaborators from Google Brain. And we also worked a lot on launching the backends that would annotate Google Photos.
5:26So Google Photos was first of its kind to offer deep semantic annotation of photos enabling search and all the photos applications that they would do where they suggest you various events and albums and so on and organize them. That was based on backends and models that we developed at the time and that we launched in the Google data centers. and around 2015 when I felt that deep learning has entered a level of maturity that is sufficient to potentially make autonomous driving a reality, I got back into the autonomous vehicle space, into robotics that I had been exposed to in my early lab days and so I joined initially ZOOX for two and a half years.
6:17It's an autonomous driving startup. I led the perception team there helping them build out their essentially 3D understanding and the machine learning around it. And then I came to Waymo, right? So that's my story, more or less. I would say I'm a machine learning person and my initial expertise had been understanding images and 3D data, geometry, and that's generalized to other applications since autonomous driving is a, I mean, it's a complete robotics stack, which includes all the key robotic systems, perception, planning, behavior prediction, simulation, mapping, localization, and so on. So by now I have this maybe more of a cross-functional exposure to all of this through the lens of machine learning.
7:16Your undergraduate was where and then at Stanford you were studying under Daphne Koller. And what was that work about? I actually did my undergrad at Stanford too, but afterwards I continued into a PhD that took six years. So I spent a total of 10 years at Stanford. It was hard to get me to leave for a while. I worked closely with Sebastian Throne because I was Daphne's first PhD student that worked on computer vision perception topics and we wanted access to robots and data from robots and at the time sebastian wasn't even at stanford but daphne somehow made a connection with him and i ended up visiting his lab in carnegie mellon and collecting data with robots and shorting one robot I'm a software person.
8:08Fortunately, there's things I learned the hard way there. But Sebastian was always a commenter of sorts, even though my primary advisor was Daphne. I think I always worked with Sebastian as well, and we have joint works, and he was a key influence, too. Yeah. And so Waymo, when you got involved with the company, where was it in its development? And where are you right now? I mean, I know you have robotaxis operating in San Francisco and Phoenix, and I hear you're going to have them operating in LA. Yeah, we have San Francisco and Phoenix today. You can experience autonomous driving. It exists at scale in these cities.
8:58We give tens of thousands of rides to users in both cities every week. And Waymo overall has given over a million of paid rides, driving over 10 million miles autonomously as well. So that is the two main markets that we have. The Phoenix area that we cover today is around five times the size of San Francisco. And next, we're looking to expand and start actually paid rides in the coming weeks in Los Angeles. We got a permit a couple weeks ago to expand to Los Angeles in charge. And so we will start doing this. And we also started doing driving driverless rides for employees in Austin, Texas. So we're going to be expanding in this fourth market that we have announced as well.
9:54My time at Waymo has been a wonderful time, seeing a transformation of the stack and getting to maturity that I think has been great to experience and working with a talented team. I think when I joined Waymo in 2018 in the summer, Waymo had hit some great milestones. One of them was in 2015, it gave the first full autonomous ride in Austin, actually. The city we're getting back to, the person called Steve Mann, a friend of the company as a blind person, the first ride involved him. and so we had done that and we were working towards our first commercial deployment and so we're evolving the stack to make sure that it will be a safe and scalable deployment and this eventually happened in 2020 when we launched our Waymo 1 service in Phoenix East Valley around Chandler and so that's been the journey so evolved to a first full autonomous service in Phoenix East Valley, then get to San Francisco and then significantly upgrade the stack to be able to handle the complex, dense urban environment of San Francisco.
11:14And now we continue to all these more cities, expanding in Phoenix. We will be working also on tackling more highways, which is important in large areas like Phoenix. Our area, if we don't go on highways, takes almost an hour to cross on surface streets today so uh so yeah that's i think maybe uh the trajectory of the last five and a half years it's been tremendous also in in terms of development of machine learning and machine learning technology so in 2015 when i joined 2018 excuse me when they joined waymo there was already a lot of impressive machine learning work um especially there was a lot of really interesting, unique work on 3D perception using LiDAR.
12:02And this was in collaboration with Alex Krizevsky and Anilie Angelo and other people from Google Brain, Alex Krizevsky being of the 2012 AlexNet fame. So he took a lot of interest in autonomous driving as an application, and he worked a lot with Waymo and developed some really interesting work, including inspiration to a first paper Waymo ever did. Even though we published it in early 2019, it's called ChauffeurNet. It's actually a trained model to predict and drive by imitating others. That was really groundbreaking, unique work at the time. And then that was still relatively early in its development and maturity, but it was very exciting as a concept.
12:52So we published that and we took all these technologies and we worked a ton on expanding them and transforming Waymo stack. I think at this point, Waymo is an AI first company. So we strive to maximize the impact of machine learning and its power across most of our applications. Yeah. I haven't spoken to anybody about autonomous driving for a while. I did an episode with Alex Kendall at Wave AI, but we were really talking about world models. In the tech stack at Waymoor, from what I've read, you have detailed custom maps for navigation. You have this suite of sensors, and then you have neural nets that are kind of wired together to process that data and and guide the car.
13:52How much of that, how much hard-coded stuff is in that stack, and how much is pure machine learning? So big parts of the whole stack of machine learning, and this spans all the core systems of Waymo. This spans things like perception system, and it also covers major models applying machine learning to planning, behavior prediction of others, and simulation, and also lots of machine learning for employed also in evaluation, extracting signal from, for example, simulation or other tests we perform on the car. So at this point, I would say all major parts of Waymo have powerful, large machine learning models underpinning them.
14:42Are all those models on the car, or what's on the car, what's in the cloud? So we have some flexibility. I think especially the whole evaluation system simulator that is in the cloud, right? Because we develop a variant of the stack and then we need to test it. And we test it in the data centers to a large extent, even though we also test it with human safety operators and we won't have specialized areas where we can stage certain rare situations. but testing in the cloud and simulation is key. I think generally we avoid requiring a human to drive our cars at any time. It's possible to ask human occasionally a specific question in some very, very rare situations when preferably the vehicle, or usually I think almost always, the vehicle is stopped at that time.
15:40can be, but that is also extremely rare. And we are trying to minimize this. So the vast majority of the system that drive the car, they're on the car, on the local compute, and they operate fully, I mean, autonomously as advertised. So how many models or neural networks are involved in the driving system? So there are several, right? I think the history of our space is one of consolidating models, right? And traditionally, many players in the space, and also when machine learning was in its relatively early stages, and there was limited compute and data, it has been beneficial to, or at least at the time to partition the tasks and have maybe one neural net to detect bounding boxes and another to segment, say, the scene into semantic segmentations of various kinds and another maybe.
16:40So traditionally that was the history. And now what's changed is we understand the power of... The models have become a lot more powerful. Their scaling properties have become much improved, especially with the advent of transformers. We understand a lot better how to make models do well on many, many tasks at the same time. That's actually been area of research by us and the whole community. And some of these AI trends can be leveraged, train very large models, so transfer knowledge from the internet and fine-tune on your tasks. So we have ways to make models do well on a lot more tasks today.
17:25And there is benefit then in efficiency and simplicity and quality usually to consolidate. So there has been a push to consolidate. That said, it's an ongoing process. There's some very interesting question whether you should have one model or several. There is challenges in having just one model in our case, especially if you need to iterate on it and ensure it satisfies, I mean, hundreds to thousands of requirements, right? You need some way to have, to be able to do work on parallel problems in this domain to ensure safety. And so there is this constraint. And there are some other constraints that make it difficult for our case.
18:11For example, compared to, say, traditional robotics. Maybe you talk to Peter Abiel or Vincent Van Oog. They have a camera on a robot, and then they have a world model, and they can train it end to end. For us, it's more challenging infrastructure-wise because we have, say, over a dozen cameras to two dozens. We have a few LIDAR, we have a few RIDAR, fitting all this data for all-in-one on the compute and scaling it is quite daunting. And that even motivates potentially training things in stages and so on. So I would say that the answer is we're evolving and we're evolving, satisfying both in the learnings in the industry and the requirements in production.
19:01and generally the trend is to consolidate. Yeah, and I know it's a very complicated system, but can you sort of walk us through how the car drives? I mean, they're the sensors and they're feeding the data into a perception model and then the perception, I mean, just kind of walk me through the pathway. So the simplest way to think is of the onboard system, right? And there's a lot of very interesting problems and questions when you build the simulator as well. But let's talk about, which is also a machine learning task to me, by the way. I'm a machine learning person. And ultimately, we want to solve problem with the data we have as much as possible.
19:46But the traditional autonomous vehicle stack takes data from sensors, can be camera, lidar, radar. And then there is a perception stack that processes this data and creates a representation of the world. In this representation of the world, one of the key aspects of it is it wants to be expressive, but also concise, right? And in that representation, then you train your behavior and potentially simulation models. And because that's a representation of the world, and then you need to, in that representation of the world, you need to plan what you do, and you need to also potentially imagine what the world is going to do in response to your plans to validate that your plans can be saved.
20:32So essentially, there's these main two parts to the stack. Perception produces a model of the world. And then in that model of the world, in plan, you validate your plans. That's the behavior system. And the model of the world is very interesting. So there is a lot of interesting question of what it can or should be. Traditionally, we have opted for at least partially interpretable intermediate representation. So one of the things that has been very beneficial in our space, and I think it's generally superior to treating all the sensors as an ordered collection of sensors, is to take their data, fuse their representations and the knowledge provided to them, and create what we call a bird's eye view model of the world.
21:17So essentially you center a grid or implicit grid on the space around a vehicle, and you essentially populate this grid with everything, all the information that you acquire from these sensors. So that's been very popular. And there is certain benefits to having interpretable representations. People understand them. They know what the problems are. They can annotate or correct them. And so the drawback of those is you need to label data to create them, typically. They're representations that we have chosen because they satisfy requirements, interpretability, testability, right? Potentially people can look at the outputs of a system and impose further constraints or corrections.
22:06That's why you asked me, is all your system just a neural net? Our system is hybrid, even though big parts of it are neural net, and we strive to solve the majority of the problem with machine learning. We still build such representation and design where you can control and steer the system by experts that can introspect the results and introduce additional constraints, requirements on top of what the machine learning models do. And that has been very beneficial for robustness and safety. I guess that's what Wave is, where they're different. They're creating a world model, and then everything happens.
22:45All the planning happens in that. The simulation and planning happens within the world model. I actually don't think they're that different, honestly. I think maybe it was oversold. You have a perception system. It produces a compressed version of the world. In our case, some of it is interpretable, but you can also have various embeddings or quantized embeddings of tokens trained end-to-end that can also be produced. It's a question of choices in that representation. And then given that representation, you have prediction models. And these are, you can call them their world models. And there is a question of which exactly world models do you build.
23:27Do you build a model that just dreams all the video of all your cameras? Well, you can, but on board, maybe that is overkill. So then which things do you choose to dream and auto-regress or predict concisely so that you still satisfy your requirements? That's where you have leeway. But we still have major machine learning models that are word models. Just the design of them may be different. And we also can train end-to-end. However, we have also put historically a lot of focus on having intermediate representations because we have found they're very important so far in actually building a fully working system.
24:14Machine learning is great, but it has limitations. There's always cases where eventually you want to impose guardrails or constraints, and intermediate representations are very powerful for doing it. Now, lately, there's been yet one more trend with large language models and the like. Now, there is a lot, I mean, the kind of language-based representations or language-compatible embeddings are very, very popular for robotics because they bring a ton of knowledge from the internet. and that said, the model, the language models were pre-trained on, which makes it, I mean, the system understand a lot more common sense and be able to interact with humans.
24:56So now that is also very interesting to us, right? So language or language compatible representations are great. That said, there is also another learning in our domain. You really need these representations to be spatially correct. So our models, the perception system needs to be very powerful at spatial reasoning, because spatial reasoning, the ability to represent the world and ultimately evaluate your plans in the world is key to ensuring safety. And so these concepts are now overlapping. And I think we have a lot of exciting, both experience, expertise, and of course, we're also experimenting with the best mix of these concepts.
25:39Yeah. Yeah. I mean, I just had Vincent van Hoek on the podcast talking about RT2 and that system. We are very much inspired by their work and follow it. And I think there's a lot of parallels to our domain. There is certain differences where we are a lot more concerned about correctly predicting and modeling the behavior of others just because in our domain, it's so important to interact correctly with all the traffic participants. You know, a lot of people don't think of cars as robots. They're absolutely robots with well-established hardware stack, yes. Is your system, could it be adopted or adapted for other kinds of robots?
26:27I think definitely it can be adapted. I think a lot of the concepts across the various robotic systems over time, I believe, are converging. There is still certain peculiarities in our domain. You can think of our robot maybe as the first manifestation, maybe the most economically, hopefully impactful and application impactful of most robotic applications. At least I believe it. And that's why I'm in this domain. But a lot of these concepts of how to build your stack, how to test your stack, how to ensure robustness, what kind of machine learning models you do. I think we're becoming more and more similar across robotic.
27:08But yet, in our case, safety, especially at high speed, we have a lot more requirements of those kinds. Like you need to have very fast reactions. You need to have very high bar to safety, right? You're interacting sometimes in parts of second with other agents who could behave also irrationally or adversarially. we need to handle all of this. And the other part is we have very many sensors on the vehicle, right? So to see in every possible direction, and in our case, we've opted for a lot of additional safety by adding LiDAR radar, which are active sensors. They ensure that even if machine learning in cameras misses something, we have other ways to capture it.
27:49So when we put all this together, there are certain differences, but if you build the stack for autonomous vehicles, we can build on top of the strengths of this stack. We can adapt it to make. At home, at work, we all know one person who's password challenged. Sticky note reminders, emailing passwords, reusing passwords, using the word password as their password. Because data breaches affect everyone, you need one password. One password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. Companies lose hours every day just from employees forgetting and resetting passwords.
28:33A single data breach costs millions of dollars. 1Password secures every sign-in to save you time and money. 1Password lets you switch between iPhone, Android, Mac, and PC with convenient features like Autofill for quick sign-ups. All you have to remember is the one strong account password that protects everything else. Your logins, your credit cards, secure notes, or the office Wi-Fi password. 1Password generates as many strong unique passwords as you need and securely stores them in an encrypted vault that only you have access to. I use 1Password and you should too. 1Password's award-winning password manager is trusted by millions of users and over 100 ,000 businesses from IBM to Slack.
29:26It beat out 40 other options to become Wirecutter's top pick for password managers. Plus, regular third-party audits and the industry's largest bug bounty keep 1Password at the forefront of security. Right now, my listeners get a two-week free trial at onepassword.com slash ionai. That's ionai all run together, E-Y-E-O-N-A-I. That's two weeks free at onepassword.com slash ionai. Onepassword.com slash ionai. Yeah. And as the system evolves, when you were saying that it's evolved a lot and then there's all this new research coming out with generative AI and large language models, multimodal models.
30:17It's a great time in our space. I think robotics is, I think progress in robotics is accelerating even though, you know, we're an example of already a successful robotics product out there in the real world. But your model, how do you, how do you integrate these new technology? But when do you decide to bring that into the car? The great thing at Waymo is we have the data and the frameworks to test the performance of our stack. And so the history at Waymo always is, you know, we've been around 15 years. And the technology in these 15 years has been dramatically evolving. And every two or three years, there's major advances in our space, right?
31:01And so we've had to continually rethink and redesign our stack to take advantage of the latest technologies. And so the core is, as in most machine learning, is we have the data and the evaluation. We have collected, as you know, we have many hundreds of vehicles out there in the world with very rich sensors. We have a lot of systems that can process this data, collect the interesting cases, potentially automatically annotate things we want. If we, for super rare things we can get, we have systems to annotate human guidance and so on, right? And so we have all this data. Now it's a question of, it's not that hard to try new paradigm or models and compare to the performance of your previous models.
31:46And anytime we can show a significant improvement, that's then we evolve the stack, right? And this has been done many times already. So, you know, we continue doing this. But the data is the key, right? So because you have all this data and you have all these testing frameworks and safety frameworks and very rich evaluations, the key is to be able to tell when you're better, right? And we worked hard to be able to do that. Yeah. There's so many questions I want to ask. One is on corner cases that arise in this long tail. I read something, this is four or five years ago, I think, of a case where an autonomous vehicle, it may have been Waymo, drove up behind one of these big tanker trucks with a highly polished back that was like a mirror, and that confused the system.
32:41How do you deal with that? Do you add that to your training data or do you tweak the algorithm? I mean, how do you deal with those corner cases? So you can annotate your data, right? And you can use it either for training or for testing. And clearly, the very first thing you want to ensure when you see very rare cases is you want a test set to ensure that you got it right. And then you look for more cases like this and you add those to training set to help you get it right. But I think in our case, because we have also radar and have LIDAR and we have camera, we will see the truck. We won't see just a car reflected in the truck.
Read the full transcript
33:23So we by design have mitigated this to a large extent compared to a camera-only system. For example, some people had an idea, oh, let's put a TV in the back of my truck and I'll show you the world in front of the truck in the TV. Well, imagine the autonomous driving car with a camera looking at that TV thinking, what's in front of me? I don't know. Is there a car there? Maybe it's not so relevant because it may look even far away in some cases, but there can be a lot of confusing things. And that's why having more sensors of different modalities is quite conducive to safety. That's not the only one, but we collect these rare examples and we mine for them.
34:02We add them to the training and testing data and we validate as that can handle them. I think it's a standard engineering machine learning approach. Can you talk about the open data set, the Waymo open data set? And as part of that, Vincent was talking about the RTX project where they collected data from all these various labs where the data has been siloed for years and then combined it and then sent it back to the labs to train on the bigger data set. and there were improvements. Were you part of RTX? And how does the Waymo Open data set fit in your stock and in your training? So let me explain.
34:47I mean, generally, we're separate from RTX. I think that data set is primarily focused at object or robots that manipulate on a table, right? And ours is focused on understanding the outdoors environment and traffic scenes. So it has somewhat different parameters. The reason why we made the Waymo open data set is based on the realizations we had in the early years, especially around 2018, when we really started the project, but even before we did research in the space. And the data sets that were available, even though seminal and great, there is a data set called Kitty at the time that was made in 2011-12 by Raquel Urtasun and I think European collaborators.
35:28So we started seeing that that data set does not satisfy our needs when we build the actual stack. And the problem there was it was typically very, very small. And when the data set is small for a machine learning person, it distorts the kind of models that win. So then that data set to do well, it encourages you to build models that either have a lot of built-in bias, or you built in a lot of overfitting techniques, which is not at all what you're going to do when you build the the real staff and so we understood that there was the set of problems they covered were a bit limited but what was even more so is the data and so the academic community is such talented great community and it's not easy to get your hands on large amounts of reasonably clean autonomous vehicle data i mean you need to have a vehicle you need to place all the sensors you need to instrument, I mean, you need to align the times and calibrations and the poses.
36:29It's just very prohibitive, right? And so we wanted to enable the community. And the best way to do this and to unlock the talent is to give them the right kind of data. And through this WAM Open Data Set, this was our attempt to push what is possible in the standards, the amount of data and the quality of data that this community would get. And I would like to think that we did a good part into popularizing and enabling a lot of very strong AI work. And in addition to releasing data and enriching it with more and more labels over five years, so we constantly improve the data set based on feedback and on our understanding about the space.
37:08And we expand the things you can do with it, expand the set of tasks you could potentially try to show state of the art in. And we also, as we introduce new tasks and labels, we create challenges. And the challenge allows different researchers across the world to compare their solutions on criteria that are informed by our experience in the space. For example, what's good 3D semantic segmentation? What is a good metric for simulated agents, right? We take the learnings we have from our space and help them essentially create leaderboards where people can compare their contributions on definitions of problem informed by us.
37:51So this has been an ongoing five-year effort, and we're launching actually next week. Today is March 12th when we're taping this. We expect next week to launch, or the week after, our 2024 versions of challenges. There are four challenges on exciting problems, 3D semantic segmentation, 3D occupancy prediction, and 3D flow prediction challenge. Then we have motion prediction in terms of trajectory challenge and embodying simulated agents. So these are the challenges this year. And the winners will get prizes and the best work will get invitations to present at the CVPR, Premier Computer Vision Conference in June at our workshop that we're organizing with other academic collaborators.
38:45So I encourage people to check it out, especially if you're interested in our domain or doing research in our domain i think there's some very exciting challenges this year and i believe this is our fourth or fifth year of doing challenges so we do our best to sustain improvements to this to the data set and uh the amount of tasks we release over time so yeah um on on the data set How is it collected? It was collected by Waymo Vehicles running the cities of San Francisco and Phoenix. The majority of the data was collected around 2018 or 2017, which is when we chose a lot of the run segments or the snippets of real data that we decided to release.
39:35And we released 2 ,000 or so run segments. They're like sequences of LIDAR and camera data for 20 seconds. And we've released since then 100 ,000 run segments, both with intermediate representation in terms of road graph and bounding boxes moving around for the various objects, but we also released LIDAR for those segments. And this year, even image embeddings. So 100 ,000 sequences is very low, very large number. It's quite prohibitive to release all the sensor data, camera data for five plus cameras on this, but we at least released compressed versions of the camera representations to enable yet more research into combining camera, lighter and intermediate representations, see how good prediction can be made.
40:25And the key for us is that this imitative prediction of agents is one of the key crutches in our domain for machine learning, because we observe how everyone drives. We observe how everyone walks. But in general, you can learn a lot about driving by watching drivers. And that's a key. That's one of those few domains where, you know, maybe for other robots, you never see how a certain dog robot should walk much in the real world. But for driving, you do. And so it's really, really powerful data, and we have been pushing a lot of research, as you can see some of our challenges this year on predicting various things, to take advantage of this key part of the data.
41:06Are there other contributors to the data set, other self-driving car companies, or other labs that have data that's relevant? I think people contribute other labels to our data set and do potentially their own leaderboards or competitions on top of that. There are other companies have released data sets as well. We're not the only ones since we started, right? So this also shows that the companies themselves understand that this is an important thing to do, right? And as the field evolves, like your understanding of the problems we want to make as challenges and the data that needed to support them evolves.
41:49And so, you know, I think we will be, as a community, and I hope at Waymo we can sustain what we're doing, but as a community, I would expect that there'll be a lot more interesting releases and challenges possible. Yeah. The data, as it's coming in from your cars, are the cars sharing data with each other in real time, or is the data going back to a central repository? Well, it's collected in the central repository. When our cars drive, they typically don't send each other's data to each other. That presumes that you will have very high bandwidth connections. It's actually a lot of data. It's hundreds of millions of pixels and lighter scans a second.
42:31So that's why the storage even for the things we're releasing is so big. It's almost a terabyte. And so even when we compress it, right? So these are the challenges. But when they're in the data center, sometimes we see the scene from several vehicles at the same time, right? Which can, of course, help us at that time reconstruct much higher fidelity models of, well, what is seen by now a collection of Waymo vehicles, not a single vehicle. This is possible, but it's not common practice in the field. Yeah. The models, part of the trajectory is to make the models generalize more and more. That's one of the main tasks of the machine learning problem.
43:14You make the most of the data you have, both to build the agent and to build this environment to test the agent. These are the two main problems as I see them. Right now, of course, you can break them down into sub-problems, but that is at the core of it. For the data you have that you're collecting and maybe keeping, how much of this can I learn with my models? The more I can learn, right? the more we tilt the system towards machine learning solutions because those just scale. Then I go to any new environment, I collect more data, it will just learn from it, right? In different parts of the world, I remember in the early days, there was a lot of talk about you train a system in Scandinavia and then try and operate it in the Philippines or someplace.
44:01Yeah, or you go to Singapore and everyone drives on the left side, right, for example. Do you need different data sets to train models for different environments, different regions, or do the models generalize enough that they can quickly adapt to a new region? So we've seen that with large models, right, you actually benefit combining the data. And this parallels the development in generative AI field, right? Like they try to train on, well, some large internet scale data sets. Or when you train translation model, often you train it on multiple languages at the same time. So when size is not a constraint, there is positive feedback across different tasks in different regions.
44:44And we've seen some of this even in the case of, so say car data helps tracks. So you need now a lot less track data to do well because you can just mix in large proportions of the car data that you have. And lo and behold, the results improve, right? And the same is true across cities and environments and conditions. For larger models, we essentially would look to pull all the data and train a much larger model and then essentially potentially distill this model or adjust it, fine-tune it to the specific cases as needed. But that seems to be the recipe today. So we aim to have one Waymo driver, and as few of them as possible.
45:25Traditionally, so far, all the cities we drive is a single set of models. It's not a fragmented set. That's really hard to maintain, and it's not scalable. That's not what you want. I don't know if you can talk about the cruise debacle. We don't have to talk about the company. But what went wrong there, and how is Waymo avoiding those problems? I mean, I would say that to me, the lesson is safety comes first. It's paramount. And I think trust is key. And so it's very easy to lose people's trust. Ultimately, I would like to think that so far Waymo has a track record of expanding thoughtfully and successfully in the domains where we have operated.
46:16And so we want to expand relative to what we believe is beneficial and safe. And I think we have a lot of data to show that our deployments, I think, compared especially to humans in various studies, is shown that we have a lot less crash rate is a lot less and our property damage claims. there is actually a collaboration with our Swiss reinsurance company showing that say injury claims went to zero with our vehicles and property that's based on 3.8 miles that they analyzed right and property damage claims went down 76 percent right and this is this is track record that we want to have and continue expanding on so we look to expand thoughtfully and I think the way to do thoughtfully is to have extremely robust safety methodologies and put what analyze what you put out there with incredible rigor and honestly uh our safety methodology is very multifaceted so there is no one golden hammer to ensure the safety of the system and there is no one golden machine learning model simulator to to prove that your system is safe our learning is you need to approach what we have with a variety of methodologies and make sure that we looked at it from many angles.
47:45So this is actually one of the most main recipes that make Waymo Waymo is that we have been working on this for 15 years. And there's a level of maturity that honestly, when I joined Waymo at the time, I was one of the things that most impressed me is that type of. Yeah. So it wasn't a matter of that there was that the data set was inferior or the the models were inferior or the there weren't enough guardrails or or hard-coded rules um it's it's a it's sort of a combination i mean you want to build a system that you convince yourself is safe to put out there right and then that's a high bar and that's not a simple question to solve right and i think at way more we tested extremely extensively in a variety of different ways and that determines how quickly we expand as our success in our uh tests right as opposed so and we we do this also thoughtfully you need to engage and work with the community right so at phoenix we have been delivering rides close to six years now right and over time now we're at a scale which is quite noticeable uh right and thousands of people or maybe tens of thousands take us every week uh but you know it's not like we we were not reckless right i don't know i it's hard for me to talk about cruise per se this is more the lessons for us i think cruises data and i think the ultimately the details of the incident that had them grounded i think a lot of the details are out there and people can judge for themselves.
49:31But you need to ultimately build trust and a lot of this trust is based on the safety record you have. Yeah. From what I understand, China has more robo-taxis deployed than the US. Do you think that they're, presumably you pay a lot of attention to what they've got on the road. Do you think that they're ahead of you guys or is everybody kind of on the leading self-driving car systems on par? And then the other question I want to ask is I see we're coming. We're going to run out of time. There's a lot of confusion about Tesla's system. How is it? Does it approach what Waymo is doing? or is it something different altogether?
50:25So maybe let me first address the Chinese company's question. I mean, clearly a few companies there have started fully autonomous deployments. Unfortunately, I don't have full visibility into their system. I have not had the chance to write in them or study them. I mean, a lot of that design is proprietary. I think ultimately one needs to experience the service and the people who are best suited to opine how we all compare technology-wise is people who experience all of these different services, right? So that said, clearly in China, it's an area of investment in maybe even government-assisted domain, right?
51:14That said, they're not here currently in meaningful deployments, so we are not head to head with them. And I don't follow them that closely, even though I understand that, you know, that is an area of rapid development. In terms of company like Tesla that deployed more driver assist type technologies. I think you know I appreciate the machine learning work and innovations they do I think to me in their nature the system is not full self driving right so they do good machine learning but as with most machine learning I'm a machine learning person as much as I appreciate it right there comes a time when you get to situations that are extremely rare and or machine learning cannot handle.
52:08And to have a real deployment out there, you need to carefully think through your stack to make sure that even if machine learning does not solve everything 100%, your full software product does. And that's a very big gap. I think that's not easy to do at all. And it leads potentially to rethinking core parts of your design if you are to go there. And so, I mean, all of our machine learning models improving yet to release a driverless service, you need to also answer the question of how to make sure the whole thing is fully robust. And that adds a level of design and complexity that a lot of those companies have not tackled yet.
52:50Yeah. Would your system ever be available for consumers to buy? I mean, in the beginning, right, in Waymo's charter, we would talk that the Waymo driver, we eventually want to enable many different kind of use cases and form factors. I think our first kind of flagship application is right hailing. But we're building a general stack, right? I mean, we discussed that maybe we can transfer a lot of the models or types of models that we build for autonomous vehicles to other mobility robots. But even before that, right, we aspire to build a stack that learns from data and generalizes across these platforms and enables various, I mean, autonomous driving applications, preferably with as similar of a system as possible or maybe the same system.
53:48That is possible, right? And so I think down the line, we have the opportunity to build solutions for personal vehicles, for trucks, potentially other form factors. But I think ultimately for a company to succeed, you need to nail your first product. And I think where we are is, I mean, I personally really enjoy riding in our vehicles. I think it's a delightful experience. Of course, I'm biased. But they're available. Everyone can try them. You can come to San Francisco and Phoenix and soon LA and Austin and just ride and judge for yourself. But it exists. It's real. It's out there. And we've proven that the product exists.
54:35Now it's on us to make sure we can have a great business doing it and expand from there. Yeah. And how long do you think before every city has a robotaxi service? and then how long beyond that before they're autonomous vehicles on the highways and, you know, all over the United States. I mean, I understand it's largely a regulatory issue, but in terms of the reliability and the safety. So I think there's a question. We are expanding every year rapidly, especially maybe you can think of it for now. We're in a multiplicative phase. so every year we go many times the driverless deployments we've had before right and even though and that's thoughtful expansion I believe so it will take a few years to continue this trend until people really feel it all over the place and to me mostly I don't count how many cities we will be in I want to build the stack that makes it economically great and scalable to do a city N plus one And so there's a ton of headroom still possible with the technologies that are being developed.
55:50And so I hope we can facilitate it. I believe it's not far out, right? So I think my personally great hope is that this multiplicative expansion of operational design domain in its scope and volume, if we sustain it for a few years, people will see these cars in many more places. Do you think in 10 years, when I'm still got my marbles, I hope? Oh, for sure. I mean, I'm an optimist generally, right? And as an optimist, I would say it should be less than 10 years and probably by some mark. Yeah, that's exciting. And where in the development of either the data or the tech stack, what are you most excited about?
56:36I'm excited about, I mean, all this great technology we have that allows us to generalize a lot better on the data we have. And I think we understand better and better how to land these kind of scaling multipliers. So the more I can learn from the data that I have, right? I'm excited about keeping up with improving our scaling multipliers on what we can learn from the data. The more we do this, the faster we will expand, right? Because ultimately machine learning, a lot of the other things, so when you get to the hybrid design of the system, then when you need humans to essentially add expert expertise or guardrails, that doesn't scale very well.
57:26But machine learning is our great scaling tool. And the more we can use it, the faster we will grow across the areas that we serve. The other part that really encourages me, actually, is maybe the thing that makes it easier to do this, which is, so what we've discovered is, after you cover some of these areas, for example, Phoenix is more, parts of it are 45 miles an hour road with certain challenges. San Francisco is dense urban with certain challenges. Maybe we cover highway. Now a lot of work on harnessing highway and starting to enable that for customers. I think what we saw is that after you have, say, Phoenix and San Francisco and you go to LA or Austin, it takes a lot less work and effort.
58:17And a lot of it is actually invalidating our system that it does well as opposed to actually making it, right? And so the learnings generalize and that's a great kind of wind at our backs. And so as long as we design the system well and we don't, I mean, again, partition every city for itself, no, we actually benefit from San Francisco to do LA and we benefit from Phoenix to do LA. And when they have more cities, the city plus one will benefit from that. And so I see it also as I think that's a great property, right? Scale begets scale in our sense. And then, you know, having these deployments, having all these experiences and the data from them will help us with that.
59:01So the ambition is that there will be fleets of robo-taxis in every city and town. And will it be, from what you can see, I know that you're not on the business end of this, but will it be inexpensive enough that people will not need to own a car? You just, on your phone, you summon a robo-taxi. It arrives in five minutes. You go where you're going. I mean, in much the way that Uber operates? So I will tell you now, right, like we're roughly the same price, but we are much better experienced personally. Again, I'm a Waymo representative, so that's my personal bias. But you're in the vehicle by yourself.
59:45You can play the music you like. There is a certain level of, I mean, it's a high-end vehicle. It's a premium experience today for the price of normal experience. I think over time, there is tremendous opportunity to optimize both our models and vehicle cost and operations to make it yet more affordable than it is. Now, there is a lot more work that needs to be done. And I think it will shift from, for now, we don't even have as many vehicles as is the demand. And over time, hopefully, you know, as we make it more affordable and yet better experience and potentially that then it's a beneficial look.
1:00:28Then we'll get even if we can succeed with this, we'll get even we'll expand the market. At home, at work, we all know one person whose password challenge. Sticky note reminders, emailing passwords, reusing passwords, using the word password as their password. Because data breaches affect everyone, you need 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach costs millions of dollars.
1:01:121Password secures every sign-in to save you time and money. 1Password lets you switch between iPhone, Android, Mac, and PC with convenient features like Autofill for quick sign-ups. All you have to remember is the one strong account password that protects everything else. Your logins, your credit cards, secure notes, or the office Wi-Fi password. 1Password generates as many strong, unique passwords as you need and securely stores them in an encrypted vault that only you have access to. I use 1Password and you should too. 1Password's award-winning password manager is trusted by millions of users and over 100 ,000 businesses from IBM to Slack.
1:02:01It beat out 40 other options to become Wirecutter's top pick for password managers. Plus, regular third-party audits and the industry's largest bug bounty keep 1Password at the forefront of security. Right now, my listeners get a two-week free trial at 1Password.com slash IonAI. That's IonAI all run together, E-Y-E-O-N-A-I. That's two weeks free at 1Password.com slash IonAI. 1Password.com slash IonAI That's it for this episode. I want to thank Drago for his time. If you want to read a transcript of today's conversation, you can find one, as always, on our website, IonAI. That's E-Y-E hyphen O-N dot A-I.
1:02:57In the meantime, remember, the singularity may not be near, but A-I is changing your world. So pay attention.
From the publisher
This episode is sponsored by 1Password. 1Password combines industry-leading security with award-winning design to bring private, secure, and user-friendly password management to everyone. Companies lose hours every day just from employees forgetting and resetting passwords. A single data breach costs millions of dollars. 1Password secures every sign-in to save you time and money.
Right now, my listeners get a free 2-week trial at: www.1password.com/eyeonai
Dive into the world of autonomous driving with Dragomir Anguelov, Vice President, Head of Research at Waymo, in this captivating episode of Eye on AI.
Drago offers a profound exploration into the technical and operational advancements driving the future of autonomous vehicles. With his rich background at Google and now at Waymo, Drago provides an insider's view on the evolution of machine learning technologies and their pivotal role in developing self-driving cars that are not only intelligent but also safe and reliable.
This conversation spans the inception of large neural networks for image understanding to the intricate deployment of robotaxis in urban environments. Drago details the challenges of scaling these complex systems across different cities and the breakthroughs that have made autonomous driving a tangible reality in places like Phoenix and San Francisco.
Drago's insights are invaluable for anyone interested in the convergence of AI, robotics, and transportation.
If you are fascinated by the advancements being made in the autonomous driving space and how they could transform our daily lives, make sure to hit the like button and subscribe for more deep dives into the technologies shaping our future.
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI




