In short
Podcast Summary: How End-to-End Learning Created Autonomous Driving 2.0 with Alex Kendall
Podcast Details
- Title: Training Data
- Hosts: Sonya Huang, Pat Grady
- Guest: Alex Kendall, CEO of Wayve
- Episode Description: Discussion on the evolution of autonomous driving technology, focusing on Wayve’s end-to-end deep learning approach in contrast to traditional methods.
Key Concepts
Evolution of Autonomous Driving
- AV 1.0 vs. AV 2.0:
- AV 1.0: Relied on hand-coded software stacks, high-definition maps, and specific hardware (like LiDAR).
- AV 2.0: Focuses on end-to-end deep learning, allowing for a more generalized and adaptable approach to autonomous driving.
Wayve's Approach
- End-to-End Learning:
- Wayve's strategy eliminates the need for separate neural networks for each application, aiming to create a generalizable AI that can adapt quickly to different vehicles and environments.
- World Models:
- These models enable reasoning in complex scenarios, enhancing the AI's ability to generalize from diverse datasets.
Market Positioning
- Partnerships with OEMs:
- Wayve collaborates with major automotive manufacturers (e.g., Nissan) to provide AI solutions directly integrated into vehicles, avoiding the need for retrofitting.
Key Discussions
Safety and Interpretability
- Safety by Design:
- The necessity for AI systems to be safe and interpretable, especially in autonomous driving, requires robust architectural designs to avoid potential hallucinations.
- Interpretation Challenges:
- While early deep learning models lacked interpretability, advancements have led to better understanding and insights into AI reasoning processes.
Generalization in AI
- Importance of Generalization:
- The ability to reason about new scenarios not seen during training is crucial for autonomous driving. Wayve's AI needs to adapt to any driving conditions it encounters.
Data Diversity and Simulation
- Data Sources:
- Wayve aggregates data from various sources (dash cams, fleets, manufacturers) to improve the diversity and quality of training datasets.
- World Models for Learning Efficiency:
- Utilizing synthetic data generated by world models to enhance learning efficiency and reduce reliance on extensive real-world mileage.
Key Takeaways
- Shift in Industry Consensus:
- The perception of end-to-end learning has shifted from a contrarian viewpoint to a widely accepted approach in the industry.
- AI's Role in Physical Economy:
- The same deep learning advancements that propelled language models are now revolutionizing physical AI applications like autonomous driving.
- Future of Robotics and AI Applications:
- Autonomous driving serves as a foundation for the development of more generalized applications in robotics, paving the way for innovations across various sectors.
Closing Remarks
- Mission-Driven Culture:
- Wayve's commitment to safety, efficiency, and reliability is reflected in its organizational culture and its approach to product development.
- Future Innovations:
- The discussion speculates on future advancements in autonomous technology, including potential networked intelligence among vehicles to improve traffic management.
---
This summary encapsulates the core discussions from the podcast with Alex Kendall, highlighting the significant advancements in autonomous driving technology through end-to-end learning and the evolving role of AI in society.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00You know, if you're building a vertically integrated robotic solution, maybe you can go deep, but our ambition is to be the embodied AI foundation model for all of the best fleets and manufacturers. is around the world. And to do that, unless we want to overload the company by building a separate neural network for each application, we need to be able to generalize. We need to be able to amortize our cost over one large intelligence and to be able to very quickly adapt to each different application that our customers care about. That's what we're trying to push.
0:45Today we're talking with Alex Kendall, CEO of Wave, about the shift from software 1.0 to 2.0, or from classical machine learning to end-to-end neural networks in autonomous driving. Wave sells an autonomous driving stack to auto OEMs, similar to Tesla FSD, but for non-Tesla automobiles. Major car manufacturers globally, like Nissan, are choosing Wave to power their AV stacks. Alex started Wave back in 2017 when most self-driving software stacks were massive hand-coded C++ code bases covering every possible edge case like navigating around double-parked cars. Alex bet the farm from the beginning on an end-to-end neural net approach to self-driving and on the use of synthetic data and world models as the ultimate path to generalization and scaling.
1:25Today, that architecture is reshaping AV and all of physical AI, including robotics. Enjoy the show. Alex, thanks for joining us on the show. Hey, Pat. Hey, Sonia. One of the things that is very special about your company is that it sort of typifies AV2.0, meaning a new architectural approach that I think is kind of demonstrated to be superior to the AV1.0 approach that people toiled with for so many years. Can we just start by defining what was AV1.0? What is AV2.0? For sure. when we started the company in 2017, the opening pitch in our seed deck was all about the classical robotics approach at the time was to take a perception, planning, mapping, control, essentially break down the autonomy problem into a bunch of different components and largely hand engineer them.
2:16And our pitch was that, okay, we think that the future of robotics is not going to be a system that's hand engineered to drive with a lot of infrastructure like high definition maps. But instead, we thought that the future of robots would be intelligent machines that have the onboard intelligence to make their own decisions. And of course, the best way we know how to build an AI system is with end-to-end deep learning. So for the last 10 years, we've been promoting an approach, next generation approach, AV 2.0, that replaces that stack with one end-to-end neural network. Now, of course, that may seem more obvious today, but it has been contrarian for many, many years.
2:54But I think today it's maybe unfair to make that basic distinction because, of course, anyone who's worth a grain of salt will use deep learning in various parts of the stack. But what you see in more incumbent solutions to autonomous driving is, of course, deep learning for perception and maybe for each different component, but still a lot of hand interfaces, still a lot of infrastructure on hide-efficient maps, and perhaps reliance on a lot of hardware. So our solution is is still somewhat moved on. But today, rather than just being an end-to-end network, today, of course, we start to talk about foundation models.
3:29We start to talk about more of a general purpose intelligence, one that can understand not just how to drive that car, but many cars with different sensor architectures, with different use cases. And so really, it all boils down to how do we build the most intelligent robot that can scale without needing onerous infrastructure? So WAVE is sensor inputs, motion output, gigantic neural net in the middle. That's right. At a very simple level. But some of the interesting things you see that are maybe different from the story we've all heard with large language models is with autonomous driving, of course, there are some interesting new factors.
4:10One is, of course, safety. The system we need to make sure is safe by design. And what that means is that we can't just pump more data in and hope that hallucinations go away. But we need to design an architecture that is still end-to-end data driven, but is both functionally safe and we can build a robust behavioral safety case. So that introduces some interesting architectural challenges. And then of course, we also need to run real-time on-board a robot, on-board a vehicle. And so dealing with the on-board compute and on-board sensor limitations make it an interesting challenge. But yes, it's the same narrative we're seeing playing out in robotics that we've seen play out in all these other AI fields, like language or game playing agents.
4:55It's that an end-to-end data-learn solution is out-competing anything we can hand code. And what we're excited to be pioneering is that the exact same narrative here in robotics and autonomous vehicles. And when you guys started this in 2017, and it was a very contrarian approach. When people from the industry said, well, that'll never work because, how did they finish that sentence? I could count hundreds of those meetings. Yeah, typical arguments were, look, it's not safe. It's not interpretable. Can't understand what it's doing or even simply it doesn't make sense. We haven't heard of this AI thing.
5:37And look, I think five, 10 years ago, it was probably reasonable to say end-to-end deep learning wasn't interpretable, but I don't think that's true today. I think today we have a lot of really great tools for understanding and responding to insights about the way these deep learning systems reason. But moreover, I think if you have the ambition to build any intelligent machine, I think it's naive to think you can build a complex intelligent machine and actually make it, you know, let's say strictly interpretable to the point where you can point to a single line of code or a single thing that causally made the outcome occur.
6:14The beauty of intelligent machines is that they are so wonderfully complex. And there, I think the way that we're going to not just design them, but understand them is through a data-driven structure. Can you say more about the before and after of the AV 1.0 stack and the billions of lines of code that goes into those systems versus the 2.0 systems today? And how quickly is that changing? Because my sense is that deep learning, large neural nets hitting the physical economy is a much more recent phenomenon than people might appreciate. Well, especially when you think about the path to distribution and deploying these systems, I mean, the automotive industry has gone through a seismic shift and bringing out software-defined vehicles and the right hardware on these cars to be able to make them drive maybe one uh common uh point of debate is is is a camera only or camera radar lidar as a sensor approach to autonomy and um uh just to be clear on on our position uh wave we we want to build an ai that can understand all kinds of different sensor architectures there's going to be sometimes where a camera only solution makes sense sometimes where camera radar, LiDAR, and we train our embodied AI model on all of those permutations from very diverse data sources.
7:31And the car we just drove in is a camera only stack. We've got other cars that we work on with partners that have radar and LiDAR. And of course, there's different trade-offs that you take there. But more generally, we're seeing mass produced cars from the best manufacturers around the world have a GPU onboard, have surround camera, surround radar, and sometimes a front lidar and what's beautiful about that is there's now the opportunity to see this ai come out and benefit people around the world i think that kind of software-defined infrastructure is happening in automotive has perhaps not yet happened to the same degree in other robotics verticals but i'm sure the market's going to move that way as well and in general having the right level of computing infrastructure in a scalable way and opening up these platforms to ai i think is is what's you know, really making this possible.
8:18And that's gone through a tipping point in the last couple of years. And, you know, your perspective of AV2.0 has flipped from contrarian to, I'd say, consensus. Maybe in the last two or three years, do you think it was FSD12 that did it? Or when did that mindset start to shift? I miss the contrarian day. But even today, I still, I was in a conversation this morning where I still see a lot of folks still say yes we need end-to-end AI they brought the you know the big tech narrative around the future of AI but they say things like we need end-to-end AI with with hard constraints or with safety guarantees and and still there's this there still can be some um you know belief that some hybrid approach is the way to go where uh where you want to try and try and take a rules-based stack and an end-to-end learn stack but often these approaches can get the worst of both worlds or just add cost and complexity so um you know i still think there is a distribution in the market of those that are leaning and moving fast and those that are uh you know are perhaps you know have some some catching up to do um but of course crediting the the breakthrough that uh all of us that have been working in deep learning that that really made this world changing and mainstream of course we've got to credit the large language model breakthroughs um i think they've inspired the world and opened up the market's mind to be curious about this technology.
9:42But also what we've been doing at WAVE, you know, a year ago, we were just driving in central London. Central London, I think, is a great proving ground because it's this unstructured, incredibly complex and dynamic city that our AI has learned to navigate around very smoothly, safely and reliably. But in the last year, we've taken it to highways to Europe, Japan, North America. Our cars were in New York City last week driving around there. And so bringing it global, being able to take it to different manufacturers' vehicles and show a product-like experience, this growth is, I think, also really opened up a lot of inspiration around the world.
10:23Why is it that you're able to launch in hundreds of cities worldwide and some of the AV1.0 companies need to actually go out and build an HGMAP? Just say a word on the difference and how technical differences are actually leading to differences, how the machine's able to learn and how you're able to roll out. Autonomous driving is all about generalization. Generalization means being able to reason about or understand something you've never seen before. Every time you go for a drive, you're going to see something new for the first time. What did we see today? We saw a road worker rolling out some carpet thing in front of the road, but on a pedestrian crossing, but not wanting to step out.
11:04And we had to reason about could we pass them without yielding, for example. There's just an example from earlier today, but you could think about all the new things you see on the roads every time you drive. You're never going to see every experience in your training data. So that means that you have to be able to reason and generalize to things you haven't seen before to be safe, to be useful around the world. and that's what has motivated our entire approach. So whether it's us, a manufacturer giving us one of their vehicles and within a couple of months us being able to drive it on the road.
11:35A couple of weeks ago in September this year, we unveiled a vehicle to media with Nissan in Tokyo. Just four months earlier was the first time we'd even driven in Tokyo and got hands on this vehicle. Four months later, we were having media drive in the car, experience it And that was a new country and a new vehicle for us. So what that showed is that our AI was able to generalize. It's trained on very diverse data from around the world. It's trained on diverse sensor sets, vehicles. And so it was able to understand that vehicle's new sensor distribution and, of course, the complexity of driving around in central Tokyo.
12:13So I think that's a really great demonstration of generalization. And if we think about, you know, if you're building a vertically integrated robotic solution, maybe you can go deep. But our ambition is to be the embodied AI foundation model for all of the best fleets and manufacturers around the world. And to do that, unless we want to overload the company by building a separate neural network for each application, we need to be able to generalize. We need to be able to amortize our cost over one large intelligence and to be able to very quickly adapt to each different application that our customers care about.
12:45That's what we're trying to push. You mentioned reasoning in there in terms of how the model is reasoning through, you know, there's construction work for what do I do now? In the LLM world, obviously reasoning is its own separate track of lots of scaling inference time computes techniques. Are you deliberately training your models to reason as an emergent property, emergent behavior of the models? Like say more about what you mean about reasoning. We are. And I think reasoning in the physical world can be really well expressed as a world model. In 2018, we put our very first world model approach on the road.
13:21It was a very small 100 ,000 parameter neural network that could simulate a 30 by three pixel image of a road in front of us. But we were able to use it as this internal simulator to train a model-based reinforcement learning algorithm uh there's a a fun blog post if you want to see the history on that but fast forward to today uh and we've developed a gaia it's a full generative world model that's able to simulate multiple camera and sensors and very rich and diverse environments you can control it and prompt the different agents or seen in it and um and that's an example of reasoning where we can train in the ability to simulate how the world works and what's going to happen next what happens when you bring this kind of representation on the road is you get some really nice emergent behavior.
14:09Like today, when we saw we were driving around unprotected turns that were included, you saw the car nudge forward until it could see for itself and then completed the turn. Or when it's foggy in London, you see the car slow down and drive to what it can reason about. And by training it with that level of understanding, it gives that level of emergent behavior that helps it really understand particularly complex multi-agent scenarios. I think that's key for getting safe and smooth autonomous driving. So the world models are really key to teaching the model how to reason through their new scenarios.
14:44A hundred percent. You mentioned earlier the diversity of your data. Say a word about where all the data comes from. It's becoming like an enormous amount of data because of course, unlike the language domain or image domain, when we're dealing with um you know a typical self-driving car that has a dozen multiple megapixel cameras radar maybe a lidar you know you're dealing with when you aggregate that up it's very quickly tens or hundreds of petabytes of data so it's it's an enormous amount of data you have to train on but it's the diversity that's really key and we've solved for diversity in two ways first one is by becoming a trusted partner across the industry and aggregating data across many different sources from dash cams to fleets to manufacturers to robot operators.
15:32And the second one is being able to filter and really understand the data. Here we've really worked hard to develop different unsupervised learning techniques to be able to cluster and find unusual or anomaly experiences. And of course, find the scenarios that our system is performing poorly at and then drive the learning curriculum on those. But yeah, today we learn from a diverse set of vehicles, a diverse set of sensor architectures of countries, and that's really one of the key things that drives the level of generalization. Does the increased growth of world models and simulated data, does that mean that you just don't need as many actual on-road miles?
16:16I think there's two sides to that question, right on the one side yes efficiency learning efficiency really matters the second you can't only rely on learning efficiency um at the limit if we take our current approach and just scale it up i'm i'm sure it'll produce generic level five driving um uh you know at the limit if you have unlimited training data this is really just a lookup data table with some some prior experience but that's not economically or technically feasible and so the question is how can you train this to be the most efficient data efficient system because i think efficiency will lead to not just improved cost but faster time to market and more intelligence so um efficiency comes from a number of different factors there's uh most importantly how the data curriculum you can place but then the the learning algorithms how do you magnify the learning you have and i think world models are a really great opportunity for that they generate synthetic data and synthetic understanding that doesn't replace real-world data, but it recombines it and magnifies it in new ways.
17:16It lets you pull in interesting insights, and I think these kind of approaches can really, really improve data efficiency. But across the board, I think working under resource constraints has forced our team to develop so many innovations. I'd also call out just the workflow, because in traditional robotics, when you're tuning parameters or algorithms or designing geometric maps and things like this, there's very well established cultures and workflows. But our team, when we have 50 model developers working on one main production model, or when we have an end-to-end net that we need to understand and introspect, or even the way that we deploy these systems to simulation or to the road and feedback, we've developed the entire culture from the ground up at Wave has been developed for embodied AI, for end-to-end deep learning for driving.
18:07the data infrastructure, the simulation, the safety licensing before we put systems on the road. This has not been a hedge or a side bet for us, but this is the entire essence of our culture. And I think doing this under resource constraints and doing this with full mission-driven conviction has led to a bunch of interesting innovations that, look, getting to where we are today, everything is about iteration speed. Speaking of your culture, I'm picturing a bunch of AI research types, machine learning engineers, that sort of thing. How does the culture of your organization differ from similar applied lab type environments given the customer base that you serve, given that you're going after the automotive industry specifically with all of its quirks around supply chain and all of its requirements around safety?
19:00And so how does that influence the culture of your business? hugely uh in fact you know for the first few years of wave uh you know we were really a group of um passionate embodied ai researchers but in the last couple of years um i'm really really proud of how our team has built out deep expertise in understanding the automotive industry but also the ability to reliably deliver to our partners there and that's a that's a different culture it's a culture i've really grown to respect because when you're building millions of cars the level of reliability and MTTF you need there is extraordinary.
19:38What have you all learned from them? I mean, I'm sure part of your job is to teach them about what's going on in the world of AI. What have you learned from them? I think some of the main things I'll call out have been efficiency and reliability. The difference between technology and a product would be some of the main themes. I mean, the level of reliability required, but also So the level of quality that is seen to really robustly prove these systems out before deployment and the pride that these companies take in that has been exceptional. Another thing has been perhaps the sense of brand differentiation and the desire for, you know, look, do you want your car to drive?
20:22How can your driving personality really match the brand's preferences? How can you provide that experience that really gives brand differentiation? And the great news is that I think we've been able to riff and brainstorm off these and come up with some really neat technical ideas down that vein. But yeah, ultimately, safe, high quality and personalizable AI has been some great feedback we've got from the industry. Can you talk about your path to market actually in partnering with the auto OEMs? How did you decide to do that? And then how do you think the market landscape will play out for how autonomy rolls out?
21:00Yeah, of course. Great question, Sonia, because since the beginning of WAVE, we've been focused on the pitch I gave around end-to-end deep learning being the approach to autonomy, but we've tried a number of different go-to-market approaches over the years. But in the last couple of years, I've been hugely energized about working and partnering with the biggest and best consumer automotive manufacturers around the world. Why is that? Well, I mentioned how they've begun to introduce software-defined vehicles. so they have the infrastructure to work with autonomy there's the market belief that this is a technology that can can really thrive and also it's the chance to get to scale far beyond what we're seeing with the city by city robo taxis we're seeing right now but moreover these are oems that are investing in the right infrastructure to go from not just driver assistance but to eyes off autonomy where you can actually you know take liability for the drive and give give the user are safe and give them time back from their driving experience.
22:02So that's awesome. I think when you think about the market, you know, there are 90 million cars built each year. And so, you know, some manufacturers that are building the autonomy systems themselves, like Tesla, built a couple of million, but the vast majority of the market, I think there's an opportunity to partner, to work with some of these innovative platforms and to bring our AI to market to make these autonomous products possible. And it will only grow from there. These manufacturers don't want to stop at driver assistance. We're working together to build eyes off and driverless robotaxi products.
22:32But the key thing is that by avoiding retrofitting our own hardware on these vehicles, by putting them in natively as a software integration, we can move fast at scale. We can build low-cost vehicles that can be homologated all around the world. I think this is going to be the path to see tens and hundreds of thousands of robotaxis rolled out around the world at an affordable price. And of course, this is all possible because of the level of generalization that this AI enables. Tesla FSD is just such a game-changing product. And my friends that have it, they can't imagine driving any other way.
23:06And so it's really cool that you're going to empower the 88 million other vehicles sold every year to be able to sell that experience as well. 100%. It's one of those things that a lot of people would jump in our car and come for a drive with some being skeptical about autonomy. but without exception, they step out with a smile on their face. It's a magical experience. And yeah, I can't wait for people to be able to try it around the world and make autonomy, not just a robotaxi tourism experience, but bring this experience to people in, yeah, eventually every city. What do you make of the sensor fusion confusion debate?
23:42You know, the one that plays it on Twitter every year or so of Tesla gets confused if there's both camera and LiDAR coming in? Sorry, radar. I think it's the wrong debate to be having. It's not the frontier question. The industry is really, I guess outside of Tesla, has really coalesced around a common architecture of a surround camera, surround radar, and a front-facing LiDAR stack. Now, this costs under$2 ,000. So it's automotive-grade components, not the retrofit robotaxi components you see today. but having a front-air GPU compute, automotive grade GPU on the car and that kind of sensor architecture is a really great platform to build L3, L4 autonomy, eyes off or driverless.
24:25It gives you the necessary redundancy. It lets you deal with edge cases that, you know, cameras alone, I agree they can get you to human level but we want to go beyond human level. And so I think this kind of architecture is affordable, scalable. It's got the supply chain for mass manufacturer and it can eliminate, I think, eliminate all accidents and really drive superhuman levels of performance. So that's what we're seeing many manufacturers bring out on their vehicles and we're integrating our AI. Of course, for a driver assistance system, camera only can work for a human level driverless system or of course, I should clarify, 90 something, you can look at different stats, but 95 % or above accidents, unfortunately are caused by human error so not only can you be human level but you can eliminate a lot of human attention and and accidents caused by that but there are still accidents that um to be able to solve would require perception capabilities that go beyond uh beyond vision and if we want to tackle that long tail um there are many ways to solve it one of the ways would be to bring in um some other sensing modalities by like radar and lidar so um you know we're excited to be working with those kind of platforms um but crucially natively integrated into uh in into the OEMs vehicles themselves.
25:41Is it the same neural net that can drive on one OEM's car and another's car? And how does that even work? Because I imagine each vehicle has, you know, slightly different position cameras, things like that. It comes from the same family. So we train a very large scale. We regularly train very large scale models. Of course, we iterate them on them month on month. But, you know, that's one model that's common to all of the fleets that we work with. but as you optimize to a specific sensor set or a specific embedded target of course you can start to specialize the model but the beauty is that 99 plus of the cost and the time and the effort is training their base model and then we can build very efficient personalization to the specific customer and so this lets us the scale but gives us the ability to you know squeeze it to very efficient real-time platforms and make it adapted to a specific specific use case.
26:36Are you going to let pat personalize a super aggressive driver model need to what driving style would you like pat yeah pretty aggressive safe very safe but you know we can do that we uh yeah we we find it's really funny when you when you build distributions around driving behavior um yeah you can you can really tell uh from the human training data we have you can really tell when it goes from being um helpfully assertive let's say to uh unhelpfully aggressive and we can we can draw clean line there yeah there you go what about you sonia how was the drive we just had fantastic it was it was it was comfortable it was it was safe it was and it felt very human actually like the way it was kind of nudging up when it couldn't see on the turn it was very human yeah well it's uh as you know complex as we can get in in silicon valley but come to tokyo or london or i was in the weekend in downtown san francisco and um yeah you really need uh the ability to predict and reason about other folks around you to be able to drive in a human-like way.
27:35And what we find is that if you're not able to smoothly go around double parked vehicles or deal with other dynamic obstacles, or even the prevailing row of traffic might not be aligned to the specific lane, but maybe there's a human-like way of driving, then what's awesome about the intelligence that we've built is it's able to reason about these things and keep the traffic flowing, keep interacting with road users in a very human-like way. I think this is going to be key for societies to accept and love robo taxis um i can't wait to make that a reality are there any specific corner cases that your cars have a hard time with today there are loads and it's it's uh uh really hard to generically talk about one because that's that's so rare yeah uh um you know if i was it's very hard to say oh it's always these types because it's always you know a corner case is a couple of edge cases coming together in a corner and it's it's always confounding factors when you get something really obscure but we're driving we've driven in over 500 cities this year uh and so when you're driving at that level of scale of course you see things that you've never seen before road signs are written in a new language um actually maybe one way to break it down is often we talk about driving broken down into safety utility and flow yeah safety being of course um safety critical behavior flow being the style of driving is it smooth is it enjoyable and then utility being the navigation and road semantics and safety and flow we found generalized exceptionally well throughout the world we get almost uniform metrics in every country we operate in in terms of safety and flow or comfort of the drive but utility has been the really interesting one as we've gone global how do you navigate and how do you deal with road signs how do you read different languages how do you deal with different driving cultures and so that's the one that's been interesting but from uh we we published some results about this when we went from the uk to the us we needed uh hundreds of hours of data to be able to drive um you know within 10 of our frontier performance um but then when we went to uh europe into germany of course we'd already learned to drive on the right side of the road coming to the us would learn to do right turns at red lights then coming to germany we had to learn to still drive on the right side of the road but of course you can't turn right at a red light there but then on the autobahn you'd like this uh you need to drive we drive today up to 140 uh so uh so pretty fast there but um yeah uh it gets more efficient each time with exponentially less data in each new market because you've seen some of those things before yeah you mentioned the beginning that large language models were part of what flipped your approach from uh from contrarian to to consensus are you integrating large language models at all into your models?
30:16And I know some of the robotics companies that are getting started now are starting from this VLA, VLM base. Is that part of your architecture? 100%. In 2021, we started working on language for driving. I remember my team came to me at the time and said, hey, we should start a project on language. I said, no, no, no, guys, startups all about focus, keep focus. But they actually gave some pretty compelling arguments. So we started to play around with these things. And a year or so later, we released Lingo, which is the first vision language action model in autonomous driving. And what was special about this model was it could not only drive a car, see the wheel drive a car, but also converse in language.
30:55And it'll let you talk to it, ask it questions. You know, what are you finding that's risky? What's going to happen next? Or even it could commentate your drive. And what's interesting about this is that, so there's a few benefits. One is bringing language into pre-training, of course, just improves the representations. power gives you know more interesting information to learn from than than just imagery alone but then second aligning the representation with language opens up a ton of interesting product features it enables a you know you to create a chauffeur experience where you could actually talk to your driver you know no longer do you need a phd in robotics to understand the system but actually you can just talk to it and and like ask it to drive uh pat if you want to if you want or race around the commute super fast, then you can demand that.
31:40But then third, it gives you a really nice introspection tool where you can start to actually, you know, you could imagine regulators or our engineering team converses with the system and language to really diagnose why it's doing what it's doing or get it to explain its reasoning. So I think these are really clear benefits, which we're really excited to be pushing. That's super cool. And you're running it on the embedded compute. We are. So we've put out demos that run off-board. Onboard's challenging with what's in the automotive market today, but some of the next generation compute, for example, the NVIDIA Thor that our next gen development vehicle is going to be with, will be large enough to run it on board.
Read the full transcript
32:13That's going to be cool. Very cool. You've talked about how autonomous driving sort of provides a path to more generalized, embodied AI. Can you paint that picture for us, how you go from autonomous driving to humanoid robots or whatever other things you might want to embody AI? I think we're going to be in the future looking at a ton of interesting use cases for robotics. What we're seeing is that mobility is becoming possible, I think, much before manipulation. Manipulation is challenging in terms of access to data, global supply chains for hardware, and actually even the hardware designs themselves.
32:55I think tactile sensing is still a really hard challenge. but inevitably it'll be a massive transformative thing but maybe you know maybe is it the maturity of where self-driving was in 2015 but today you know our system is rapidly becoming a general purpose navigation agent giving it an arbitrary a sense of view and a goal condition it's able to produce a safe trajectory so I think we're going to see a rapid advancement from not just consumer automotive robotaxis you think about trucking and other applications but you know this AI will enable manufacturers and fleets who want to build robots in any kind of mobility application and of course we we're really excited to be working with uh um you know frontier developers and applications over time as you as you go out across that robotic stack um and i expect we'll see more maturity in the coming years from manufacturing and manipulation use cases as well but in the end i think the benefits of having a large foundation model that certainly in In automotive, I think we have access to the largest robot and data supply chain.
34:01And so we're really lucky in that regard to be able to push forward the intelligence there. But generalizing that intelligence to new applications, I think there'll be benefits from the model being able to experience multiple different verticals. And it'll only make it more general purpose. Any applications you're excited about? I mean, I'm psyched to have humanoid robots walking around. Yeah, me too. I think they're going to be neat. you know whichever form factor i think humanoids will play a big part i think other forms of locomotion as well and then manipulation there's some some really interesting challenges in those space but i think the same story is going to play out uh you know a working on a narrow application you know like when self-driving went to phoenix arizona and put in a ton of infrastructure and expensive hardware to make it work is is going to i think have limited runway but working on general purpose lean low-cost hardware stacks that really focus on making the system most intelligent and robust you know i think this is the recipe for scale um so yeah let's uh let's watch that space yeah do you think there are major research breakthroughs needed to reach kind of physical agi so to speak and if so what do you think is the most promising direction absolutely i do i think there's so much more ground to scale up the current approaches and we'll do that but uh i think we'll get compounding returns from i always usually talk about four factors that that drive performance there's of course data and compute but then also the the algorithmic capabilities and the embodiment the what is the hardware and capability on the on the robots and i think we need to push all four now on the algorithmic side um there is so many opportunities for for growth i think um a key one is measurement uh how do you actually uh measure and quantify these systems how do you respond quickly find regressions be able to really have a simulator that that closes the the real world gap at scale and can run efficiently i mean it's no secret that these generative world models are very compute intensive but having a good measurement system will just drive efficiency and iteration speed so that's a that's a key one people often talk about being a chicken and egg if you have a perfect simulator you've solved self-driving and vice versa and i i really believe that um i think alpha go showed that when you have a perfect simulator you can just solve problems through monte carlo tree search um and so i you know i i think that's gonna be the case in robotics as well um so yeah one is one is measurement uh another pillar is building more generality into the model how can you build out more modalities uh and align those different modalities in their reasoning um i think this is going to open up new use cases, particularly when it comes to human-robot interaction and navigation.
36:48I was going back to the utility problem before. Some of these things I'm super excited about. And then the last one is just engineering efficiency. I mean, training these systems and the data requirements is extraordinary. And so I wouldn't understate, I think the most sexy part of this problem is the efficient infrastructure to train and serve these models and getting that right. It's, yeah, I think it's a real competitive advantage or disadvantage. We started by talking about AV 2.0. Someday, I imagine we might be talking about AV 3.0. What could AV 3.0 look like? If you go 5, 10, 15 years in the future, are there any other big leaps in this industry that you think we'll see?
37:34You said that with such deadpan. uh av so the whole premise of av 2.0 was all about putting the intelligence on the car and not needing infrastructure and a ton of uh you know ton of overcooked hardware but really making the system intelligent uh and so i think we're seeing that emerge now with the system that can generalize to the world with all of the onboard scalable um intelligence and compute if i were speculate where av3.0 we haven't thought we haven't um sort of thought about in depth uh lately but one idea could be taking the intelligence outside the car so i mean when you start to have you know majority prevailing autonomous vehicles you could imagine a ton of new things you could do when they start to communicate when they start to interact with each other you know why do we need traffic lights in the future if they can coordinate why do we need uh you know all these sensors if you can actually just communicate with the av in front of you to be able to see around corners i mean of course i'm speculating here it opens up tons of interesting cyber security questions um communication latency questions things like that but uh i don't know i'm uh i'm all up for embodied ai and um uh if we can build a safer and more accessible system by taking the intelligence uh not only in the car but but beyond maybe um maybe that's a path let's see i think that's really interesting if av3.0 is the point at which it's sort of a mesh network and you know At that point, maybe humans aren't allowed to drive because they can't communicate with the mesh network the same way that the robots can.
39:06Or maybe there are special places that humans go to drive just for recreational purposes. Always trying to do a taste drug. It's all autonomous. Yeah, interesting. How do you hire and how do you attract people with how hot the AI market is these days? I love that question because at the end of the day, our team is our product. Our team are the most important thing to making this possible. And we talk a lot about at Wave about being a place where you can do the best work of your career. And what that means for me in Embodied AI is having a set of colleagues around you that inspire and excite world-class in what they do, having the right resources, the right culture unblock you.
39:44But I think uniquely at Wave, we are able to bring together really a frontier AI environment with a near-term product opportunity in automotive. So if you want to work on intelligent machines and see your system brought out with the scale of impact of ChatGPT in robotics, I think this is a place where we can do it. The other thing is that we've gone global. I mean, we have teams in London, Stuttgart, Tel Aviv, Vancouver, Tokyo, Silicon Valley, and, you know, wherever there's almost some of the major AI and automotive hubs. and we're really looking to build a global culture that can bring this product to the world, work with customers around the world and most importantly, collaborate with the very best people.
40:35Yeah, so anyone who's interested in pioneering embodied AI, pushing the frontiers and actually turning it into a game-changing product, come chat, we'd love to speak. Wonderful. Alex, you've believed in the future for end-to-end neural nets in self-driving and in the physical economy for longer than just about anybody. And it must be incredibly fulfilling to see that vision start to come to life. Congratulations and thank you for joining us. Thank you, Sonia. Thank you, Pat. It's such a privilege.
41:19Thank you.
From the publisher
Alex Kendall founded Wayve in 2017 with a contrarian vision: replace the hand-engineered autonomous vehicle stack with end-to-end deep learning. While AV 1.0 companies relied on HD maps, LiDAR retrofits, and city-by-city deployments, Wayve built a generalization-first approach that can adapt to new vehicles and cities in weeks. Alex explains how world models enable reasoning in complex scenarios, why partnering with automotive OEMs creates a path to scale beyond robo-taxis, and how language integration opens up new product possibilities. From driving in 500 cities to deploying with manufacturers like Nissan, Wayve demonstrates how the same AI breakthroughs powering LLMs are transforming the physical economy.
Hosted by: Pat Grady and Sonya Huang




