Wayve CEO Alex Kendall on Making a Splash in Autonomous Vehicles - Ep. 209

6 Dec 2023 · 32 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

NVIDIA AI Podcast: Episode 209 - Wayve CEO Alex Kendall on Making a Splash in Autonomous Vehicles

Episode Overview

  • Host: Katie Burke Washabaugh
  • Guest: Alex Kendall, co-founder and CEO of Wayve
  • Focus: The emergence of AV 2.0, a new paradigm in autonomous vehicle technology characterized by comprehensive AI models that integrate perception, planning, and control.

Key Themes and Discussions

  1. Introduction to AV 2.0
  2. Definition: AV 2.0 signifies a shift from the previous approach (AV 1.0) dedicated to perfecting vehicle perception to a more integrated system that enables real-time decision-making.
  3. Embodied AI Concept: The central idea is to equip AI with a physical interface to interact with its environment, enhancing its ability to learn and adapt to dynamic situations.
  1. Wayve’s Role in AV 2.0
  2. Mission Statement: Wayve aims to revolutionize autonomous vehicle development through deep learning and AI.
  3. Historical Context: Founded in 2017, Wayve demonstrated its AI-powered driving system that generalizes across various cities in the UK.
  1. Hardware and Software Considerations
  2. Hardware-Software Problem: Kendall emphasizes the importance of considering hardware and software separately yet collaboratively. High-quality sensors alone are ineffective without the right AI software.
  3. Sensor Setup: Initially focusing on a camera-radar approach, Kendall believes this is the most scalable and effective method for AV deployment, eschewing complex setups like LIDAR for now.
  1. Generative AI and Synthetic Data
  2. Role of Generative AI: This technology allows for the creation of synthetic driving scenarios, which helps in training the AI system on rare edge cases not commonly encountered in data collection.
  3. Example Use Cases: Generative AI can simulate diverse environments (e.g., snowy conditions with crowded streets) to enhance the robustness of the AV systems.
  1. Innovations by Wayve
  2. GAIA-1: A generative world model designed for developing autonomous vehicles, capable of creating realistic, dynamic simulations.
  3. LINGO-1: An AI model that allows passengers to interact with the vehicle using natural language, enhancing the explainability and trustworthiness of the AI systems.
  1. Scaling and Safety
  2. Expansion to New Cities: Wayve’s ability to scale operations across different urban environments without extensive data collection fleets is attributed to the AI's capacity to generalize learned behaviors and concepts.
  3. Safety and Trust: Kendall argues that improving safety measures and building public trust are crucial for the successful deployment of autonomous vehicles.
  1. Future Outlook
  2. Investments in AI: Wayve plans to continue investing in AI innovations, focusing on enhancing training data, compute capabilities, and generative AI technologies.
  3. Five-Year Vision: The expectation is that in five years, AI will be able to operate vehicles in a trustworthy manner, with the capability for natural language interaction, effectively integrating AI into daily life.

Key Takeaways

  • The transition from AV 1.0 to AV 2.0 represents a significant shift in the understanding and development of autonomous vehicles.
  • Embodied AI plays a crucial role in enabling vehicles to operate safely in dynamic, real-world environments.
  • Generative AI is critical for enhancing the training and validation processes of autonomous vehicles, particularly for rare edge cases.
  • Wayve is leading innovations that can dramatically improve the scalability and safety of self-driving technology.
  • The future of autonomous vehicles is promising, with advancements in AI likely to redefine transportation and user interaction with technology.

Additional Resources

  • To learn more about Wayve and its initiatives, visit [Wayve's website](https://www.wayve.ai).
  • For more insights on NVIDIA's Inception startup accelerator program, check [NVIDIA Startups](https://www.nvidia.com/en-us/startups/).
  • Alex Kendall will be featured in an in-depth session at GTC 2024, focusing on the impact of generative AI on autonomous vehicle development.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Welcome to the AI Podcast. This is Katie Burke Washabah, your host for all things autonomous We're more than a decade into the development of autonomous vehicle technology, and the industry is undergoing a shift in its approach to this highly complex AI problem. Referred to as AV 2.0, this next generation of autonomy focuses on large, unified AI models that control multiple parts of the vehicle stack, such as perception plus planning and control. Earlier development, on the other hand, focused on these layers separately. This transition to a unified architecture is mainly fueled by generative AI.

0:46Autonomous driving company Wave is at the forefront of AV 2.0. Founded in 2017, its mission is to reimagine the approach to AV development with deep learning and AI. In 2019, Wave first demonstrated its AI-powered driving system, which has since been able to generalize to different cities across the United Kingdom. Joining me today is Alex Kendall, co-founder and CEO of Wave. Prior to founding Wave, Alex was a research fellow at the University of Cambridge, where he earned his PhD in computer vision and robotics, as well as numerous awards and recognitions for his work. Hello, Alex. Welcome to the podcast.

1:25How are you today? Hi, Kasi. Thanks for the kind introduction. Very excited to have you on the show with Wave standing at this intersection of autonomous vehicles and generative AI. It's been fascinating to see how these two transformative technologies are coming together, which is a long way of saying that I have a lot of questions for you. So we will get started. We'll start off with something hopefully on the easier end. Where does the name Wave come from? I thought your introduction was great. And one of the things you said about autonomous driving being an AI problem, it's interesting how we may take that for granted today, but it's really not been how people have thought about the problem.

2:12I mean, when we started and said, look, we've got to take an AI approach to solving self-driving, it was seen as a very different and contrarian idea. So I'm just delighted to see the progress that we've been able to make in AI this year, and it should be a great conversation. As for the name, one of the hardest things to decide when you're starting a startup. But look, I am a very passionate amateur surfer and we're based in London. So there aren't any beaches here, sadly. Maybe there are where you live. But I thought the next best thing would be able to ride a wave home as you're going home from work.

2:49And maybe we'd be able to recreate that experience with an autonomous vehicle. Well, I'm based in Michigan, so we need a few waves here as well. We'll see what we can do about that. So reading through your company materials, you mentioned the term embodied AI to describe your AV platform. Could you walk me through exactly what you mean by embodied AI? Yeah, embodied AI is the idea of embodying or giving AI a physical interface, an interface to interact with the real world. So in practice, it's about not just creating AI that is in a software sandbox or it's something that operates on the internet, but something that's deployed in the physical world, whether this is in robotics or other applications.

3:30And this is really the way that we think you should think about self-driving cars. It opens up a number of new challenges, like with embodied AI systems, you can't just train them on a static train and test set, data set that's collected offline. You need to be able to learn to interact, understand, and to deal with the dynamic world that we live in. So I'd go as far as saying I think it's going to be the next forefront of AI and certainly been how we've been thinking about things in autonomous. So speaking of autonomous, what drove this shift from AV 1.0 to AV 2.0? And how does WAVE's approach to this problem differ from what we've traditionally seen in the industry?

4:13Yeah, good question. Six years ago when we started, I was looking at the autonomous vehicles at the time, and they come from some really pioneering work almost a decade ago in the DARPA Grand Challenges. And these were systems that everyone believed at the time that perception was the problem. If you create perfect perception, you can create perfect driving. So how do we solve perception? You just add new sensing modalities for every edge case, a different sensor for every problem, or maps that tell you exactly where to drive and what to look out for. And that was really where the industry was going.

4:43But if you take a step back and think about what do we expect from our embodied AI systems in the future, we should expect that they have the intelligence on board to make their own decisions. We should expect them to be able to operate safely, that we can build a level of trust with them, that they can interact in the dynamic worlds that we live in. Not systems that are only going to operate in an affluent region where there are grid-like streets where the sun always shines and you've been able to tell the car how to behave, but systems that can really be available to everyone everywhere. When you have that perspective in mind, it gives you the opportunity to rethink how to build technology to create that.

5:18And when we started at a time, I was fortunate enough to have worked on a bunch of machine learning technology and be inspired by some of the things that were becoming possible. It was becoming possible to build computer vision systems that could see the world for themselves and not have to rely on an HD map. It was possible to build systems that could learn to make decisions that are more complex than what we can hand engineer. And that's what really led to the genesis of being able to build an AV 2.0 system that could learn to drive in a way that generalizes to new cities, new use cases, and new vehicle types and bring the benefits of autonomy to the world.

5:51So as you mentioned earlier, this isn't just a perception problem. Both hardware and software have become equally important to developing the vehicle. What have you found is needed or required in terms of both hardware and software for this AV 2.0 approach? Yeah, the interesting thing is that I think you need to, and embodied AI is a hardware-software problem. You need to consider these things separately. There's no point just having a fixed hardware platform and then trying to develop the AI, or even thinking about how intelligent a system is by just platform specs alone. You know, a good example is the best eyes in the animal kingdom are from the mantis shrimp, quite a primitive animal.

6:34Their eyes have such insane resolution, high dynamic range, all of these kind of factors, yet their perception intelligence is relatively poor compared to humans where we have worse eyes but a much richer set of intelligence. And so it kind of illustrates that you've got to consider hardware and software together when you're thinking about the overall intelligence of the system. And the great thing about an AV 2.0 approach or a machine learning approach to this problem is that we can learn to deal with different modalities of sensing, particularly given we train these systems through self-supervision.

7:07They're largely trained on predicting what's going to happen next. If you feed in different sensing modalities, camera, camera radar, or even a LIDAR, you can learn to extract what information is useful from some versus others and in what situations you should rely on them. But it can also let you get more informative information out of fewer sensors. So it allows you to get a richer understanding from things like camera sensors. And it can allow you to build a safer system with fewer or less complex sensing setups. So what that means practically for us is that we think the choice of sensors is really a question of what is the most safe and scalable platform today.

7:46You want to have some redundant modalities to give you superhuman perception capabilities and some modality redundancy. But working on sensors that are unproven at scale is risky right now. So for us, that means a camera radar approach is the best way to get started. It gives you different modalities that give you active and passive sensing and that manufactured into most modern vehicles today. So it's already shown how to build them at scale. And with the power of AI, we can make a system safe and robust enough without needing to go to LIDAR, for example. But when these things change, when sensors improve, when new technologies come along and they are able to be scaled, of course, we'd love to bring them into the stack and to keep improving the system as long as it doesn't add a burden to the complexity and requirements on building fail-safes and things like that.

8:36But camera radar, I think, is going to be the most effective and scalable first scalable deployment of autonomy. So in addition to sensors, what does Wave look for in a compute platform to run its software? Yeah, it's interesting on the compute platform, right? I think what NVIDIA has built with Oren and what's soon going to be coming with Thor is just at the head of the game in terms of the most forward-looking future-proof technologies that allow you to run big neural networks. I think on an Oren, we can run a 10 billion parameter neural network in real time. It allows us to put the power of large language models, embodied AI, and systems that can generally understand dynamic driving scenes live on a vehicle that can run in real time.

9:23So certainly you look for these kind of factors, but also the way that it integrates into the system. You need to have seamless, low latency integration with your sensors. And when I say intelligence is a hardware-software thing, it's no use if you've got all the tops in the world to run your stack. but then you have to go through some intermediary hardware to be able to access your sensors. Then you're just going to add latency in your system and your system becomes sluggish and therefore not intelligent. So having that automotive integration, the platform, and certainly I'm a big fan of the Hyperion reference architecture that NVIDIA is putting together.

9:59I think those are great places to start. But to answer your question, I think you want to be investing in a high-tops environment to be able to run these technologies that are only going to grow in capabilities. And that's what we're looking for. Okay, so I'm going to put you in the hot seat for a second. With this AV2.0 approach, do you think that the AV challenge is solved? Or will this continued increase in sensor resolution and software complexity require us to keep pushing compute performance, keep sort of building that ladder to the moon? Yeah, if it was solved, you know, we could all go home and go and rest and relax.

10:42But I'm still, you know, there's still really exciting problems to still tackle. But what I would say is the things that we are seeing in terms of the graphs of performance versus data or performance versus compute that show the trajectory that these models are on. We're pretty excited about the things that we've demonstrated this year, whether it's Lingo bringing the power of language models into driving or Aya bringing the powers of generative AI to world modeling and understand the dynamics of the scenes or even the performance that we're seeing on the road, generalizing to different cities or vehicle types.

11:13These results are remarkable. And in the next year, we are going to be growing the compute, the parameters and the training data that are behind our neural networks, behind our AI models by 100x, two orders of magnitude in each of those factors. And, you know, the plots that we're seeing of performance versus data or performance versus compute are just going on a rocket ship straight up. And so, you know, what I'm excited about is the emergent behavior, the performance that we're going to see scale. And I think this is a narrative that isn't new to the world, right? We've seen the exact same narrative play out in large language models as they went from GPT-1 to 2 to 3 to 4.

11:52As these things scaled and grew, we saw incredible capabilities come out. And my take is we're going to see the exact same thing in embodied AI. That segues nicely into my next question. We've kind of touched on this here and there earlier, but what exactly is the role of generative AI in this new approach to AV development? Yeah, what is generative AI? I mean, if you take the technical definition, it's when you're generating content from some prior distribution. And that might be generating an image, it might be generating some text, or for us, it might be generating a motion plan for the vehicle to drive.

12:31So it's a bit of a broad catch-all phrase. And for vehicles, it's about generating a way to safely navigate a vehicle through an environment to achieve its goal. But you can also think about the applications for synthetic training data. We're a big believer at WAVE that synthetic data is going to play a huge role in both learning and validating the level of performance that we need to deploy these vehicles safely. And I think generative AI opens up some huge opportunities. Let me give you a few examples. We see lots of rare things on the road, whether it snowed in December in London. First time we saw snow of decent quantity in the four or five years we've been operating here.

13:13And so we got some data from those days. Or I was out driving the other day and we were going through a very busy Portobello Road Market Street in the centre of London. We saw crowds of pedestrians around us and it happened to be sunny that day. What Genitive AI can do can allow you to take those scenarios and combine them in new ways. You can take crowds of pedestrians and snow and bring them together and create a snowy crowded pedestrian scene. And that allows you to not necessarily generate things that are just completely net new, but it allows you to recombine things in new ways to become more robust to the environment.

13:43It allows you to recompose scenes in ways that allow you to experience or become more robust to the long tail of edge cases, the rare scenarios that really matter when it comes to self-driving. And this is the big problem that's challenged the industry. Prior AV 1.0 solutions that required brute force validation and armies of engineers that would meticulously look after every single last edge case. This allows us to move beyond that and use a data-driven approach to be able to create This generative flywheel that creates experience and proves safety for the long tail of edge cases. And so that's the opportunity that we're excited to be chasing at WAVE.

14:24Yeah, so with AV 1.0, we've seen so many companies grow and scale to cities outside of their home base by deploying large data collection fleets before moving on to operating their test vehicles in a new city. But there's a great video on your site, I encourage everyone listening to check out, of Wave being able to expand your operations to new cities without that sort of underlying data collection step. So could you talk about how you were able to kind of scale the system so quickly? This all comes back to, you know, the idea of, are you building an autonomous vehicle that's told how to drive and told how many lanes there are, which lane it should be in, and how she can navigate through an environment?

15:14Or are you building a vehicle that has the onboard intelligence to understand things for itself? And we've set off to make the bold step to choose the latter. And that means, you know, when we come to a traffic light, not telling the vehicle which lane it should be in or where the traffic light is, but giving it the ability to see the lights for itself, to understand the lanes, and if the navigation prompt it's been given is to turn left, to know it should be in the left lane. Or if the road markings say it can be the left two lanes, then for it to be in the left two lanes. And what we've seen is as we train these capabilities in our system through London experience, it learns these kind of concepts that generalise in a remarkable way.

15:52So it means that if we take the car to, you know, we've taken it to over 15 UK cities now. Let's say we go to Manchester or Cambridge or Liverpool or Leeds. We go to one of those cities and we saw that the car could take that knowledge in it, or you go to a traffic light it's never been to before in one of those cities, and pull up to it and spot the traffic light, understand its color, know where to stop, know which lane to be in, and all of these kind of things that you expect because it's learned that, you know, that generalized behavior. And it was really interesting to see, not just with different cities, but also with different vehicles.

16:21We started off operating on Jaguar I-Pace passenger vehicles. And, you know, as we've been able to explore and deploy some commercial trials in grocery delivery, we moved and created some vehicles as light commercial vans, 3.5 ton vans that we wanted to operate. And it was remarkable to see that we could also generalize from a passenger vehicle to a van platform with about, I don't know, 2.5 to 3 % of the training data. We were able to transfer that knowledge. It's a bit like how large language models learn to go from English to French to Mandarin. They take that general structure of language and learn how to make an interface or project to those different languages.

17:02You know, for us, it was learning the capability that could abstract away from the specific number of sensors or positions there on the vehicle, but learn to understand how to take in sensor data from around the vehicle and use it to make sense for that vehicle alone. An interesting thing was that, you know, it used to take probably a thousand hours of training before you see a behavior like overtaking a double park bus occur. And we saw after only 80 hours of training data on the van, it would learn to overtake a double park bus because it could bring that knowledge over from the broader training context.

17:32But also, because the van is larger, longer vehicle, when it was doing corners, it would also learn the specific van-related behavior, like taking a wider corner or dealing with the larger geometry of the vehicle. So it's really good at picking and choosing those things and learning what generalizes and what doesn't and how to adapt to new scenarios. I think that's what's required to bring autonomy to scale to the world. In June, you unveiled Gaia 1, a generative world model for developing autonomous vehicles. How did Gaia 1 come to be? Knowing what's necessary for AV simulation of the high fidelity to the real world, the necessity for assets to be physically based, how are you able to produce something like that purely generatively?

18:18Yeah, we've been working on these kind of things for a long time. I remember back in 2018, we put our very first model-based reinforcement learning system on one of our autonomous vehicles. It's tiny. It was like a thousand-parameter neural network that compressed a road scene down to something like three or four dimensions. It was a very tiny autoencoder at the time. We went and showed that, and the advantages of building a world model were clear to us at the time. It allows you to explicitly reason about the dynamics of the world. If you think about, take a large language model today, if you ask it what is, I don't know, four plus five, it might tell you four plus five equals eight.

18:57If you then put in a prompt and ask it, was that correct? It might go back and say, look, four plus five. No, it wasn't. It's actually nine. Let me change my answer to nine. And you see that kind of ability where it will put out answers in an auto-aggressive way, but it won't be able to understand the implication of its answers until you ask it to check. What a world model lets you do, or what a model-based system lets you do, is it lets you actually simulate forward the implication of your decisions. It lets you say, look, if I'm going to drive through here, what might happen to me and others in the world around me?

19:26And it lets you reason, understand, and check that behavior before you actually execute the decision. Now, that's not as important when you're dealing with large language models, because it's not safety critical. You can hallucinate a few things, and it's often not critical at the end of the day. But with self-driving cars, you know, you hallucinate something or you get something wrong and it can be a life or death situation. So that motivated us that we really wanted to build a system that would be robust, safe, and not just throw out, generate decisions on the fly, but actually understand the implications of its decisions.

19:59And that motivated us to build a model-based approach, a world model to policy learning to driving. Since 2018, we've had multiple iterations of it. Every year at the biggest computer vision conference, CBPR, we've brought back a new paper with new ideas, new iterations as we've gone from a small-scale learned thing to something that could understand full dynamic urban traffic scenes to something that could be multimodal and generate probabilistic futures to where we are today. Gaia 1 just represents the culmination of this journey. It's a truly incredible system, nine or 10 billion parameter system that can generate photorealistic, diverse, multimodal, probabilistic futures, understand dynamic interactions between the ego vehicle and other vehicles, model very thin structures like road signs, traffic poles, pumps in the road, things like this and just render this incredibly realistic and consistent 3D scene with all of the dynamic agents in it.

20:57And that level of capability, plus, by the way, it's controllable in terms of you can bias it to be in different weathers or ask for different, prompt for different actions or interactions. These kind of things is just an extraordinary step up for us and something that we are excited to see form the backbone of our synthetic data generation and world model understanding of while we're driving. You also recently announced the driving commentator LingoOne. Could you talk about how this works for those listening who may be unfamiliar? Yeah. A couple of years ago, this has also been a multi-year journey, but a couple of years ago when my team came to me and said, look, let's go build language for driving.

21:40I remember when Jamie and BJ came and said that to me, I said, look, this is going to be bizarre. This is why are we working on language? Look, let's stay focused on embodied AI. And they sort of talked it through. And the arguments they made were, look, explainability and understanding of a system is really important. It's important we can understand and these AI systems can be transparent about what they're thinking about. But also language represents a huge opportunity for tapping into more data, to learn from the vast knowledge we have online or in other sources, or to more directly interact with the robot.

22:18And in fact, we believe that the future of robotics interaction is likely to be through language. It's a very natural and accessible interface. You don't need a PhD in robotics to understand it, but you can just chat to it like you and I are chatting right now. And one of the things that I think has been really important in growing wave is yes, we have a very driven culture to deliver and improve and ship along our roadmap, but also a science team that's able to take a step back and work on challenging problems, things that might not necessarily work. I think our last step was something like 56 % of our science projects had a positive outcome.

22:55The others were negative results or things we learned from. And so all of these To be able to have the culture that can take risks, celebrate failure, look at long timeline things that aren't just looking to get a result for the next month's release. But, you know, in language, we started working on it two years ago and Lingo came out this year. It was a result of a two-year effort. It didn't start off with a big team. It was one or two people and it's sort of grown from there. But that kind of culture, I think, has been crucial to us to be able to achieve the results that we have now. And, you know, the things that I'm seeing coming down the pipe that are going to go beyond Gaia and Lingo are just another step up from that.

23:30And we're continuing to invest in the cutting edge science that makes embodied AI possible. So you mentioned the importance of being able to communicate with a robot without a PhD. Why is this type of communication or narration from the vehicle's AI so important in the context of autonomous vehicles? I remember the first time I went in an autonomous vehicle with no one in the front seat driving as an L4 vehicle, and it was mind-blowing. It was so cool. It's something that you very quickly become comfortable and familiar with, and it's a remarkable experience. I'm really excited about that future, but the thing that I noticed sitting in that vehicle is you have very little control.

24:16It's almost like you're driving on invisible railway tracks. You sit, you strap in, and then you're off. And it's different from, I don't know, being chauffeur driven or even driving your own vehicle where you really have more authority on what's going on. And so I think what language gives us the ability to do is to have more authority and control and ultimately build a better sense of trust with the robots that we're entrusting to delegate various parts of our lives to. So, for example, if you're worried about something, you'll say, why are you doing this? Or what do you think about this? Or have you seen this hazard?

24:50Or let's say you want to, can we take the next right? Or can we drive this way? Or I've spotted a better way to go. And all of those kind of things that empower us further, or even in a non-autonomous driving setting, but in a domestic robotic setting, if you want to customize or tell the robot how it should behave, right now your options are either to do a teach and repeat, where you put on a VR suit and transcribe the instructions through actually doing it yourself, or you go and you program it and then you do need a PhD in robotics or something like this to be able to do it. And language just democratizes this.

25:25You can use things that are natural to you and I. You can say, look, go pick up the dishes from the dishwasher and stack them in the cupboard. Very simple instruction, very accessible. And I think this is just going to create, make robotics to be a huge value add in our lives. This is what will unlock that kind of future. So you mentioned that the Wave team is continuing to work and experiment on different generative AI applications. Could we get a peek behind the curtain of what's in the pipeline and what you all are focusing on? Yeah, what's coming next? Well, the next steps for Lingo and Gaia are super exciting.

Read the full transcript

26:02Of course, right now they are offline tools that let us understand and explore synthetic and generative data. The results are pretty exciting when you integrate them live on the vehicle. And so integrating them into a way that they can affect and control the way we drive. And of course, there's a ton of scale to go from here. The scale and performance of these algorithms is where we're just getting started. So watch this space. Yeah, so much has happened in the past year in this space. What do you see the next five years looking like? How is this all going to continue to progress and shape up as AV 2.0 continues to develop?

26:44Yeah, well, the exciting thing is I think there's an opportunity to see this technology deployed in the real near future because we're now seeing vehicles being produced that have surround cameras, forward-facing radar, a single GPU like the NVIDIA Orin in these vehicles. And all of those vehicles being produced by the leading manufacturers around the world have the necessary hardware to support a stack like this. Now, of course, when you start to bring together the hardware and software, the things you can and should be improving on top of that. But the fact that it's almost like the automotive industry's roadmaps and AV 2.0 roadmaps are now colliding and it's still very far away from the AV 1.0 roadmap that requires all of those LiDARs, HD maps, sensors and supercomputers in the back of them to run.

27:30But the fact that these systems can be deployed on those vehicles is hugely exciting. This opens up the door for real value to be delivered to people. It opens up the door to start improving safety, to building public trust and expectation of these systems. And of course, the data and the revenue that's going to be required to continue to invest and grow these capabilities. So we're pretty excited about these opportunities. And from our perspective, continuing to build the premier data, whether it's the data we get from our commercial fleet partners or vehicles we deploy in, compute, you know, continuing to scale training with our partners, Microsoft, Azure, or on the algorithm side, continuing to innovate beyond Mingo and Gaia and what's coming next.

28:14These kind of investments are things that will continue to grow. And what I would say is that the rate that we are growing our investment in these areas, the growth from, say, last year to this year, I think we're going to see an even larger growth from this year to the next and continually go on. That kind of trend is just accelerating. And, you know, perhaps the one that's caught me by surprise has been the speed that language has evolved. Large language models have been a huge unlock for bringing together different modes of data and to be able to pull in the intelligence from internet text, as well as create these kind of interfaces.

28:50And I'm sure there's going to be a couple of more breakthroughs of that magnitude. On the AI front, it's been interesting to watch. Like 10 years ago, all of the breakthroughs came from computer vision, whether it was AlexNet, residual connections, batch normalization, all of the big deep learning breakthroughs were in vision. Then we saw the advent of language and we've seen transformers and other technologies come there. And so it's been interesting to see how the problem domain has shifted. But I think embodied AI over the next five, maybe 10 years is going to start to become the premier way because it pushes you and challenges you on things like safety, criticality, trust, transparency.

29:27It pushes you on dynamic environments that are never the same train and test set, but they require you to generalize to new environments. And I think the richness of the problems in the space, for all the leading AI researchers out there, I would really encourage you to look closely at embodied AI because I think that's where those next big breakthroughs are going to come. It's where the problems lie. It's where the challenges that push data, that aren't just exploiting what larger language models can do today or brute forcing more scale, but the problems that push you to think how to build what's next, build better, build further.

30:00I think that's all embodied AI. And so it's going to be exciting for us to have this conversation in five years. I think we will have AIs that can drive vehicles in a way we can trust, interact with them with language, prompt them, ask them to do certain things. or even delegate parts of driving to them in a way where they can take control. This is a pretty cool future. It sounds like we've only really scratched the surface of this revolutionary technology. To learn more about Wave, please be sure to register for GTC, happening in person and virtually March 2024. Alex will be hosting an in-depth session on the impact of generative AI on AV development.

30:43Thank you so much, Alex, for joining me today. This was a really fascinating and enlightening conversation. Thanks, Katie. I can't wait to come to GTC next year. Cheers for the conversation.

31:44¶¶

From the publisher

A new era of autonomous vehicle technology, known as AV 2.0, has emerged, marked by large, unified AI models that can control multiple parts of the vehicle stack, from perception and planning to control.

Wayve, a London-based autonomous driving technology company, and a member of NVIDIA's startup accelerator program, is leading the surf.

In the latest episode of NVIDIA’s AI Podcast, host Katie Burke Washabaugh spoke with the company’s cofounder and CEO, Alex Kendall, about what AV 2.0 means for the future of self-driving cars.

Unlike AV 1.0’s focus on perfecting a vehicle’s perception capabilities using multiple deep neural networks, AV 2.0 calls for comprehensive in-vehicle intelligence to drive decision-making in real-world, dynamic environments.

Embodied AI — the concept of giving AI a physical interface to interact with the world — is the basis of this new AV wave.

Kendall pointed out that it’s a “hardware/software problem — you need to consider these things separately,” even as they work together. For example, a vehicle can have the highest-quality sensors, but without the right software, the system can’t use them to execute the right decisions.

Generative AI plays a key role, enabling synthetic data generation so AV makers can use a model’s previous experiences to create and simulate novel driving scenarios.

It can “take crowds of pedestrians and snow and bring them together” to “create a snowy, crowded pedestrian scene” that the vehicle has never experienced before.

According to Kendall, that will “play a huge role in both learning and validating the level of performance that we need to deploy these vehicles safely” — all while saving time and costs.

In June, Wayve unveiled GAIA-1, a generative world model for developing autonomous vehicles.

The company also recently announced LINGO-1, an AI model that allows passengers to use natural language to enhance the learning and explainability of AI driving models.

Looking ahead, the company hopes to scale and further develop its solutions, improving the safety of AVs to deliver value, build public trust and meet customer expectations.

Kendall views embodied AI as playing a definitive role in the future of the AI landscape, pushing pioneers to “build better” and “build further” to achieve the “next big breakthroughs.”

For more on NVIDIA's Inception startup accelerator program, visit https://www.nvidia.com/en-us/startups/

More from NVIDIA AI Podcast

All 115 episodes
Wayve CEO Alex Kendall on Making a Splash in Autonomous Vehicles - Ep. 209NVIDIA AI Podcast · 32 min
Listen in VO