In short
Eye On A.I. Podcast Episode Summary
Podcast Information
- Title: Eye On A.I.
- Host: Craig S. Smith
- Description: A biweekly podcast focusing on individuals making significant contributions in artificial intelligence, providing context and considering global implications.
Episode Details
- Episode Title: #152 Alex Kendall: How Close Is AI to Taking the Wheel?
- Description: Craig Smith talks with Alex Kendall, CEO of Wavye, about advancements in autonomous vehicle technology using AI, particularly focusing on the company's world model approach, called Gaia One.
Key Topics Discussed
- Introduction to the Guest
- Alex Kendall is the founder and CEO of Wavye, a company focused on developing autonomous vehicle technology leveraging AI.
- Background in engineering and a PhD in computer vision from the University of Cambridge.
- World Models in AI
- Definition: A world model helps an AI understand the state of the world and predict changes based on actions taken.
- Importance: Essential for safety-critical applications like autonomous driving to ensure accurate decision-making.
- Contrasts with Large Language Models (LLMs): World models can simulate future scenarios, which is crucial for understanding the implications of decisions.
- Autonomous Driving Technology
- Wavye employs an AI-driven approach that is distinct from traditional robotics methods often seen in current autonomous systems, which rely heavily on predefined rules and infrastructure (like HD maps).
- End-to-End Neural Network: Wavye utilizes a single neural network that processes inputs to produce motion plans.
- Challenges and Innovations
- Data Challenges: Discussed the high-dimensional data from sensors and the need for efficient data processing.
- Generalization in Driving: The system is designed to generalize across different environments without needing extensive retraining.
- Unsupervised Learning: Wavye's approach minimizes the need for labeled data, improving efficiency.
- AI Model Architecture
- Gaia One's Structure: The model is built on a transformer that predicts future states based on tokenized inputs from various modalities (video, text, actions).
- Integration with Robotics: The technology is expected to adapt to other forms of robotics beyond just vehicles.
- Future of Wavye and Regulations
- Commercialization Efforts: Wavye is actively partnering with major UK fleets for grocery delivery and other applications.
- Regulatory Engagement: The company is working with UK regulators to develop a framework for autonomous driving.
- Compute Requirements
- Emphasized the difference in compute requirements for training and inference, with the majority of inference running on vehicles rather than in the cloud.
- Partnership with Microsoft Azure to handle extensive data processing demands.
- Future Directions and Open Source
- While the research paper on Gaia has been published, the actual model is not open-sourced yet.
- Future aspirations include expanding the technology to humanoid robotics.
Key Takeaways
- World models are transforming how AI can be applied to real-world scenarios like autonomous driving.
- Wavye is pushing the boundaries of AI technology with a focus on efficiency, safety, and generalization.
- Collaboration with regulators and industry partners is crucial for navigating the complex landscape of autonomous vehicle deployment.
Conclusion This episode highlights the innovative approaches taken by Alex Kendall and Wavye in the development of AI-driven autonomous vehicles. The conversation sheds light on the importance of world models in enhancing decision-making capabilities and the ongoing challenges in the industry.
For further insights, consider exploring Wavye's progress and the implications of their technology on future AI applications.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00The world model for me is a model that can understand the state of the world and predict how it's going to change given an action you put into it. So mathematically speaking, it's a function that takes your current state, your current action and predicts the next state. Essentially a simulator. It's a model that can allow you to understand how the world will evolve given different things you might want to do or interact with that world. The result of this is a world model that can simulate the future. And if you want to, you can take that state and decode it back to video so you can produce the video output of what's actually going to happen.
0:36But you might not want to if you want to keep this real time and efficient. You can just stay in your embedding space and use it to drive your car. Hi, I wanted to jump in and give a shout out to our sponsor NetSuite by Oracle. I'm a journalist and getting a single source of truth is nearly impossible. If you're a business owner, having a single source of truth is critical to running your operations. If this is you, you should know these three numbers. 36 ,000, 25, 1. 36 ,000 because that's the number of businesses that have upgraded to NetSuite by Oracle. NetSuite is the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.
1:2825 because NetSuite turns 25 this year. That's 25 years of helping businesses do more with less, close their books in days, not weeks, and drive down costs. One, because your business is one of a kind. So you get a customized solution for all of your KPIs in one efficient system with one source of truth. Manage risk, get reliable forecasts, and improve margins. Everything you need, all in one place. As I said, I'm not the most organized person in the world, and there's real power to having all of the information in one place to make better decisions. This is an unprecedented offer by NetSuite to make that possible.
2:17Right now, download NetSuite's popular KPI checklist designed to give you consistently excellent performance, absolutely free at netsuite.com slash ionai. That's I-ON-A-I-E-Y-E-O-N-A-I, all run together. Go to netsuite.com slash ionai to get your own KPI checklist. They support us, so let's support them. I'm Craig Smith, and this is Eye on AI. This week, I speak with Alex Kendall, CEO of Wave AI, to understand Wave's innovative approach to autonomous vehicles using a world model called Gaia One. Alex explains the advantages of world models, which we've explored before on this podcast, with Jan LeCun and how they can be used in AI agents.
3:21The discussion offers a unique view on the progress, promise, and obstacles in developing AI to act in the physical world. I hope you find the conversation as fascinating as I did. I'm an engineer at heart. I've loved building things ever since I was growing up and did so in the back garden. But I got the chance to work on a bunch of robotics in my childhood in New Zealand, whether it's building drones to chase some sheep around the field that I grew up in or other projects I did at university. But one way or another, they ended up taking me to the University of Cambridge. I spent many years there doing a PhD in research fellowship in computer vision.
4:06I'm fortunate enough to publish some of the first work that applied deep learning to scene understanding algorithms like semantic segmentation, depth motion, other forms of scene understanding. And yeah, I think that work really inspired some of the ideas that I've had to be able to build machines that can make decisions for themselves. Interestingly enough, a lot of my PhD work stopped short of understanding the future or doing future prediction, which is, I think, one of the big topics we've been able to address with world models and Gaia, which I'm looking forward to talking about. Yeah, well, that's fascinating.
4:48And as I said, I just had Jan LeCun on the podcast talking about world models. He mentioned Gaia. and so where do I start I'm interested in in world models as an alternative to large language models and Gaia won your model uh Jan says it's a little different uh than his JEPA architecture that he's using to research world models. So can you tell us a little bit about Gaia One? And then I have a lot of questions. I'm interested in marrying this tech with robotics, because the big challenge in robotics beyond the hardware is building an AI brain that can plan and make decisions and that sort of thing.
5:53And there's a lot of talk right now about LLMs being able to play that role, but I also think there are a lot of problems because of the hallucinations or the fact that large language models don't have a very concrete underlying model of the world. So why don't you talk about Gaia, how that came about, what it is,
6:27both generally and then how it's built, and we'll go from there. Well, taking a step back, maybe some background. I lead Wave, an autonomous driving company, and how we've set off on a different path to build autonomous driving systems that have the onboard intelligence to drive different vehicles in new places, including places they haven't been to before, and understand the complexity and the long tail of situations that you see on our roads. And taking an AI approach to autonomous is driving is quite contrarian and different to how people usually look at this problem. When we started six years ago in 2017, we set off to build an end-to-end neural net that could learn to drive, take the data as input and output a motion plan to control a vehicle.
7:21And this end-to-end AI approach, I guess, why are world models interesting here? So a world model for me is a model that can understand the state of the world and predict how it's going to change given an action that you you know you put into it so mathematically speaking it's a function that takes your current state your current action and predicts the next state so it's a it's a essentially a simulator it's a it's a model that can allow you to understand how the world will evolve given different things you might want to do or interact with that world why is this important for self-driving? Well, the first thing you might look at when you're building an end-to-end neural network to drive a car is to build something that's autoregressive.
8:03Build something that creates a function that takes your input state and produces the motion plan that you should drive with. And that's kind of what people did in large language models. And the problem with self-driving is it's a safety critical application. If you make the wrong decision, you're not just going to put out some hallucinated text, but it's life and death decisions of driving on our roads. So for that reason, it's really important that you are aware of the implication of your decision and you can understand the dynamics of the world. So that's really what motivated us to start off with world models as a concept.
8:35And in 2018, we actually published a blog with one of the first examples of doing this. On an autonomous vehicle, we published a model-based reinforcement learning system where, actually only on a quiet country road, but we learned to drive a car with a world model. So this system had never driven an actual car, it only learned in its imagination in a world model, but it used this model that it trained of the dynamics of the world to be able to learn to operate this car and drive it down a quiet country road. And I guess over the last six years, I can talk more about Gaia, but over the last six years we have scaled that approach to the point it is today where with the latest and greatest in generative AI, we can now understand the full, diverse, rich, dynamic urban scenes that we operate in today, like central London.
9:22And building the world model, I mean, the autonomous driving systems to date are, help me out here, but they're primarily reinforcement learning systems that are taking in data from various sensors and have trained on what the best approach is in any particular situation. Is that right? Or what is the current state of autonomous driving systems? Well, we're having an AI conversation and fundamentally autonomous driving is an AI problem. It's a problem of complex, high-dimensional decision-making. And so So you'd assume that you're going to use a data-driven method like reinforcement to do it. But actually, that's not the case.
10:15If you look at all of the large autonomous driving, if it's out there today outside of WAVE, you know, primarily the approach is a traditional robotics one. It's one of, yes, you use deep learning for perception, but once you have the state of the world, it's very much a, you know, hand coded optimization approach to produce a motion plan that's aided by a set of infrastructure like HD map, high definition maps that tell the car where and how to behave. So it's actually not an AI approach. And what we've done is, you know, I think the first time an AI system has actually driven on the roads at this level of scale.
10:51So that's not how the industry has traditionally approached things. When you bring in an AI approach, of course, you know, the challenges there are how do you understand what it's doing? How do you make sure it's safe and making the right decisions? And so that's brought up some of these challenges that led us down the road of world models. So, well, that's interesting. I didn't know that about autonomous driving systems. So they have all these sensors that data is coming into a central decision maker. And you're saying that decision maker is a traditional control system and not probabilistic AI system?
11:33Broadly speaking, yes. I mean, a lot of the systems running around San Francisco today, for example, are of that approach. Now, more and more machine learning is being used throughout over year on year in these systems. But it's not an end-to-end neural net. It's not a large transformer that decides the whole decision making. And that's the step that we've taken to replace that entire stack with one big neural network that learns how to drive end-to-end. Right. And I've seen a lot written recently about using the reasoning power of large language models to play that role, to decide on actions.
12:17And then, you know, with some other piece of software to translate that action, to execute on that plan. And can you talk about, well, first of all, for Gaia, so there's a world model that's building a state of the world in its, I guess, its weights. and then a reinforcement learning model that learns to act on that state of the world. Is that right? Maybe you can describe the architecture a little bit. Yeah. Let me jump into some of those details. We've got a research paper online that talks about them in great depth. But one of the interesting things for me is that, you know, if you look at the three big major trends that we've seen in large language models this year, I mean, at the start of the year, it was all about scale.
13:23Everyone was talking about how many parameters, how much data, how much compute are these models trained on. In the middle of the year, it became about multimodality. We pushed scale to some degree, and now it's about how do we understand across different modes. and a lot of image text systems came out, for example. And then more recently, it's about synthetic data. The benefits of synthetic data are clear. You can control the bias in your training data. You can ensure that the training data is equally sampled across the things you care about. You can control the distribution of your training data.
14:01And you can often get information that is harder to understand and from noisy real data alone. And so I think those three trends that have really driven the state of the art and say large language models, the interesting thing is we've seen the exact same thing play out in robotics. So for us at WAVE, we've been pushing the scale of our neural network that drives the car. And in the next year, our roadmap is going to be pushing this in terms of parameters, data and compute by 100 X further, two orders of magnitude further. And so the results, the emergent behavior we're seeing come out is just remarkable.
14:38The ability for the car to nudge its way through crowds of pedestrians, to do complicated unprotected turns, to predict the behavior of other agents cutting in or moving around our vehicle, all of this kind of thing emerges at that level of scale. The second trend on multi-modality, you know, that's where I think it's really important to be able to learn to understand between different modes because ultimately if you're training a self-driving car just off the video data it has, it's going to be intelligent but you know perhaps it'd be more intelligent if it doesn't only have that video data but also you know text and other information sources it has when you and i learned to drive um i learned to drive when i was 16 uh and you know i had what maybe my my mom and dad probably had had 20 or 30 hours in the car with me uh maybe not that long maybe i maybe i learned something like five or ten hours i think i hope i was a fast learner but you know that kind of length and time to learn how to drive um but But it wasn't just that that allowed me to drive.
15:36It was probably the 15 or 16 years of experience I'd had of learning how the world works, learning what objects are, how things might behave on the roads. And it was that observation that gave me the intelligence to drive. I think that's seen the same as true in robotics. We can train our autonomous driving system now, not just on the video data of it driving, but also internet, video and scale video and text. You know, we can literally feed it the PDF document of the highway road code that the government writes and give it that as, you know, further context to understand. And so I think multimodality is becoming really important to bring together different sources of information and improve the intelligence of your system.
16:14But then secondly, with tech specifically, you know, I think the future of how we interact with robots is going to be through language. We are going to be talking to our robots, interacting with them. you know there's a reason why you and i have evolved language as a way we communicate is because it's the most efficient way to get information across that that you know that we understand uh um and uh and so for that reason or maybe not most efficient but most natural way and so for that reason i think the accessibility of robotics will be greatly improved by us being able to just literally converse with it you should be able to be in your self-driving car and say take the next left take the next right drop me off here or i'm worried about this why are you doing that and you should be able to you know build a sense of trust through it and we've done exactly that at wave we've produced a system called lingo which is a first a vision language action foundation model that combines those modalities of video of action and robotics and language that allows us to talk to our autonomous vehicle and ask it why what it's doing and then finally uh the third trend on synthetic data and this is where gaia comes in uh gaia our world model um not only as a system that allows us AI to understand the implications of decisions it's making, but also produce synthetic data.
17:28Generative AI is very good at recombining data in new ways. And so we have lots of experiences of foggy scenes on the car. We have lots of experiences of jaywalking scenarios, but we have very few foggy jaywalking scenarios. And Gaia can not only allow us to understand how the world's going to evolve, but it can create new examples we haven't seen before. And again, we can do that by connecting vision, language and action. We can prompt it and say, literally give it a prompt and say, I want an example of a jaywalking pedestrian in the fog. Or we can take a scene that exists in the real world and change it and ask it to recreate it with new features and things like that.
18:07I didn't really get into the architecture guy and I'm happy to, but I just thought those three trends have been really powerful for us in ai and uh and you know uh exactly applicable to robotics as well yeah uh so you can tell from my questioning i'm a journalist uh not a practitioner uh so my knowledge is fairly surface but i understand large language models uh to a degree i understand the transformer algorithm and and And the large language models are predicting the next token in a sequence. And because of the volume of training data, it does a very credible job most of the time. A world model, I understand Yen's JEPA architecture, the joint embedding predictive architecture, to a degree.
19:08he says your model is is something different so can you kind of walk us through at a very high level of uh how uh the the model is trained and and what it's doing it's predicting the the the next state of the world whether it's video or text or whatever is it doing that with uh a transformer algorithm uh how is it doing it just i'm sure the the audience would also like to hear absolutely uh so i mean yan and i share a lot of common belief around these systems we think that um to go beyond the autoregressive nature of language models are just predicting the next word and getting to systems that can understand and be safety critical we need to have world models um we share the vision these should be unsupervised they should be able to be trained through self-supervision uh through you know whether it's signals like contrastive learning or building energy spaces or things like this um you know i think we share a lot in common there uh um i mean jepper is is a great architectural approach uh i you know i think um uh i think there's a lot in common these systems in terms of there's some representation space um you want to train it with with unsupervised learning and and so for our approach in particular with with Gaia what we first do is we take um you know we tokenize up the different inputs whether it's images action or language and uh essentially it's a well today it's a large transformer but you could use whatever your favorite flavor of neural network or let's say machine learning system that you want to use.
20:58But essentially you take those inputs and you learn this dynamics model, this ability to take current state, current action and predict next state. And then the result of this is a world model that can simulate the future. And if you want to, you can take that state and decode it back to video so you can produce the video output of what's actually going to happen. But you might not want to if you want to keep this real time and efficient. You can just stay in your embedding space and use it to drive your car. And, you know, right now we're working on Gaia 2 and Gaia Drive. And these are systems that very much are going to see Gaia embedded in the vehicle, able to actually increase the intelligence, understanding, and improve the safety criticality of our system in a production setting.
21:45So that's really the guts of the architecture. Today, it's a transformer that's able to predict future states. And again, forgive me. I'm sure you'll cringe at my repeating this back to you in sort of super layman's language. But you're encoding the data coming in or tokenizing it and turning it into embeddings in, I guess you call it a feature space or embedding dimension. And you're making predictions based on that at that level. So they're, and then to see the video, you decode it into pixels. But in that space, you can predict the, is it, how specific are those representations? When I said embedding space, I guess I meant representation space, yeah.
23:01One of the other big challenges of self-driving compared to large language models is the amount of data you have as input is enormous. I mean, take our latest vehicle, for example, it's six or seven cameras. They have eight megapixels each, and then you care about multiple frames over time, and not to mention if you consider an imaging radar as well, or choose the sensors you want to use. Just in cameras alone, the data, eight megapixels, your RGB value is there, so So, you know, 24 bytes, 24 million bytes alone from those, multiply that by the six cameras, there you've got about 120 million bytes, and then multiply that by multiple timeframes.
23:44I mean, you're talking about gigabytes of data there alone. And that, you know, that data, what it does is it makes it, you can't have an embedding space where you're dealing with gigabytes of data. It's just too much to process. And so to make these models practical, you need to be able to take that extraordinary amount of input data, all those videos, and compress them into an embedding space that you can reason about. So I guess the question is, how many factors do you think you care about for a driving scene? You know, you can list them out. You care about the positions of the cars in front of you, the pedestrians, the cyclists, the direction they're facing, the way they're going to move, the weather conditions, the traffic light, all these kind of factors.
24:27Now, interestingly, if you go down the path of trying to list them out by hand, you end up in the AV 1.0 or the classical robotics approach to autonomy. And that doesn't scale because it's very hard to enumerate all those factors and to reason about them a priori. So we can do it as a thought exercise, but I wouldn't advocate for that approach. But the point is, is that it's not, you know, hundreds of millions of numbers. There's probably a much smaller set of things that you care about there. And so we want to learn that. you know you don't care about the pixels in the sky except you want to know the weather but you don't really care about all those things it's there's a lot of redundant information in in the signal compare that to large language models you have sentences of text that are really high signal to noise ratio you know text is precise at saying this means that and impacts this right it's a direct description of what you care about compared to videos where most of the pixels in an image you don't care about clouds you know second story of a building you don't care about when you're driving So the point is that you want to be able to take this data and embed it in a very efficient space.
25:31And the way we do that is through end-to-end learning about, you know, what do we care about for driving? What actually is going to impact how the world's going to evolve? And that's what we look to learn. So we look to build a transformer and a self-supervised learning approach that learns an embedding that is really efficient, is as small and compressed as possible, but has the information that we need to understand the safety critical natures of the scenarios that we're driving through. So that's the primary task of that embedding model and that learning model of the scene representation. Yeah.
26:07And then right now that, you know, I've seen some remarkable videos that you've done and maybe I can, you know, splice one into this podcast. But they're awesome when you go drive in the car. I go out most weeks and when you see it, learn new behaviors week on week. Like just yesterday, I went for a drive down in part of London in Notting Hill and we went through Portobello Road Market. It's this crazy area with loads of pedestrians all on the road. the fact I've never seen our car sort of nudge its way through a crowd of pedestrians before, but it did so really safely in a way that, you know, if you just stopped and waited for the pedestrians to clear, you'd be stuck there for an hour.
26:53So these kind of new behaviors over time, there's tons of videos that we have online of this stuff, but it's pretty amazing seeing AI operate in the physical world like this. Yeah. The videos, so you're then decoding the representation space into pixels in a video, that's useful, I presume, for creating training data. But when you're actually driving in real time, how is that state of the world being translated into action and that's where i think you said there's a an rl engine or agent that's that's learning over time how to act on that can you talk about that part yeah absolutely um i mean one of the amazing things about gaia properties are that they can create photorealistic and diverse scenes that controllable by text you can modify these environments and you can also create multimodal futures i think the really powerful thing is when you only observe the past you know you can't predict how the future is going to unfold in a driving scene and the fact that gaia can generate diverse multimodal plausible futures is a really important a really important factor but uh yeah in order to actually control the car so this is what we call Gaia Drive.
28:35It's not just decoding to images, but also, well, there's many ways that you can go about thinking about incorporating Gaia into a driving system. For example, you could generate future data and use that synthetic data to actually just train a system. You could use it to predict the future and use the information it learns about predicting the future to improve the driving representation. Or you could actually bring it into a full on, you know, model based reinforcement learning, a model predictive control or some kind of learn simulator. What that means is that let's say you're in a driving, you're at a green light and you want to decide whether you drive through the intersection.
29:16What the system can do is it can perceive, run its world model, run Gaia for a few seconds ahead and see what might happen. Maybe it runs it a few times and sees how various different things happen. and then it can make a decision based on how it thinks the future is going to play out we do that in in our brains and in our hippocampus we have mechanisms that are you know most famously referred to as sort of thinking fast and thinking slow thinking fast the reactive decision making you don't really plan ahead you just you just do whereas thinking slow you take a step back and use a reason what might happen should i do this and you sort of play chess a few a few steps forward and we can do the same thing in robotics robotics has typically had you know two levels of control there's like a low level control that runs it over 100 times a second that controls parts of a robot and you have a high level control that operates typically around 10 times a second that does the high level reasoning but actually what these kind of models that you do is maybe move one more level up abstract and have a three-tiered system you can have a thinking slow thing which can involve a large language model to interact and to reason and to plan it could involve a world model to actually understand the implications of these decisions and that might happen at anything from one hertz you know one time a second to maybe even one time every 10 seconds it's quite an infrequent high level if you think about when you're driving that sort of high level kind of topological task planning you do can happen at the highest level then that middle tier you know you're deciding you're designing the motion plan that you want follow to ensure you don't hit things and that you follow road rules that's more reactive and it runs it at that kind of 10 times a second player and then 100 times a second is sort of the minute changes and breaks and and steering to make sure that you actually achieve that plan that you've set out and that three-tier uh system i think is uh uh so i think world models can play a really great role in that top one a new level of higher level reasoning into into robotics and uh the advantages of that over uh agents built purely off of large language models is that they're they're they require a less compute i would imagine but also um their their their decisions or their predictions are grounded more firmly in reality where as a large language model even if it's tokenizing pixels in a video it's it's only predicting one token ahead as it goes along is that right yeah look i think large language models and world models can be complementary right like large language models give you the ability to under well give you a text interface they give you a ability to interact through language but they give you the ability to learn a really incredible understanding of the world through internet scale text.
32:14And world models, the advantages of world models is they give you the ability to understand the implication of your decision, whether you are making a driving task decision or whether you are outputting sentences and text, you know, whatever it may be, it allows you to understand, okay, what is the implication of that decision in the environment I'm operating in? Is that a good thing or not and that understanding can help you do things in a safety critical environment so i think world models become really important in an application like self-driving maybe not so much in like an internet search problem but when you have safety critical applications whether that's in you know medical correspondence of language models or self-driving for for our application that's when uh world models give the ability to do that the other advantage i'll describe those is not just at runtime, but also at training time.
33:03World models give you ability to learn more efficiently. We do this in our brains as well. As people, we daydream or nightdream. When we dream, we actually go through a process that solidifies our actions and lets us replay experiences to learn to do them better next time. Yeah, like if you're learning to play tennis and you you hit a ball once you know you're not going to hit a ball with every single permutation of the angle that your racket might be to learn how it might go you might only hit it a few times but from that you need to learn the general way of how to hit a ball to get it in the right part of the court to be able to play tennis and what we do is from the ways that we hit the ball we actually um you know replay this in our internal world model many times when we daydream to be able to learn how to do that more effectively and the same is true with machine learning when you have a world model you can get more out of your training data you can replay recombine reconfigure and use that to understand your training data and and learn a lot more performant effective or safe policy or decision making system from your training data because of the fact that you can learn and recombine those experiences in new ways.
34:18So world models are really powerful, both training and inference or testing time. So, and you have this system not only creating training data, but acting as an agent to drive a car. How developed is that system? how far is it from commercialization for example how many cars do you have it in and how many road miles have you logged and that sort of thing yeah we've been we've been spending the last uh six years building a different approach to autonomy and where we're at today is we've been able to demonstrate that it can do a lot of things that have been blocking the industry for many years. It can drive on the kind of equipment that's on modern vehicles today, a single GPU computer, surround cameras, maybe a forward-facing radar.
35:14It can drive in different places it's never been to before, and it can drive different vehicle types. So we are very excited to be in the process of commercializing this technology now, to seeing it deployed across the world's most innovative to fleets and vehicle manufacturers, and to see this deployed in a way that can realize value quickly and accelerate the growth from assisted autonomy through to full autonomy. And so we are in that process right now. So far, we've been excited to partner with some of the UK's largest fleets, fleets like DPD, Asda, and Ocado Group. And these are large fleets that do things like grocery delivery here in the UK.
36:01And those partnerships have been wonderful. We've been delivering groceries throughout this year with our partners as, for example, in London, it's showing some of the value that this autonomy technology can bring to society. Yeah, just on that, I imagine you still have a safety driver in the car. How is the regulation? I don't want to get lost in that regulation discussion, but how are you managing to operate? As you know, crews in California has run into all kinds of trouble, but how is the regulatory environment allowing you to operate fleets, for example? Yeah, that's a really important question.
Read the full transcript
36:51So today we operate with safety operators, but we have a two-pronged approach here. The first is that we want to see the ability to build value and see deployment of the system still as a driver assistance system or with safety operators, as you say. And there's extraordinary value that can be brought there, whether it's the helping support safety or improving the efficiency of operating vehicles. There's a big opportunity there that we can see today through driver assistance. But then on the other hand, of course, ultimately we want to get to level four autonomous driving at scale. And we've been really excited about the work that we've done with regulators around the world, but most specifically here in the UK.
37:35We've had a number of UK ministers for rides to show them the technology firsthand. We've sponsored the UK's parliamentary working group on autonomous driving technology and we've helped support bringing legislation to be considered to make this technology legal and we've been offered and are working on a 1.9 million pound grant with the government to help understand and put forward a safety framework for AI systems. There's many more activities that we've been doing in the space but we're deeply engaged with regulators and believe that empowering regulators to understand AI technology and to, yeah, empowering them to understand it and to be thoughtful about how best to manage those risks is the best way forward.
38:27So we've really looked to do our part on that front and have been excited about the results so far, the traction, I should say. Yeah. Can you, I have questions about the compute intensity of this and the the the amount of training data but setting that aside for a minute uh mercedes-benz has i think a level four autonomy in germany if i'm not mistaken and the regulations there allow them to drive hands-free on certain roads and that sort of thing. How is that system working, and how does that compare to Gaia 1? And, yeah, answer that first. So my understanding is that I haven't seen a level four system from Mercedes-Benz or other automotive manufacturers outside of the trials we've seen in China or in some parts of the US with some of the technology companies.
39:48I have seen a limited level three system which is where the vehicle does take control and liability of the vehicle for certain scenarios and I believe Mercedes have a product that allow you to do this at low speeds on a highway in in you know rush hour traffic but you know in general the automotive industry is very we're yet to see as an automotive industry we're yet to see technology at scale which can do productionized which can do driving you know any vehicle anywhere whether it's not just highway at low speeds but urban suburban uh different cities around the world and that's very much what we're interested in solving a way we want to build a drive an ai driver an artificial intelligence system that that has the on-board intelligence to drive vehicles through all of the kind of scenarios that you know we expect in our daily lives and the commutes and the travel that we do we want to build this kind of technology to help assist people to free up their time, to make it safer on the roads, to give them a more sustainable drive.
40:59All of these kind of benefits, we want to ensure this technology can be brought out quickly and broadly. We think that this is a unique opportunity in this space and something that, of course, the world... But, yeah, I think it's a really important next evolution that we need to be able to deliver to lift up, quite frankly, the safety and performance of all of the cars that we're putting on the road today. Yeah. And for training this system, I mean, one of the challenges, as I understand it, of other autonomous vehicle control systems is they depend very heavily on supervised learning. and you have to label enormous amounts of data so that the system recognizes corner cases and there's this very long tail of those cases.
42:06And for example, you can train a car to drive in California, But if you take it to Norway or something with a vastly different climate, it's going to run into trouble if the snow hasn't been labeled into the system and that sort of thing. how much how do you train Gaia and can it generalize across environments in real time or does it do you need to train it in advance for every kind of environment the interesting thing is that when you are driving you're generalizing all the time you will never see the same thing twice on the road. Every time you go driving, the weather's going to be slightly different than you've ever seen before.
43:10You're going to have cars and other agents around you in different locations. So even if you're on the same road that you've driven, you're commuted every day of your life, the specific things that you go through will be different in some way from what you've seen before. So I guess the first point to make is that autonomous driving is all about being able to generalize. We do it every time we operate these vehicles on the road. The question is, how far can you generalize? It's one thing to generalize driving on the same road every day. It's another to drive on one road in the UK and then be able to drive in the US where you have new things like four-way stop signs, right turn at red, other local driving cultures.
43:49And so we want to build a system that can effortlessly generalize to new environments, to allow us to bring it to everyone around the world. And how we've done that is our system is trained through unsupervised learning. So it doesn't need boxes drawn around objects or things labeled in the scene for us to learn how to understand the scene. It's all unsupervised. We watch driving data and learn from that how the world works, learn how to predict the future, train models like Gaia and these kind of things. So it's all through unsupervised learning. And the key thing is that it becomes more efficient the more data you get.
44:25So we might need a certain amount of data to train in the UK for us to train a system and and for us to generalize a system to another environment like the US, we will need a fraction of that data because a lot of the experience is shared, right? The same rules of physics apply in both countries. People tend to behave in a similar way, although there are some differences. You need a little bit more data. We found that going from a passenger vehicle to a van, for example, generalizing to a different vehicle, going from a Jaguar I-PACE passenger vehicle to a 3.5 ton delivery van, that took us about 2 to 3 % of the training data to be able to achieve similar level of performance on the van.
45:08And so generalization with end-to-end neural networks can be very efficient. It's like large language models learning to generalize from English to German to French to Mandarin to other languages. And so those are some of the things that we think about when we look to scale our technology and generalize it to new driving scenarios. and and what about the uh the compute uh required uh to to train the models and then to uh operate the models and do the models uh operate uh through a connection to the cloud or are they is there a GPU running in the vehicle? Yeah, so we don't train the models live on the car.
46:01They are safety assured, validated before they're deployed on the car. Rather, we take the experience we get driving all the time and we feed that back to the cloud in order to improve the new models that we'll be validating and deploying. So it's all trained at scale, and we've been partnered for many years with our great friends Microsoft and we work with Azure to be able to train at scale. Azure has been able to provide us with extraordinary compute power, but more importantly, the innovation that you need to be able to deploy to train these kind of models. The tough thing with training video scale foundation models like we do is that the data requirements, I talked about the size of the input state of these models, it's truly genormous.
46:48and to train on this kind of data, you can't have it all sitting on the local compute nodes. It's tens of petabytes, maybe hundreds of petabytes, and so you need to be able to stream it to your GPUs. This means you need a different set of infrastructure from when you're training a large language model. You can't have your data stored locally on the GPUs. You need to be able to stream it from storage that you can do fairly random access across, and that's quite a hard infrastructure challenge. So we've been thrilled to be able to do a lot of pioneering work in the space with Microsoft to make it possible to train models at this scale.
47:25And is that in order to get the video into a... Do you vectorize it in order to tokenize it? Or how does that work? Yeah, we take the video data or the experience for all kinds of data. And yep, it's tokenized. It's fed into the transformers at scale. And that's how it's trained. And it's not just video data. It's also the navigation prompt. It's the other robot sensors like GPS, you know wheel speeds things like this all these kind of input data we you know we we use and and feed into the AI. Yeah I mean the reason I'm asking about the compute requirements is you know these large language models are amazing they're people that are scaling them 10 times or more from what we've seen so far but there's a limit in the availability of gpus for the time being and for the foreseeable not the foreseeable future but for for at least a few years uh and then uh because of the that constraint the there's a limit on uh the amount of inference anyone customer can use the model for, you know, these rate limits.
49:08Does a world model face those same constraints? Sorry, Craig. Can you clarify your question? Constraints around rate limits? Yeah, just on the availability of GPUs, both for training and for inference. Because, you know, it costs an enormous amount of money to train an LLM, but it also costs an enormous amount of money on the inference side. And people are accessing these models through APIs, and because of the cost and the compute constraints, they can only process so many requests a minute or so many tokens a minute. And that is really limiting the enterprise scale applications. So does that same kind of problem run through a world model or is a world model less, fundamentally less compute intensive and so you can get around those problems?
50:17Well, the great thing about our space is that at inference time, most of the compute runs on the car. And so assuming you've got vehicle assets to deploy on that are there and fleets and are generating value, then you're in a great spot because you can run all the inference you need on the vehicles themselves. So our challenge is really a training time. Yes, there's some inference costs, but it's more manageable because it's a hybrid edge cloud model. and primarily you need to have most of it on the car because the car should be able to have all the intelligence it needs to be safe and make the kind of decisions it needs to operate in an environment on board the vehicle.
51:01So training time, it really matters. And there, yeah, I mean, we are hungry for all of the compute and data storage that we can get to be able to power these models. And I feel fortunate to be writing Moore's Law and other year-on-year improvement curves that just bring down the costs and increase the availability of increasing orders and magnitudes of these systems. But no, we certainly do have appetite to lift what we've got another 100x to next year in terms of scale of data and compute and parameters in this model. Training computers are really, really important factor for us. Okay, then quickly, two questions.
51:43Is Gaia open source And I've been writing a lot about why we don't see humanoid robots the way everyone wants to see. And one of the hardware problems, but the other problem is having a reliable AI brain. could this model be applied to other forms of embodied AI robots? That's an awesome question. I'm really bullish on humanoid robotics being a part of the future. The technology we're building, we want to see this kind of embodied AI system empower all kinds of robots, whether it's manufacturing, domestic robots, or self-driving cars. But I agree with you. I think self-driving cars will be the first big application of embodied AI at scale because there is data, there are hardware platforms, there is a business case and so it can be built today.
52:42I think getting the data and the hardware platforms for humanoid robotics is harder, but I hope that the scale of embodied AI we can build through self-driving can make that easier by taking that technology and adapting it to humanoid robotics in the future. And I would love for Wave to be part of that once we've got our self-driving embodied AI systems to scale. That's really the path we see. So I agree with you on that one, but I'm very bullish on, in 10 years' time, AI not just being chatbots and co-pilots, but being all kinds of physical embodied AI systems in the worlds that we live in. Okay.
53:26And then the question on open source, is Gaia open source? So, Gaia, we have written an extensive research paper around it, and we're very much fans of openly engaging with the AI community and sharing ideas. We haven't open sourced the model itself for now. And that's something that we're continuing to iterate on and develop internally and to use to deploy in our fleets with our partners. Hi, I wanted to jump in and give a shout out to our sponsor, NetSuite by Oracle. I'm a journalist and getting a single source of truth is nearly impossible. If you're a business owner, having a single source of truth is critical to running your operations.
54:12If this is you, you should know these three numbers. 36 ,000, 25, 1. 36 ,000 because that's the number of businesses that have upgraded to NetSuite by Oracle. NetSuite is the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more. 25 because NetSuite turns 25 this year. That's 25 years of helping businesses do more with less, close their books in days, not weeks, and drive down costs. One, because your business is one of a kind. So you get a customized solution for all of your KPIs in one efficient system with one source of truth. Manage risk, get reliable forecasts, and improve margins.
55:05Everything you need, all in one place. As I said, I'm not the most organized person in the world, world, and there's real power to having all of the information in one place to make better decisions. This is an unprecedented offer by NetSuite to make that possible. Right now, download NetSuite's popular KPI checklist designed to give you consistently excellent performance absolutely free at netsuite.com slash I on AI. That's I on AI, E-Y-E-O-N-A-I, all run together. Go to netsuite.com slash I on AI to get your own KPI checklist. They support us, so let's support them. that's it for this episode i want to thank alex for his time if you want to learn more about the conversation we had today you can find a transcript on our website i on ai that's e-y-e hyphen o-n dot ai and remember the singularity may not be near but ai is changing your world so pay attention.
From the publisher
This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.
Download NetSuite's popular KPI Checklist, designed to give you consistently excellent performance - absolutely free at NetSuite.com/EYEONAI
On episode #152 of Eye on AI, Craig Smith sits down with Alex Kendall, founder and CEO of Wavye, a company building autonomous vehicle technology using AI.
In this episode, Alex provides an inside look at Wavye's approach to autonomous driving, which leverages world models and reinforcement learning to create an AI "driver" that can understand complex urban environments. Alex explains how world models allow an AI system to imagine multiple futures before acting, enabling safer decision-making, and shares Wavye's progress on deploying autonomous delivery vehicles with partners in the UK.
We also dive into the differences between world models and large language models, the unique data challenges of perception-based AI, and Wave's ambitions to expand this AI technology to new applications like humanoid robotics.
If you enjoyed this podcast, please consider leaving a 5-star rating on Spotify and a review on Apple Podcasts.
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview, Introduction and Netsuite
(03:45) The Advantages of World Models
(10:12) Developing Autonomous Driving Technology
(16:35) Partnership with Microsoft and Training Challenges
(21:10) Decoding the Concept of Generalization in Autonomous Driving
(27:08) Compute Requirements and Infrastructure
(32:45) The Role of Azure in Training
(37:12) What Is Tokenization of Data
(41:47) Addressing Compute Constraints
(46:12) Future Applications of World Models and Open Source




