Autonomous Driving, Visual AI, and the Road Ahead with Porsche and Voxel51 - Ep. 267

30 Jul 2025 · 41 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

NVIDIA AI Podcast Episode Notes

Episode Title

Autonomous Driving, Visual AI, and the Road Ahead with Porsche and Voxel51 - Ep. 267

Hosts

  • Noah Kravitz (Host)

Guests

  • Tin Sohn (Technical Lead for Vision-Language-Action Models at Porsche)
  • Brian Moore (CEO and Co-Founder of Voxel51)

Key Themes and Discussions

Overview of Autonomous Driving

  • Modern vehicles are heavily reliant on computing power and data, particularly for autonomous driving systems.
  • The transition from rule-based systems to data-driven, end-to-end models is significant in the automotive industry.
  • Synthetic and simulated data are increasingly used for safety-critical testing.

Importance of Data in Autonomous Vehicles

  • The shift from pre-specifying scenarios to training AI with large datasets is crucial for developing autonomous driving systems.
  • Collecting billions of kilometers of driving data is necessary to safely navigate unknown scenarios.
  • Unlabeled data and the need for automated labeling pipelines are emphasized as essential for training effective models.

Levels of Autonomous Systems

  • Level 2: Assistance systems that help a driver (e.g., lane-keeping, automated following).
  • Level 4: Autonomous systems that can operate without human intervention in specific domains.
  • Level 5: Fully autonomous systems capable of handling all driving situations without human input.

Role of Voxel51

  • Voxel51 focuses on making data central to visual AI projects and helps companies like Porsche manage and analyze their data effectively.
  • The platform provides tools for data annotation, evaluation of models, and the generation of actionable insights from data.

Simulation's Role in Safety

  • Simulation allows for testing various driving scenarios that are difficult or impossible to replicate in the real world, enhancing safety.
  • High-fidelity simulations can model complex interactions and scenarios, such as rare events (e.g., a helicopter landing on the road).

Challenges in Autonomous Vehicle Development

  • Different sensor limitations (e.g., LIDAR, radar) complicate the accuracy of autonomous systems.
  • Domain generalization remains a challenge; AI struggles with unfamiliar scenarios that were not included in training data.

Future of Autonomous Vehicles

  • Anticipation and reasoning in driving behaviors are necessary to ensure safety.
  • Foundation models may help move towards more sophisticated reasoning capabilities, enabling vehicles to learn and adapt from broader contexts.

Technological Innovations

  • Recent advancements in AI (e.g., transformer architectures, multimodal models) are improving the capabilities of visual AI.
  • The integration of synthetic data generation is becoming essential for filling gaps in real-world data.

Trust and Explainability

  • Building trust in autonomous systems involves making their decision-making processes transparent and explainable.
  • Systems need to be capable of reasoning in a human-like manner, improving trust between human drivers and autonomous systems.

Key Takeaways

  • The automotive industry is evolving from traditional manufacturing into a data-centric software environment.
  • Safety and trust are paramount in developing autonomous driving systems, with data quality and simulation being critical components.
  • Real-world validation of autonomous systems is essential to ensure their safe deployment.
  • Future models must incorporate multimodal data understanding to enhance adaptability and reasoning.

Resources

  • [NVIDIA AI Podcast Website](https://ai-podcast.nvidia.com/)
  • Voxel51 Website: [Voxel51](https://voxel51.com)
  • Porsche Research: [Porsche Research GitHub](https://github.com/porsche)
  • [Google Scholar for Porsche Research](https://scholar.google.com)

Closing Thoughts This episode emphasizes the significant progress and future potential in the realm of autonomous driving, highlighting the intersection of AI, data management, and vehicle safety. Collaborative efforts between companies like Porsche and Voxel51 are pivotal as they navigate the complexities of this evolving field.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. It's often said that modern cars are computers rolling along on wheels. From performance and safety systems to in-vehicle infotainment, computers control and oversee many, many functions in today's vehicles. Autonomous driving systems, of course, are no exception. The quest to build self-driving cars depends on compute power and, yes, data. Lots of data. Here to delve into the inner workings of autonomous vehicles and the increasingly vital role of data in modern car making, are Tin Son, technical lead for vision language action models at Porsche, and Brian Moore, CEO and co-founder of Voxel51, whose visual AI and computer vision data platform, 51, is used by customers across a range of industries, including, you guessed it, Porsche.

1:00Tin, Brian, welcome to the NVIDIA AI podcast, and thank you so much for taking the time to join. Thanks for having us. Thank you for having us. So maybe we can start with each of you introducing yourselves and just kind of talking a little bit about what you do at your respective companies and a little bit about how that relates to autonomous systems. And of course, we'll get into it. So, Tim, maybe you can start. Right. So I'm Tim. I'm a PhD student and tech lead at Porsche AG, like you said. I'm dealing with vision language action models for autonomous driving. And in my research, I want to turn cars into embodied agents that can understand space, time, and physical properties of the real world so that they are able to act within it and interact with the driver through natural language or through pose, facial expressions, or gesture.

1:48Fantastic. And Brian, tell us a little bit about Voxel51. Yeah, so as you mentioned, I'm Brian, the co-founder and CEO here at Voxel51. First, my background. So I'm a geek, a nerd by background. I have a PhD in machine learning from University of Michigan. Co-blue? Exactly. That is where over 10 years ago, I met my co-founder, Jason, who is a faculty at Michigan. We started off doing some consulting work, of course, being located in Ann Arbor, just down the road from Detroit, the Motor City. We had the opportunity and great pleasure to collaborate with a number of automakers, 10 plus years ago in early versions of autonomy, where we kind of reached that key insight that, in theory, it's all about models and algorithms.

2:28In practice, it's all about data and data quality and data strategy, which led us to the opportunity and need to provide our product 51 to help solve some of those data challenges. Very cool. And we'll get into, as I said, as we talk, we'll get into a little bit more about what 51 is and what Voxel 51 does with customers like Portia. But maybe, Tin, let's start with you. On the podcast, we've said, I think it's actually an Andrew Ng quote originally, but we like to say that AI is like the new electricity. It's sort of there in the background, powering more and more, you know, everything that we do across industries and research disciplines, and it's kind of providing power that we run on.

3:09Why is it so important in the automotive industry and when we're talking about autonomous vehicle systems. Why is it so important to organize and understand these huge amounts of data, whether they're coming from real-world capture or generating synthetic data? Why does that play such a big part in developing autonomous driving systems? So to deal with this increasing scenario space in the open world and the industry is shifting from pre-specifying scenarios and pre-specifying everything which can happen in modular pipelines, which are only partially supported by AI models towards end-to-end pipelines where AI plays the major role or plays the sole role.

3:49And instead of specifying all the scenarios, we are training data and directly mapping them to outputs, to actions which the models have to perform. So we are removing the inductive bias and removing all the rules and knowledge we have and training from the data. This is mainly because we are moving from level two assisted driving to fully autonomous systems that have to operate in conditions which are partially also unknown to us, which we are not able to specify fully. And in order to act safely under these conditions, we have to collect lots and lots of data, so billions of kilometers of driving data which need to be collected.

4:28And not only that we need to explore this whole scenario space to provide safe driving solutions, but also we have long-tail distributions of different traffic scenarios, of different interactions of agents. And not all data has the same value for the model. We have different kinds of redundancies and imbalances in the training data for these models. And this is where we need to explore the data, where we need to capture the data, which is mainly unlabeled. We need to implement automated labeling pipelines in order to understand which are new scenarios, important scenarios having impact on the agents and which are scenarios which are less important, which are redundant, which happen more often.

5:10For instance, we have lots of scenarios with walking pedestrians on zebra crossings, but there's not so much situations where helicopters have to perform a landing on the Eagle Lane of an autonomous vehicle, for instance. Right, but it could happen. Yeah, exactly. The other thing which is important in that regard is that there's not that much ground truth data for different modalities, for different sensor setups. So in order to be able to train models, we need different kinds of data from different modalities. For instance, spatial data or data from multiple sensor setups, from multiple embodiments.

5:45Or in terms of our research, also agent interactions with the physical world or visual question answering data are sparsely found in available real world data. So we need a data curation platform like Voxel 51 provides in order to harness different methods, in order to create meaningful pipelines, meaningful data labeling pipelines, which are relevant for our training, for the training of the models, but also for the validation during the operations of the models. We need to validate and continuously observe because the scenario space is so large that we have the obligation to observe it during the operation too.

6:23So to go back to a couple of things you said, first, just kind of the level set. You mentioned the levels of autonomous vehicle systems and the difference between level two and then going to a fully autonomous system. Can you just briefly give an overview of what the levels are and kind of where today's technology is at? Yeah. So in current vehicles, we mostly have level two systems, which are driver assistance systems. and these systems don't operate on their own. They cannot act in the environment. They just support the driver in certain tasks, for instance, lateral or longitudinal movement.

7:02And in the future or in the near future, we will have fully autonomous systems. So these are like the lane keeping systems? Exactly, lane keeping systems and automated distance systems, which keep the citizens. Like following behind another car? Following behind. Okay, and so that's level two. and then getting up to, is it level five that's considered fully autonomous? Yeah, so level five is fully autonomous. This is basically a human agent which can act in the real world. So if we have an AI which is able to interact fully autonomous without any separation from what a human agent would act in the environment, we have level five.

7:40But we are not striving for level five yet. We are striving for level four, which is that in most domains, and autonomous agents can act and interact. But there are some borders, some boundaries, which we call operational design domain boundaries, boundaries of the system we need to consider. And the system needs to be able to identify these boundaries and to have fallback loops where the driver can engage with the system. Got it. So, Brian, where does Voxel 51 fit into this? How do you work with Porsche? And then if you like also, So you can talk a little bit more about the company, the platform, and how it serves other customers in general.

8:20Yeah, absolutely. So, you know, like Tim mentioned, our platform, our product exists to put data at the center of all of his team and indeed all visual AI projects work. For context, what we see today is that models and weights, while very important, are becoming increasingly commoditized or publicly available. And therefore, the key differentiator between the success and failure leader versus follower status of a product or company lies in their ability to turn the data they have access to into what we call actionable insights or intelligence for their systems. So in other domains, let's say more legacy domains, use cases like structured data, data that fits in spreadsheets and tables and so forth, there's quite a lot of mature tooling that exists out in the market to help you make sense of that data, run computations, process that data.

9:10However, in our experience, when it comes to visual data, image data, video data, 3D data, all the associated metadata that needs to be understood and processed by AI systems, There was really a dearth of tools in the market that played that role of being that key platform that puts data at the center of all the development work. So we experienced that ourselves through our research and our consulting and ultimately identified that a product needs to exist that's open and extensible in the market that teams like Tins at Porsche can adopt, feed their data into, and really facilitate the end-to-end development of their systems.

9:47So that encompasses data annotation or labeling, data curation, and importantly, evaluating models and then turning that flywheel whereby you identify failure modes or gaps in a model's performance and turn that into the right decisions about what new data should I go out and gather to address a specific failure mode. Whether it be real data that's available through an autonomous fleet of vehicles in the world, or maybe synthetic data that's increasingly becoming important to fill in gaps that are unusually hard or difficult to get your hands on. The helicopter landing next to me on my daily commute.

10:24Exactly. You may not have many examples of that at your fingertips. However, in order to truly get to L5 systems, we need to have confidence that these, you know, agentic systems are going to be able to respond and react to those kind of extreme, but nonetheless important events that could occur in practice. Yep, yep. I mean, less unlikely, right? But I always think of kids darting out into the road on a bike, on foot, whatever. And yeah, I want my fully self-driving car to understand how to avoid the kid, the helicopter, 100%. Exactly. Yeah. And so just to sum up what our product offers, it's exactly that.

11:00If you want to, for example, deep dive into the performance of an autonomous vehicle in situations where there's a crowded intersection at night with low light, and you want to develop confidence or trust that that system's going to respond correctly, you need the ability to perform that query, analyze that model's performance. And if it's not up to par yet, take the right actions to build better data sets so that you can get to where you need to be in terms of performance. Fantastic. So, Tin, in the real world, Brian was talking about building that trust. How do you, how does Porsche go about making sure that its autonomous driving systems are safe, are road ready before releasing them?

11:41How do you test? How do you, how does simulation come into play? We're talking about the role of data and everything. Can you talk a little bit about your process? So when it comes to simulation, it is an increasingly important and valuable tool for us. And we have talked about the increasing complexity of the systems, moving towards end-to-end systems, moving towards maybe embodied agents which interact with the driver. And we have talked about the increasing complexity of the scenario space, where we have more and more scenarios which have different parameters. And simulation really gives us the ability to ask the question, what could happen.

12:15So we don't only have recorded data, we also have interaction between agents in the environment. We have different models where we can evaluate the feasibility in all of the state and action space, not only in the recorded state and action space we have from our real test drives. And this is really, really valuable because this gives us the opportunity to have models which generalize over the whole environment and everything which can happen. And simulation enables us to do that by becoming increasingly high fidelity, by providing increasingly realistic environments, increasingly realistic sensor models, and also behavioral models for different agents.

12:59And this is important for us because we really want to have safe systems on the road. And only through simulation we can capture this because in the real world, there's so much which can happen. And there's also some scenarios which are not so easy to replicate. For instance, if a helicopter lands on the road and we cannot create a test case in the real world without harming someone or without having the danger of harming someone. So we need simulation and also synthetic data generation to capture these scenarios. And we have the ability to do that with an increasing level of realism. In the latest developments, there also have been additional aspects of simulation like improving the fidelity with generative models by feeding in scenarios into a generative model and getting more realistic outputs which resemble almost video scenes and you see that also in nvidia cosmos so we are really excited to see the developments which are happening currently in that area which improve our realism and improve the fidelity of the simulation and therefore also the validity to test in simulation environments.

14:07You mentioned that accounting for the unknown or the unpredictable is a big issue with developing safe autonomous driving systems. Are there particular known scenarios, you know, things that happen in the real world that are just really difficult problems to solve when it comes to building an autonomous system to deal with it. You mentioned low light, crowded intersection, nighttime, those kinds of things. Are there particular scenes or even particular variables that just really pose a challenge for working on these systems? Definitely. So we have certain physical limitations of sensors. For instance, radar and LIDAR sensors have certain limitations.

14:51LIDAR sensors are limited by, for instance, rain or water reflections or reflective surfaces. While radar sensors can be noisy in the presence of certain magnetic fields or metals. And camera, of course, we know when camera becomes noisy in many weather situations. So we have different physical limitations and we need to also to model them and to understand in which situations agents can act. Because it's not a physical limitations and where these agents have physical limitations too. And simulation also provides, in the latest development, through realistic sensor models, also the ability to evaluate virtually in that regard.

15:31When it comes to agent interactions, we have a lot of limitations because of the problem of domain generalization. We have different types of scenarios where agents need to interact. And if the data is not present in the training data, we have the issue of learning concepts and applying concepts like human driver do. For instance, if I'm a human driver, I see certain entities in the environment, which I never saw, but I can anticipate them. But an AI agent has difficulties anticipating them because when it didn't see it in the training data, it cannot generalize over the unseen event. And this is also where these methods come into play.

16:19We need to capture also these situations. One thing I would add on the synthetic data front, one thing we're seeing is that it's becoming increasingly important to be able to generate variations of a scene. For example, we're excited to integrate with NVIDIA's Cosmos World Foundation models in our product. And one thing that enables is you can import a realistic scene, parametrize it, and then tune different things. What would it look like if that vehicle was a different shade of beige? Or what would it look like if there was another vehicle or a pedestrian in the same scene? Or maybe let's change the weather conditions.

16:52And that ability to kind of play around with a realistic scene is important to kind of develop a trust that the system is going to be able to deal with all the different variations that it might see of that scene in practice. Right, right. Is that a manual process or are you automatically generating the variations to put them back into the system? Yeah, that's one of the exciting things about the software integration we have with the Cosmos Foundation models as an example. users of our product can identify a scene and then click a button to automatically generate all of those variations and pull them back into their training data set and indeed automate the analysis or assessment of all those different variations.

17:34So we think that having humans in the loop is definitely important. It would be a mistake to sort of ship a system without having a human's eyes on the scene to identify or spot check or build trust in performance in key scenarios. But of course, the volume of data that you need to get to the performance that you need is immense. And so you have to leverage automation whenever possible. Brian, as Voxel's worked with Porsche and other leading automakers, what's been surprising to you along the way? Maybe unanticipated challenges, maybe getting past a hurdle in a surprising way. What are some of the things that stand out to you as you think about working with automakers?

18:15Yeah, I think that the main thing I would say, as Tin was describing all of the very interesting software or related systems challenges in bringing autonomy to market, what we're seeing is that the leaders in the space really are reinventing themselves not as automakers, but really as software companies. Right. And that's the type of sort of skill and expertise that's going to be needed to really solve these problems and bring a differentiated product to market. There's a trend today where, you know, if you think of the software component of autonomy as something that you can procure off the shelf, maybe through a vendor, then it can get you to a certain level of performance.

18:54But the key, you know, as I argued before, the key to, you know, really industry leading performance is to harness the data that your company has access to. You really need to bring the development and iteration of that software in-house rather than just outsourcing it to reach kind of leading status. Yeah, from a consumer's perspective, when you start talking about software and automakers, I think of infotainment systems just, you know, as a driver, as a passenger, kind of the first light up thing that I see, right? And I've been following a little bit from afar, but, you know, kind of this almost like dance between the automakers and some non-automotive software makers and, you know, mobile phone makers in particular, as you get into plugging your phone into the car.

19:36And then maybe like with Apple CarPlay or Android Auto, it takes over the in-car infotainment, that kind of thing. But then listening to you talk about it and thinking, well, OK, let's move from infotainment to something I can't even imagine how complex it really is as an autonomous driving system. And it makes perfect sense to me, Brian, what you're saying, that to really develop a top-notch system, you can't just grab something off the shelf and plug it in. Like, you've got to be, you know, shaping it to fit your vehicles, all the data you have, all of that kind of stuff. Yeah. And it's also important, I think, to not take it all in-house and say that you can do everything.

20:13I think an effective pattern that we've seen in the market, and of course, Tim would be an authority on this more than me, is, well, let's first focus on being able to validate the performance of systems so we can truly understand if we're going to work with a vendor for a certain piece of technology, can we truly understand its performance and develop trust in it? And then we can evaluate whether it makes sense to bring certain aspects of the system in-house so we can fine-tune it with our own data and so forth. So that focus on validation and evaluation is definitely important in the short term.

20:43Brian, you may have said this at the beginning, so forgive me, but how long ago was Voxel 51 founded? Yeah, so Voxel started over 10 years ago now as a kind of a consulting partnership between myself and my co-founder, Jason, as a sort of venture-backed software company. that that journey started in 2018. So seven years in market. Okay. So along that path, almost a decade now, or, you know, seven years to market in the decades that you started, have there been particular technological breakthroughs that have really allowed Voxel to just take your practice to the next level and, you know, do things, offer things to customers you couldn't?

21:20And particularly when it comes to, I mean, this is what you do, but when it comes to wrangling and managing and understanding these vast qualities of data, Are there particular technological advances along the way that have really, you know, made the work that you do possible? Certainly, there's been just immense technological innovation on the model and algorithmic side. Yeah, it's one of those questions where I'm sort of like, have there been innovations in the past 10 years? Maybe a couple. Yeah, yeah. So one of the, you know, so there's been clear advances in, you know, technologies like the transformer architecture, which kind of leveled up another kind of order of magnitude of performance potential.

21:56And then, of course, we had all of the advancements in chat GPT and large language models. And now we have models in vision that are more multimodal in nature, vision language models that can pull in information from text and audio and fuse that with vision. The long term, I think, vision in the space is that we need models that can go directly from pixels or sensor inputs directly to actions or decisions. And that kind of end to end system is kind of the holy grail. and it's very exciting to see all of that develop. The lessons learned from us, it always requires more data, more data, more data to organize, to understand, to sift through, to find kind of the needles in the haystack.

22:37And so the number one feature request we get from our customers is definitely, hey, you know, as I'm thinking about my plans and my goals for next year, it's involving an order of magnitude plus more data, the need and ability to connect to more GPUs to compute on that data. And so we've definitely benefited from the rapid pace of innovation and distributed computing and related technologies that our platform can plug into to help deliver that scale to customers. To dig in on something you said real quick, and correct me where I'm wrong here, but I think you said that the holy grail, as you put it, is moving to a system where a model that can go from pixel to action directly.

23:14Where are we at now? How do you get from pixel to action currently? Yeah. First of all, just to unpack the historical context there, we refer to the space as visual AI. what you may have referred to it in the past as is computer vision. And that historically has represented kind of very low-level tasks, like taking an image and classifying it as a certain animal or drawing a box around a certain object. So that's a very kind of low-level task. It's important information, and the system needs to understand the content of an image or a video stream in order to reason about it. However, I think the lesson that we keep learning, even on sort of the language models, is that to the extent that we can push the system to be more end-to-end and, you know, have the authority to do a lot of reasoning itself and go directly from raw inputs to the decision, there's a capacity for more sort of intelligence.

24:06Are we at that step yet? Certainly not. But I think that's where the leading edge is in terms of research and the work that Tim and his team at Porsche are doing. It's very exciting times. Gotcha. A long time, well, not that long ago, but a while ago, I read a novel that was kind of a, you know, not quite cyberpunk, but along those lines, tech heavy. And one of the little threads in the book was about self-driving cars and near future freeway system in the United States where the cars talk, you know, automatically to each other and to the toll taking systems on the roads and all that kind of stuff.

24:38But one of the themes that came out was, you know, if every car on the road was autonomous, that would be potentially the ideal situation for human safety. Because if all the vehicles were self-driving and they're all performing at a good level, they're going to be able to make much better decisions than human drivers just because of the raw ability to compute and deal with all the data and just everything humans can't quite do. Given where we're at now in reality and level two systems and all the stuff you've been talking about, how do we use technology? How can technology help improve, help build human trust in autonomous vehicles?

25:16So the question you asked is how to establish trust towards the human and where human has most trust in our system, which are explainable and are acting similar to humans. So one big limitation also current end-to-end systems have is that they are not able to describe why they are taking certain actions. They do not derive their actions from basic concepts like humans do. humans can understand from several hours of driving how to drive in the world while autonomous systems can't do that and they have to have big amounts of data in order to derive actions and where the shift needs to happen is towards systems which are able to derive and generalize based on certain contexts derived from human knowledge like derived from different basic concepts like we learn in driving school how to act in the environment and also how to explain and describe these actions.

26:20And only when we are able to do that, and this is also done in the research on foundation models in the area of autonomous driving, only when we are able to do that, we have trustworthy systems which are also able to interact with the driver, with the co-driver in that case and to explain their actions and also to explain when certain boundaries are met. Like human drivers also have certain boundaries. They also cannot perform in all situations. So this is the shift we need to make. Right. Brad? Just to add a few other angles to that. So I try to hold two things at once together. One is the excitement for the long-term future of fully connected, fully autonomous vehicles, which are orders of magnitude safer than humans and can benefit from other information that human drivers don't like the ability to directly communicate with other vehicles or other sensor types like LIDAR that offer information that humans don't have.

27:17At the same time, I think we'd be remiss not to be very excited about all of the very concrete benefits that are already in the market today, specifically around safety, right? You know, the fact that my vehicle today has L2 systems that can automatically detect, you know, maybe I've lost focus on the highway and a car in front of me is stopped in traffic, my vehicle can already react to that and keep me safe. And I think there's so much more potential to roll out in more of an incremental way. AI enabled advancements that focus specifically on trust and safety. And, you know, Tin painted a great picture of what's needed from a explainability and transparency standpoint on the technology, but also very excited about the very concrete benefits that everyday drivers can see over the next five years while we continue working towards the grander future and the sort of potential of a fully autonomous network is that kind of excitement technologically that keeps pushing the face of innovation forward.

Read the full transcript

28:14And we're happy to help do our part to enable that to happen at Voxel. But definitely appreciate both aspects and tension. Right. Tin, along those lines, what excites you the most about the future of autonomous vehicles at Porsche? So I'm probably biased because I'm conducting research in that area. But what really excites me... I hope you are. What really excites me is turning cars into Knight Rider, if you know the series Knight Rider. Oh, yeah. So turning cars into embodied agents which are able to anticipate the environment similarly how humans do. So we can leverage foundation models, leverage the knowledge of the world, of the web context we have.

28:55in order to not have direct mappings of actions from inputs like end-to-end models have, but to have some kind of reasoning process in between and to generalize over different scenarios regardless of embodiment. Like end-to-end approaches are maybe error-prone and sensitive to different embodiments, to different sensor setups. We need to retrain them based on, if we change the sensor setup, we need to retrain the whole system in many cases. And foundation models promise to have some kind of intermediate representation, which incorporates web context and incorporates world knowledge and gives us the ability to generalize on different kinds of scenes and act in these scenes, even if not previously trained on that specific scene, because we have this reasoning process.

29:46And another part of the whole puzzle is that we have completely new types of interactions. We have vision language navigation. Me as a driver, I can tell the car what to do and the car knows about its abilities and can choose certain actions which it knows and create new experiences for the driver. So for instance, looking for a parking lot in the shadow or any other situational request the driver could have, we can create new actions which are not previously trained by any model but are derived from the context and derived from the world knowledge which is pre-trained on those models. So there's a lot of potential in there, but there also needs a lot to be done in regards to that.

30:31In our research, we have investigated what foundation models need to be able of in order to fulfill this task of vision language navigation. And together with also in close collaboration with Voxo51, we have identified four areas. And this is semantic understanding, which is classes, affordances and attributes. Spatial understanding, which is the locations, the orientations of objects to watch each other. Temporal understanding, which is the development over time in the past and the future. And physical understanding, most importantly, which is the world model, the physical rules like forces or gravity applied to the environment.

31:09Or in very specific terms, also, for instance, vehicle dynamics, which have to be considered when planning actions. And we have identified that current foundation models are very good in semantic understanding, like deriving these basic concepts of scenarios. But we need to improve them in spatial, temporal and physical understanding in order to really grasp the task of vision language navigation. And we're really, really excited what's happening currently there in this environment. Also in research done by NVIDIA and other players in that area. And we are looking forward for the research. And on our own, we also train models and create models which are hopefully able to grasp spatial temporal understanding and also physical rules of the environment in order to interact with it.

31:54It's good that you guys are the ones working on safety because, Tin, as soon as you started talking about the driver requesting the car to do something, I immediately thought, man, if I had a Taycan, I could just ask it to get me places fast. But I don't think that's quite where we're getting at here. But Brian, along the lines of safety, where do you see the biggest opportunities for improving autonomous vehicle safety through simulation, through access and tools to help you use better data? Yeah. So for us, it always comes back to data and the important role that data plays in the success of AI.

32:29And fortunately, there's some pretty exciting technological advancements happening in the data space. So historically, one of the most onerous and costly and time-consuming aspects of an AI project was data annotation. The need to gather a data set and have it labeled to sort of teach the model all the information that it needs to know. Interestingly, that was historically done in sort of an outsourced way where the data was shipped off to human teams to do the rote work of labeling that data. You can imagine that's very expensive and creates an artificial bottleneck on the amount of data that you can feed to your systems.

33:04Actually, these days, with the emergence of these generalized foundation models, on our team, we've done a study in comparing the performance of AI systems that are developed on human annotated data versus automatically labeled or auto-labeled data from foundation models. And of course, the benefit of leveraging automatic labeling is that it's far more efficient, lower cost, and so forth. And the question was, well, how did the performance compare? And interestingly, we found that you can achieve comparable performance replacing human annotation in many situations with auto-labeling, which kind of opens up the valve of the potential quantity of data that you can feed to specific systems in use cases like autonomy.

33:47And so we're excited to bring auto labeling to our users in the 51 platform. We call it verified auto labeling. The verified part is important because it's not just about feeding tons of data to a system, but being able to verify the correctness of that information. And so we've developed some technology internally that can kind of help you get the most value out of the knowledge that foundation models have while also prioritizing verification and trust in your users understanding the performance. And then of course, there's the simulation piece, because if you can't get your hands on exactly that scene that you know that your model has a weakness or a failure mode that needs to be addressed, tapping into simulation techniques, as we mentioned previously, to fill in those gaps and build higher quality data sets, very exciting in terms of the potential to push us to the next level of performance.

34:41In the LLM space, we hear today things like, oh, well, we've already trained on all of the information available on the public internet. And so now we need to resort to synthetic data because there's nothing left, right? Now, I'm not sure that that's exactly true, but it's definitely the case that synthetic data generation techniques have different complementary capabilities to real data. And increasingly in the future, we see it being an integral tool to the, you know, toolkits of teams that are putting data quality at the center of their development efforts. As we wrap up, let's hold on to that future forward mindset for a moment here.

35:17And I'm going to ask you if you can think ahead five years from now. So 2030, wow. So if we're in 2030, what do you hope will have changed, will have progressed and developed in the world of autonomous vehicle safety and simulation five years out? Well, I'm definitely excited in general about the opportunity that autonomous driving plays and shedding light on the powerful capabilities of visual AI and multimodal models. I think from a geek standpoint, it's a perfect testing ground to test the value of multimodal systems that can understand not just the visual inputs or LiDAR inputs, but also things like you mentioned before, a fully connected system where we can take information about the intentions or behaviors of other vehicles to bring us to that next level of safety or feed in audio signals to build a more holistic understanding of a scenario, I think it's going to unlock the next level of safety.

36:15When the model can understand the urgency of my swearing behind the wheel, it'll know how much danger we're really in. Exactly, yeah. So maybe I can add on that. Yeah, please. So I believe that safety will shift from like a holistic view where we have to really test all of the necessary scenarios for a specific domain towards more situative safety because we can have models which are able to reason on different situations, which are able to derive from specific basic concepts and are able to interact with the driver so that there can be safe conditions even in unsafe situations. So that the driver, for instance, gets requested to a takeover by a model which talks in natural language to the driver Or also the model can anticipate different situations based on the concepts behind it as an unsafe situation where it can act accordingly.

37:12And like Brian said, the multimodality plays a major role. So we will have models which are not only able to reason on visual data, not only able to reason on camera data, for instance, but also on spatial data, on point clouds, by radar and LIDAR sensors, and also on map information and other information which is available in the environment. And therefore, we believe that the future models should be able to have more of a situative understanding of safety, like humans also have. They do not need to capture all of the situations, but they can be able to anticipate the situations and act safely even in an unsafe situation.

37:52Anticipate also unsafe behavior by other agents, and this is where we have to go. Yeah, that whole concept of anticipation is such a big part of living life as a human, right? And so it makes sense, the importance of it within these systems. But I will leave the complexities of getting it to work to folks like you. It's fantastic for us as users or enthusiasts of automobiles that there's companies like Porsche out there, because I'm sure that they'll balance the safety and automation with the fun factor of being a driver behind the wheel. Whether or not you have to be driving, I'm sure it'll be a first class experience with Porsche involved.

38:29No, that's a great point because there are, I think anyone who's driven a car has probably taken at least one drive just for fun or to relax or to get away from it, clear your head, you know, and so just that driving experience is something that I hope we can hang on to going forward. Absolutely. Ten, Brian, this has been a great conversation. I've learned a lot about the current and future of autonomous vehicles. So thank you both for joining. For listeners who would like to learn more, would like to dig a little deeper into what Porsche is doing with autonomous vehicles and everything else, into what Voxel 51 is all about, where would you point them to go on the web to get started?

39:08Brian, I assume the Voxel 51 website, but where can they go? Definitely check out our website. For those technologists in the audience, check out our open source project, 51, completely free, openly available. Download it, kick the tires, test out this data-centric view of developing your next visual AI system. Fantastic. And Tim, is there a Porsche research blog? Is there a part of the website devoted to autonomous vehicles? Where's the best place for a listener to start? Yeah, so we also have a lot of our research open sourced on the Porsche free and open source website. For instance, a benchmark for evaluating foundation models and the task of vision language navigation based on the four capabilities we just mentioned.

39:49And also a lot of research papers you can grasp and just check out our Google Scholar and our GitHub page where you can dive into what we're doing and foundation models for autonomous driving. Fantastic. Again, thank you guys so much. And, you know, anytime you want to talk cars and autonomous vehicles and all that stuff, give a call and maybe we can catch up and do it again in the future. Looking forward to it. Thanks, Noah. Looking forward, Noah.

40:35Thank you.

41:04Thank you.

From the publisher

Tin Sohn, technical lead for vision-language-action models at Porsche, and Brian Moore, CEO and co-founder of Voxel51, explore how AI, data, and simulation are shaping the future of autonomous vehicles. They share insights on the industry's transition from rule-based systems to data-driven, end-to-end approaches, the growing use of synthetic and simulated data for safety-critical testing, and how foundation models can enable cars to reason, act, and even interact like human drivers. Learn more at ai-podcast.nvidia.com.

More from NVIDIA AI Podcast

All 115 episodes
Autonomous Driving, Visual AI, and the Road Ahead with Porsche and Voxel51 - Ep. 267NVIDIA AI Podcast · 41 min
Listen in VO