Driving Safer AVs Faster with Smart Simulation, Neural Reconstruction, and Data-Centric Tools - Ep. 289

11 Feb 2026 · 45 min · 19 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

NVIDIA AI Podcast Episode Notes

Episode Overview

  • Title: Driving Safer AVs Faster with Smart Simulation, Neural Reconstruction, and Data-Centric Tools - Ep. 289
  • Description: This episode discusses how autonomous vehicle (AV) teams can effectively manage vast amounts of data to improve safety and performance by utilizing smart simulation, scenario-driven data curation, and NVIDIA-powered tools.
  • Hosts: Noah Kravitz with guests Rohan Bhasin (Fortelix) and Dan Gural (Voxel51).

Key Themes

Introduction to Guests

  • Dan Gural: Head of Technical Partnerships and Machine Learning Evangelist at Voxel51, specializing in data platforms for physical AI and AVs.
  • Rohan Bhasin: Senior Solutions Engineer for Sensor Simulation at Fortelix, focusing on synthetic data generation and sensor simulation tools.

Understanding AV Simulation

  • Definition: AV simulation involves creating realistic driving environments to train and test autonomous vehicles, enhancing their safety and efficacy.
  • Evolution: Shift from camera-centric solutions to end-to-end systems that incorporate generative AI for improved data fidelity.
  • Importance of Data Quality: Emphasis on the adage "garbage in, garbage out," underscoring the necessity for high-quality training data.

Sensors in Autonomous Vehicles

  • Types of Sensors:
  • Cameras (most commonly used)
  • LiDAR (light-based sensing)
  • Radar (used by some manufacturers)
  • Variety of Data: Autonomous vehicles utilize a range of sensors to gather diverse data, which is crucial for accurate model training.

Challenges in AV Development

  • Data Overload: AV systems often face challenges managing petabytes of data effectively.
  • Gap Identification: Companies need to identify gaps in their data, focusing on critical edge cases rather than collecting more nominal data.

Innovations in Technology

  • Neural Reconstruction: Provides unmatched fidelity in simulation compared to traditional physics-based rendering methods.
  • Scenario-Driven Data Curation: An automated process by Fortelix that labels events in driving logs, allowing for more efficient data analysis and scenario creation.
  • Physical AI Data Engine: A tool by Voxel51 designed to streamline data access and management across various AV teams, enhancing collaboration and data utility.

The Role of Generative AI

  • Foundation Models: These models facilitate quick scenario generation, allowing teams to create diverse variations in simulations without extensive real-world testing.
  • Efficiency Gains: Focus on reducing the time it takes to train models and validate scenarios, which is critical for speeding up AV deployment.

Evaluating Realism in Simulations

  • Realism Measurement: High fidelity in simulations is crucial for ensuring that the model behaves appropriately in both training and testing environments.
  • Trade-offs: While realism is important, the ability to generate useful training scenarios quickly is also valued.

Collaborative Tools and Team Dynamics

  • Streamlining Processes: Emphasis on creating streamlined workflows to ensure effective communication and data sharing across teams.
  • Future Team Structures: Possibility of flatter team structures that reduce redundancy and enhance collaboration, driven by advancements in simulation technology.

Key Takeaways

  • Data-Centric Approach: Autonomous vehicles can benefit significantly from a data-centric approach, focusing on curating and refining existing data to improve safety.
  • Technological Advancements: Innovations in neural reconstruction and generative AI are reshaping the landscape of AV simulation, enabling faster and more efficient development cycles.
  • Continual Learning: The field is rapidly evolving, making it essential for teams to adapt and leverage new technologies to enhance AV safety and performance.

Conclusion The discussion highlighted the rapid advancements in AV technology and the necessity for effective data management and simulation strategies. The integration of neural reconstruction and scenario-driven data curation is paving the way for safer and more efficient autonomous vehicles.

Additional Resources

  • Fortelix: [fortelix.com](https://fortelix.com)
  • Voxel51: [voxel51.com](https://voxel51.com)
  • NVIDIA GTC Conference: [nvidia.com/gtc](https://www.nvidia.com/en-us/gtc/)

Feel free to reach out to Dan Gural or Rohan Bhasin on LinkedIn for more insights and discussions on AV technology and simulation methodologies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Guest Introductions and Backgrounds

1:07 to 2:44

Introduction of guests Rohan Basen and Dan Goral and their expertise.

“we were one of the premier data platforms for physical AI and AV, allowing you to help find all the kind of problems and curate your data set to make sure you're always focusing on the most important data.”

Overview of AV Simulation

2:44 to 4:54

Discussion on the evolution and importance of AV simulation.

“I'm sure Rohan can give a better answer for this one.”

Sensor Data in Autonomous Vehicles

4:54 to 7:10

Exploration of sensor types and data used in autonomous vehicles.

“Famously, Tesla does not have LiDAR, just calling that out.”

Challenges in Data Translation

7:10 to 8:36

Discussion on the gaps in data translation from physical to digital.

“or that drive that you were looking for in the real world.”

The Role of Neural Reconstruction

8:36 to 10:38

Insights on how neural reconstruction improves AV simulations.

“Can you guys talk a little bit about foundation models and neural reconstruction and how these two technologies in particular have changed expectations around simulations, AV simulations?”

Addressing Edge Cases in AV Development

10:38 to 12:32

Discussion on the significance of edge cases in AV safety.

“I mean, just to show you where we were five years ago, this is not a joke.”

Closing the Safety Gap in AV Systems

12:32 to 14:00

Examination of the challenges in achieving high safety standards.

“What are some of these, you know, either most important, most difficult edge cases to test?”

The Challenge of Safe Autonomous Driving

14:00 to 19:23

Explore the challenges of achieving high safety standards in self-driving cars.

“There's just no way in the world you would ever release a self-driving car that only drives safely 90 % of the time.”

Realism in Simulation and Data Utilization

19:23 to 27:04

Discuss the importance and methods of realism in AV simulations and data usage.

“Yeah, I completely understand that perspective as well.”

Ensuring Consistency Across Autonomous Driving Models

27:04 to 28:00

Learn about the need for consistency across different components in AV technology.

“And sort of is there a commonality amongst the data?”
Show all 19 chapters

Exploring Data Language and World Models

28:00 to 29:10

Learn about the challenges of data formats and the importance of effective communication in world models.

“This is what we're pushing out here to understand, hey, I'm not saying that I'm going to push out this new data language.”

NVIDIA Platforms and Neural Reconstruction

29:10 to 30:25

Discover how NVIDIA's platforms enhance simulation capabilities for AV teams.

“So I don't have to ask as many questions.”

Maximizing Business ROI with Old Data

30:25 to 31:35

Understand how past data can be repurposed to improve autonomy and safety in AVs.

“And I can take this old data, doesn't matter how old it is, doesn't matter if it's sitting on a hard drive for years, and now start to generate these augmentations or variations that are really interesting.”

The Physical AI Data Engine Explained

31:35 to 33:00

Learn about the Physical AI Data Engine and its impact on AV workflows.

“Is you're basically painting a 3D picture, right?”

Data Auditing and Enrichment Techniques

33:00 to 37:45

Explore the processes of data auditing and enrichment to enhance data quality for AVs.

“trying to wrangle all of this data at once?”

Future of AV Teams and Simulation

37:45 to 41:42

Examine how advancements in simulation will reshape AV team structures and roles.

“We always like to wrap on kind of a forward-looking note, but in this case, it's going to be kind of a very specific question kind of following everything we've been talking about.”

Exploring Voxel51 Resources for AV Simulation

42:05 to 42:55

Learn about the open-source resources available for self-driving examples.

“just the state of AV simulation and everything we've been talking about.”

Connecting with Experts on LinkedIn

42:56 to 43:38

Understand how to connect with experts in the field for further inquiries.

“So if you ever have any follow-up questions if you're interested in chatting more or just interested in connecting I recommend you reaching out to me on LinkedIn.”

Future of Autonomous Vehicles and Events

43:39 to 44:24

Discuss the fast pace of AV development and upcoming industry events.

“and see if there's anything that interests you there.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:10Welcome to the NVIDIA AI Podcast. I'm Noah Kravitz. NVIDIA GTC is this March 16th to 19th, online and in person. Visit nvidia.com slash GTC to learn more about the premier global AI conference and to register now. Today we're talking about autonomous vehicles and AV simulation, using AI-powered systems to make driving safer and more efficient. With me are Rohan Basen and Dan Goral. Rohan is Senior Solutions Engineer for Sensor Simulation at Fortelix, and Dan is head of technical partnerships and a machine learning evangelist at Voxel 51. Gentlemen, thank you so much for taking the time to join NVIDIA's AI podcast.

0:51Welcome. Thanks for having us here. Thank you for having us. So before we get into talking about all things AV simulation and driving, maybe we could start with kind of self-intros. Dan, maybe you can start. Just kind of tell us briefly what your company does, what your role is there. Yeah. Like you said, my name's Dan. I'm from Voxel 51. we were one of the premier data platforms for physical AI and AV, allowing you to help find all the kind of problems and curate your data set to make sure you're always focusing on the most important data. One of the things we always love to say around here is garbage in, garbage out, right?

1:27So if you're not training on the right amount of data or the right data, you're not going to be able to get that model that you're hoping for. Personally, my background is in all things physical AI. Before we called it physical AI, I spent a lot of my time working on edge devices such as drones, cars, robotics, and many such. I have a very strong background in hardware and things like that as well. But now, of course, as we evolved more towards this physical AI world, I can now kind of slap that label on it and everyone should know what I work with. Awesome. Rohan? Hi, I'm Rohan. I'm a solutions engineer at Fortelix, and I help our customers work with sensor simulation tools from a variety of providers to meet their synthetic data generation needs.

2:09My background is in sensor simulation as well. I was part of developing Ford's internal simulation toolchain for their Level 2 and Level 3 driver assistance stack. Very cool. Maybe we can start with a little bit about what AV simulation is. We've talked about it on the show before. A lot of listeners will be familiar. But maybe one of you guys could give just kind of a brief overview of what AV simulation is, why it matters in your work and in the broader mission of making autonomous vehicles real and safe and worth using. And then you can get into talking about what the ecosystem is like today.

2:47I'm sure Rohan can give a better answer for this one. So I'm going to pass this one to him. Okay. Yeah, definitely. I think the AV simulation ecosystem has definitely changed a lot since I started working on this six, seven, eight years ago. A lot of focus then was on camera-based solutions and using bespoke neural networks for individual perception tasks. So think tasks like object detection, lane detection, semantic segmentation, and fusing those together in the vehicle. Now we see a drastic shift away from that into kind of more end-to-end solutions where we have one stack that does all the perception controls and outputs steering throttle and brake commands to generate a driving trajectory that's safe.

3:27I think a big reason for that has been developments in generative AI, right? We talk about 3D Gaussian splatting or diffusion-based models, and both of those have contributed significantly to the shift in sensor simulation data and the fidelity of sensor simulation data that we're able to get and use for downstream training or validation of our AV stacks. I'm sure there's not a one-size-fits-all answer to this. But when we're thinking about, for the listener who's thinking about, you know, sensors on an autonomous vehicle and gathering all this data and what does this all mean, on a typical, you know, car or truck, whatever works, how many sensors are we talking about?

4:05Roughly, how many sensors does each vehicle have? And, you know, what are the kinds of data that you're talking about when we're thinking about all these different sensors pulling data? Yeah, that's a great question. and it's always fun to get that question because it really depends on who you ask, right? If you go walk into Waymo, they'll tell you one answer where we have 25 different sensors and this thing and that thing. If you go ask Tesla, they're going to tell you a completely different answer, right? Okay. Just cameras. In reality, what we typically see is you're going to have a mixture of almost always camera data.

4:38That's things like videos and images. And then we call LiDAR data, which is, if you're familiar with things like radar and such, it's laser-based, pointing to understand how far away things are. Basically, your car shoots a laser at the car in front of it. That laser comes back. I know the car is 10 feet in front of me now, right? Famously, Tesla does not have LiDAR, just calling that out. But for the majority of everyone else, they do have either LiDAR or radar on their cars. Now, something that Rohan touched on as well is we're seeing this new adoption of even additional sensors on top of this.

5:09Traditionally, when we were doing lane line detection or actor prediction, things like that, We only really relied on the video and the LiDAR in order to get these predictions. But now we also have a lot of additional sensors to measure a lot of the physical components of the car that is being fed into our models as well. Things like very things you'd be common to find on every car, like a speedometer, of course, to understand how fast we're going, but also very complex physics-based sensors to help understand exactly how the car is kind of positioned, where it's moving, which way it's turning, which way it might be getting pulled and such as well.

5:41And so when we're kind of moving from physical world, real world, I hate saying the real world, but the physical world and then simulations, where are the gaps between what's available in the tools and the workflows that devs work with in these simulations and what AV devs actually want or even need for the work that they're doing? Is it a matter of getting more and more accurate simulated data? Is it something entirely different? Where do the gaps exist that you're trying to address? Yeah, this is a great one. And I'm sure Rohan and I, we can bounce off of this one because he's really good at the second half of this answer I'm going to give.

6:22And then I'll cover the first half here, right? So the first thing I want to touch on is, is it just more data? The answer is absolutely not, right? If it was just more data, we would have self-driving cars, right? If you're not familiar with this, most very serious self-driving car companies, your Wamos, your Waves, your Zeekses, they're on the 100-a-petabyte scale, typically, or somewhere close to that. If you don't know what a petabyte is, 1 ,000 terabytes is one petabyte. So we're talking about... It's a lot of data. A lot of data. Exactly, right? So where I see a lot of the issues today with a lot of users that I interact with, a lot of problems, is something that you touched on there.

6:59Is the actual translation from the physical to the digital. Did you do everything correctly? Is everything calibrated correctly? because it's all well and good if you've got exactly that scene or that drive that you were looking for in the real world. But if you can't properly communicate that back into your model, you're going to have a lot of issues, of course, right? If it doesn't see it the same way it actually happened, your model is going to degrade because of that performance. And that's a lot of what Voxel 51 is supplying to our AV users and such. It's ensuring that digital to, or physical digital translation is as clean as possible.

7:33Now, there is a lot of other things that people are going to be expecting whenever you do these translations and what happens after the translation, right? Things such as, well, I don't want to be doing this just for pulling out of the garage, for instance, right? There's scenes that matter more than other scenes, right? And this is where Fortelix really, really shines in their tool on how they can help identify these cases. Thanks for that. Yeah, exactly right. I think for us, as Dan said, there's a lot of data already out there and Fortelix's goal is to help AV developers make sense of their data, right?

8:03Identify what data they have and not necessarily generate or collect more data, but use that data smarter and fill in the gaps with exactly what they need, right? Not scenarios of you pulling out of the garage, but niche edge cases that are safety critical. You know, we have tools that allow you to generate smart replays and create unique and interesting scenarios which are going to stress your stack and test whether it's actually up to the mark. And that's the kind of, you know, smart and I guess niche data that you need to generate that little bit of additional model performance rather than throwing hundreds more petabytes than what you already have.

8:41Right. Can you guys talk a little bit about foundation models and neural reconstruction and how these two technologies in particular have changed expectations around simulations, AV simulations? Yeah, neural reconstruction definitely is something that's created like a step change in what AV developers look for in simulation. The level of fidelity that we can get from neural reconstruction is unmatched by any kind of physics-based rendering engine. It's just not even a close competition. You know, this technology is still developing. It's only a few years old, and there's still improvements to be made.

9:14But the goal is much more approachable now than it ever was before. And with Fortelux's Smart Replay technology, you can take a relatively uninteresting snippet and generate variations on that. We can have pedestrians crossing in front of you, other vehicles driving in behaviors that you don't expect, breaking the rules and creating interesting scenarios with a level of realism that was unimaginable even just five or seven years ago. Yeah, I mean, I think to add to that, I think the most important thing to understand here, I like to make this point very often, is that the most valuable resource when you're developing an AV system is not GPUs, it's not data, it's not people, it's time, right?

9:56How can I train as fast as possible? How can I test as many cases as possible? Right. And so every time you train a scene that it's not going to make your model better, you just waste it a lot of time. Right. And so with neural reconstruction and foundation models, we don't have to go drive and find the snowy days. I'm not checking the weather to figure out when it's going to rain next time. Right. I can use models like NVIDIA Cosmos Transfer is a really great example to just generate, say, hey, this is the scene I'm interested in. But how do I do when it's raining? How do I do if it's snow? oh, you can use a tool like Fortelex to change the actors.

10:31What if I didn't turn left? What if I turned right, right? And the main goal here is to save time, right? Both because these neural reconstructions render faster, it's easier to move things around. I mean, just to show you where we were five years ago, this is not a joke. This is not an exaggeration. We literally had people playing GTA V and crashing into other people to capture data, right? This was not a joke. This was state-of-the-art even at the time, right? And so think about how much time we're saving and how far we've come in only a few years where now we can just imagine these things with foundation models and neural reconstructions, Gaussian splatting, where we can skip not just driving in the real world, not playing GTA, but actually just doing this in a matter of a couple prompts or clicks in order to figure out where we need to go next.

11:16I don't know if I'm more amazed by the fact that that was only five years ago or that like we've come that far in five years ago you know what I mean I'm kind of like absolutely oh man that was a long time now it's only five years you know but but I mean things have advanced so much is it I shouldn't say I don't know if is it still is the right way to ask this but Rohan you mentioned this the edge cases of, you know, somebody, a kid running out onto the road, you know, what have you, sudden weather change or an object in the road, that kind of thing. Are these the kinds of things that are still amongst the most, I don't know, sort of time and bandwidth intensive problems to solve when it comes to, you know, building a safe AV stack?

12:03Is it still kind of solving for these unknown edge cases? or what are, I see you guys nodding, people can't, you know, listening. Of course, yeah. Yeah, yeah. What are these things? Because, you know, I was going to ask you about this process of, and again, you talked about it a little bit, of, you know, creating this high-fidelity 3D reconstruction and then you can kind of turn it into all these different edge cases to test for and to kind of stress test on your stack. Is it these, you know, again, like the unexpected person or obstacle in the road? Is it weather conditions? What are some of these, you know, either most important, most difficult edge cases to test?

12:40Yeah, I think these are definitely amongst the most safety critical cases. Obviously, the tails of the distribution are very long and it's impossible to capture all of them. But I think the hope, you know, for us, for Voxel 51 is that we can capture enough of these events in your model that your model understands what to do, even if it's not seen the exact scenario before. Right. As humans, you know, we learn that whether it's a dog running in front of us or a cat or a human, we know that we should stop and kind of create that same learning in your models. Again, with like the high volume of data that's being collected by a lot of different companies, right, they're collecting a lot of nominal cases of you driving smoothly along the highway where nothing's happening.

13:19And that data is not necessarily incrementally useful to your model. And so at Fortelix, definitely one of our key offerings is being able to create a huge variety of edge cases just from one relatively nominal log and allowing you to understand how your stack reacts to each different event. Right. And just to add to that, I'll say it in a different way, right? Where if I gave you, Noah, a billion dollars in a year and a half and said, build me a self-driving car, right? I am highly confident that a person of, you know, even minor technical, not to undermine the great work that people are doing, but even someone with minor technical ability can build a self-driving car that drives correctly 90 % of the time.

14:01There's just no way in the world you would ever release a self-driving car that only drives safely 90 % of the time. Of course, this is how we think about it. Anyone that works with machine learning models know how difficult this is. Closing the gap from 95 % to then 97 % to 99%. And now we're in this 99.9 % of the time kind of region, right? And still not being able, it's not safe enough, right? We need to make sure it's safe enough. That's where the real problem lies, right? Yeah. Each of those little jumps is like an order of magnitude, more effort, more time, more data. You're the last mile problem.

14:37No, for sure. When you're working so deeply in simulated environments, how do you measure, and maybe there's a better word than this, but how do you measure the realism of scenarios, of objects of behaviors within these simulations? And how can you come up with benchmarks, other ways to measure things that will translate into safer outcomes, you know, out on the road? I want you to go because I have a really big devil's advocate answer to this question. Okay. Oh, it's a teaser. I like it. I can go first. Yeah, for us, realism falls broadly into two categories, right? And it depends on what you're trying to do.

15:12If you're trying to train your stack or if you're trying to test your stack's performance. In each case, you know, obviously better realism leads to better AV stacks and safer driving behavior. For training, we want the synthetic data that we are generating, right? Whether it's through neural reconstruction or diffusion-based, you know, generation like Cosmos, we want that data to increase the model performance of your end-to-end driver stack, right? And there's, you know, whether that's making sure that you're following the rules, minimizing disengagements or collisions, essentially. We want training data to improve the model performance.

15:48And we can kind of use intermediate like task-based perception models as proxies to make sure that this is something that's happening. That's typically what you would do with like data generated from Cosmos for Spark. For testing, we want to see the same performance in the reconstruction that we saw in the real world. That means we want our stack to fail in the same way, right? Whether that's predicting false positives or false negatives to understand or ensure that our simulation is behaving exactly like the real world. It's not exactly applicable to end-to-end models, but again, using task-based perception models can be a good proxy here.

16:25It gives us confidence that adding reconstructed snippets back into the training set with labels will improve the driving behavior. When you're able to see new scenarios, add those to your training data set, and hopefully learn not just, again, like I said, that a person is running in front of a car, but have the same reaction if it's a dog or a cat or something else unexpected. Yeah. To add to that, right, I'll give the marketing answer first, right, which is very important first, which is we do the same thing very similar, right? And Rohan and I, our companies serve different purposes inside of these orgs at work, right?

16:56So Rohan's very focused on safety and make sure it fails in the same way. These are all absolutely crucial. It's not to say that I don't plus one everything you just said. But in BoxL2D1, we focus a lot on how the machine learning model interprets the context of this, right? Does the model think it looks like the real one? And there's many different metrics and stuff you can do, such as embedding models. You can extract embeddings from your model that is driving your car. So whenever we do a scene and then we reconstruct it synthetically or I make it rainy or if I make it snowy or something like that, does the model still think it's driving in the same neighborhood?

17:29Because in the past, we see this with not the case. It looks like a butterfly to me. The model does not think it looks like it. It looks like it's completely different. That's exactly what we're trying to avoid. There's also a lot of classical kind of photogametry, videogametry methods you can apply here, which is like, hey, this doesn't really look like a normal video to us, right? Like the pixels are in the wrong place. It's not really moving in the right directions. It's a lot of generative AI noise, right? That we don't want to have. That's throwing our car off and such. Now, the devil's advocate answer here is I don't think it actually matters as much as you think it does.

18:00How so? The goal of training a self-driving car is to make sure it doesn't crash, right? Right? So I actually don't care if it doesn't look exactly like the real world, as long as my car gets better at driving. Right? And this goes back to my point earlier about time and things like that, right? Where if you can make me a synthetic scene that looks 90 % of the way like the real world, and my car can drive it in, it learns from it, it learns from its mistakes. And that we see it's reproduced in the testing scene that Rohan is talking about. And that's great. Right? Now, if you spend 10 times as long to get to 100%, and now I have all my RTX shaders on, I have all my ray casting, my shadows are perfect, I look at the puddle, it reflects my reflection, all these kind of things.

18:48Studies have shown that it doesn't matter. That the car doesn't care if it sees its shadow or if it sees a puddle and things like that. The important part here is that it's learning how to drive better. And so there's always a trade-off here in terms of we always want to be as real as possible, of course. We're getting closer and closer to that goal every day. But there is this acceptable little zone here where as long as you're being efficient in this process, your car is driving safer, and you're always maintaining that real-world data inside of your testing and validation loops, then some level of not-realness is acceptable.

19:22Right. That makes sense. Yeah, I completely understand that perspective as well. A little bit different of machine learning engineers, the data side of things. But yeah, there have been studies and papers published that show mixing real data and synthetic data, even if that synthetic data isn't 100 % realistic, still leads to improved model performance, at least on traditional task-based perception models. So it's very reasonable to think that that should apply to end-to-end models as well. And to reemphasize, the digital twins, Gaussian splatting, the on-verse new rack, is still way better than what we were a year ago.

19:55It's way more realistic and everything like that. That's not to say that we're personally thinking about it or anything like that, right? Of course, we're getting so much better at it every day. I'm speaking with Rohan Bossin and Dan Gorel. Rohan is Senior Solutions Engineer for Sensor Simulation at Fortelix, and Dan is the Head of Technical Partnerships at Voxel51. Rohan, I want to ask you about a Fortelix technology or an early step in the pipeline that you guys do called Scenario-Driven Data Curation. Can you unpack that, talk a little bit about what it is and how it differs from traditional ways that AV teams pick drive logs for simulation and for retraining?

20:37Yeah, absolutely. Scenario-driven data curation is what we call our automatic scenario labeling tool. It allows us to automatically label temporal events like approaching a stop sign or turning right on a junction when pedestrians are crossing in logs that we ingest and we add them to our database. It allows us to search for individual events or a combination of overlapping events to identify scenarios of interest. We also have dashboards and metrics to see what data you've collected for your stack and what data is missing and how, you know, my role is to help our customers fill those gaps. Right.

21:12And then I understand you guys also have a product called Fortify that helps when customers are scaling up their programs to big fleets and using larger models. Talk a little bit about what Fortify does. Yeah. Fortify is, I guess, like our overall tool chain. Okay. And this is a step within that. I think in the way that it's different from, as you were asking before, like the way it's different from kind of traditional ways, AV teams pick drive logs, that this is like an automatic process that we have when we ingest logs, and you don't have to go searching within your logs for certain events. Oh, great.

21:47Okay. As Dan said, you know, you've got hundreds of petabytes of data. Digging through that is difficult. Yeah. process. And our scenario matching tool helps kind of label those events and let you search within each category to understand what data you have in your ODD and what's missing. Fantastic. Dan, from Voxel's side of the house, how do your neural reconstruction and data curation tools change the way that AV teams think and go about operating? Yeah, right. And if I haven't heard it or said it already, it's all about efficiency here, We're trying to be as efficient as possible, as well as holding as much context as possible for the right job.

22:27So Fortelix, for instance, as we just heard, takes a really interesting approach, a really great approach about scenarios and scenes. And can you label and find the temporal aspects such as left turns and U-turns and pedestrian crossings and things like that? For us, it starts earlier in the pipeline from data curation based off the videos or the 3D assets there to help understand one of two things. One, is this the right thing I'm looking for? I'm really looking to see where is the model uncomfortable. And I link that kind of ambiguous because it is an ambiguous thing where we don't always want to make and project what we think the model needs to learn.

23:04We need to hear and listen to our model to understand where we're going to learn. right so using different methods of embeddings and curation understanding model performance and evaluation help us understand exactly kind of where the model feels uncomfortable so that way we can make sure that we're only reconstructing the training on the seams that are going to give us the most value right that we can then pass with tool like foretellics and say hey this is the subset of that petabytes of data that we think are the most interesting let's label these first and then we'll move back down the chain and we'll you know gradually grow our data from where we go from there.

23:37Additionally, it's all about considering how accurate your translation is from physical to physical, right? We can always get better, right? There's always going to be microsecond differences, centimeters off. There's always going to be additional machine learning context you can add to it. Similar to how Rohan talked about those perception models in the middle to evaluate kind of the performance of your autonomous driving stack, you can actually do this before training as well to understand things like if I'm really interested in traffic lights that are turning yellow, for instance, right? Well, I can use a machine learning model that finds specifically yellow traffic lights and then only grab those seats to see, hey, did my cargo through the intersection?

24:16Should it have stopped? How close was I to the intersection, right? This is a very interesting scene from an AV scenario perspective and something that we can do inside of the box of the tool to make sure that you find the best suits what you're looking for. Right, right. Amazing. From the engineer's point of view, how much do things like interactive exploration and visualization of data sets, how much do those help in helping the engineers trust these reconstructed 3D worlds as they use them to simulate things? Yeah, it's a great question, right? I think visualizations are always going to be important, right?

Read the full transcript

24:52A lot of computer vision, and I'm very thankful we've gotten to this point because it wasn't always like this in the beginning. a lot of computer vision starts with your eyes, right? We need to look at the data. We need to understand data. We need to understand that like, hey, whatever I'm about to go train my model on, I've actually looked and checked it and understand, right? And have that human element to it, right? And so as the field evolves, these visualizations evolve with it. Rohan talked about all the Fortellix dashboards, understanding things like evaluation metrics, such as where do my sim scenes seem the most realistic?

25:22Where are they drifting the farthest from real scenes? Do I have enough coverage? for instance, how many of my scenes are in parking lots versus highways versus cities. These are all visualizations. You can't go and scroll through all of your scenes to find this information. This needs to be kind of brought to you in a way where you can easily kind of understand not just the fact that I don't have enough parking lot data, but oh, I can also see the correlation of like, look, my model is actually a lot worse than parking lots as well. Right. And so now this makes sense, right? Where I have a gap in my data, I see the gap in my data, and that's really the name of the game for an engineer here, right?

25:55I'm going to obviously use the Voxel 51 Fortelix example here, but I get my data into Voxel 51. It's all translated great. I passed all my checks. I'm ready to go reconstruct. I pass the reconstruction to Fortelix. Fortelix labels it. They scan it. Now, so we did this over 100 scenes. It's going to be able to tell me where those gaps are. Now that I know where the gaps are, I can go back to a tool like Voxel 51 and find if I have that data somewhere in my data lake of data or I can use generational methods like Cosmos and things like that to generate those methods. It's always going to be a cycle.

26:27It's always going to move it continuously in a circle. But a lot of that is driven through today. Visualizations, dashboards, understand exactly the right correlation to kind of find those causations, right? Yeah. As Dan said, you know, I think, yeah, that's where Fortalex and Voxel work really well together, right? We identify the scenarios of interest, but, you know, at a more kind of deeper level, identifying individual events, like Dan gave a good example of traffic lights that are turning yellow, identifying those scenarios and, you know, making sure that they look realistic. You can understand how your data looks, whether it's from a realism lens or how your model is interpreting it.

27:03Both of those are super relevant and kind of point towards a partnership that works well for both of us. As these stacks, the AV stacks, use larger and larger models, world models and perception models, How do you ensure that, you know, as you go through training and model training and then evaluation and then working in simulations, how do you make sure that all the different components can understand each other? And sort of is there a commonality amongst the data? Is there kind of a common data language that need to get things talking? How do you sort of manage that level of understanding, you know, throughout the stack?

27:40Definitely an interesting question. It's something that we're still trying to understand the best answer for today, right? There are some things that we can definitely answer today. One is that there is no common data language. This is an issue. I don't think it's an issue that is easily fixable as well. Now, Voxel 51, this is one of our biggest physical AI workbench. This is what we're pushing out here to understand, hey, I'm not saying that I'm going to push out this new data language. Everyone get on. That's not practical. People have data. It exists already. It's in their own format. That's not going to happen.

28:11right? So it's almost better to confirm not the language you're speaking, but that I'm translating it correctly, right? So it's going to come in all these different forms. There's open source formats, but a lot of people tend to actually just have their own formats internally. Let's just ensure that whatever we're communicating to that world model, that foundational model is being said correctly, right? And that is the important part here. In terms of how we evaluate them afterwards, words, it's tricky, right? This is something we're learning, right? I would make the argument that we're a little early today, right?

28:44But before we were evaluating so many different modules in the stack, right? We talked about reception models, reception models go to tracking, tracking goes to prediction, prediction goes to planning, and planning goes to control. That's like six different modules that we would evaluate. And I would know how each one performs, but I didn't really understand the cascading effects and waterfall effects of errors, right? With world models, we can skip a lot of steps, right? A lot of times it's end to end. It's input, output, how well did you do, right? So I don't have to ask as many questions. I can just look at one model and say, are you doing good or are you doing bad?

29:16So the problem gets a little bit easier, but the question of are you doing bad and why is still a little bit of an open question today on how we solve this. So you guys are building on NVIDIA platforms and as you've talked about at length, using neural reconstruction technology. What's possible now building at this scale on this technology that just wasn't feasible before when your teams tried to do it on smaller, maybe bespoke stacks. What does the current scale allow you to do? I think for us, the current scale actually is how many variations we're able to render quickly on different scenarios.

29:57As Dan said, time is the key currency here. And with the level of compute that's available now, we're definitely able to generate a huge variety of scenarios, render them quickly, and then also pass that down into our training set, evaluate our stack faster, and it's sped up the iteration of each model training and testing cycle that's hopefully going to lead to safer autonomy faster for everyone. Yeah, and just to add to that, I think there's two interesting angles. One, from the business ROI standpoint, there's so much data that these companies and teams have collected over the last five ten years even and have just been sitting around like on a drive more or less useless like you know as much as you think that you know these cars are driving around and we're training on those constantly like that's typically not the case right like a lot of times it's mostly useless right however narrow reconstructions have made this useful all of a sudden right because i can go back and take all these logs and say like hey, well, my car didn't crash in this case, but it almost did.

30:59So what if it did? And I can take this old data, doesn't matter how old it is, doesn't matter if it's sitting on a hard drive for years, and now start to generate these augmentations or variations that are really interesting. From a pure business ROI, thank goodness all this data we collected is not useless, right? That we can actually do stuff, and this is going to solve a lot of the problems, we hope. From the other perspective, and this is really focused more around the Omniverse New Rec and kind of technical advancements here. We talk about Gaussian splatting. If you don't know how Gaussian splatting is, I'm going to give you a really easy, simple, over-the-top basic example of what it is.

31:36Go for it. Is you're basically painting a 3D picture, right? You basically have your paintbrush, and I'm painting the picture, but it's 3D, right? So I can decide how far away my strokes are. But that's all it was for a long time, right? And so when we think about driving cars, well, like cars move, right? So I need to be able to move things in the painting, but I have to be able to then communicate that in a very complex way to simulation environments. I won't go on to all that detail. The idea is now with Oni vs. New Rec and some of the other available models that are coming out there, I can now kind of specify, hey, this truck that I'm painting right here, it's moving.

32:14It's going to continue to go down the scene. I might want to pick up that truck and move it somewhere else, right? Whereas all these other points out here that I've proven, all these brush strokes for keeping the painting analogy, they're static. They're not going to move anywhere. Don't move them as the truck moves. And that small advancement of being able to specify, hey, this is a moving object versus, hey, this is the background, has been probably the biggest step change that we've seen in the last six to eight months. Yeah, yeah, yeah. Fantastic. So, Dan, I want to ask you about another Voxel 51 innovation product thing you've announced recently.

32:47Something called the Physical AI Data Engine. This is also powered by NVIDIA Tech. Can you explain what it is and how it changes day-to-day workflows for, you know, AV teams who, as you've been talking about, are just working with, you know, petabytes of data, trying to wrangle all of this data at once? Right, absolutely. So something that we're very proud of that recently came out, our Physical AI Data Engine, it helps to solve a lot of problems that we've been seeing. And this is something that we created not because, you know, we thought it was a good idea. this was something that we created because we actually went out and talked to these AV companies in person and began to hear the same pain points over and over and over again.

33:26So it was just a matter of time of like, hey, we just need to build this. If we want to move on to whatever that next stage of self-driving is going to look like, this is something that just needs to exist. So we might as well just build it. And it's important to remember that AV teams, even today, it's not a bunch of people sitting at one table kind of doing their work. When we think about AV teams, you have to keep in mind that there's a self-parking team, that there's a L2 team, an L3 team, there's a fully autonomous team, right? And these are all people who work in different regions of the world.

33:55They might be remote. They might work on different cars. There's many different factors there, right? These people don't always, and most often than not, don't talk to each other, right? Including this, the data team doesn't talk to them either. So a lot of this has just kind of been throwing things over the wall back and forth to each other for the longest time, right? And so this causes a lot of issues, right? The first issue is that you're not getting key advancements both internally at your company as well as externally from developments outside in the field, right? Your data pipeline has become so Frankenstein for your specific use case that, unlike the self-parking team, for instance, I actually only care about the backup camera.

34:35We're only backing up into spots or something like that, right? I don't need the side cameras or something along those lines. And you can extrapolate this out to all the teams, right? So now we can put them into the physical AI data engine of saying, hey, you know, the data team, the ML team, when the data comes off the car, we need to make sure that everyone has access to this data in the same format, in the same workspace. That way, if you have this key innovation, you have this great new model that parks the car the best or drives the car the best or predicts how far away a car is the best, right?

35:04Whatever it might be, that that can be easily shared across all your teams, right? In a way that you don't have to rewrite every single code. You're not throwing files or Python files over the wall, hoping for the best, right? Now you have a direct streamlined access all the way from when that data comes to the car, all the way till it gets to that ML engineer and everyone can see exactly every step of the way when this happened. And the way this is done actually in practice is through two main steps that we do today, with a third step being the NVIDIA tools that we're building on today. The first step is when a data gets ingested, we want to maintain that physical, the digital translation.

35:41The only way to ensure this agnostically in a way that every company in the world can make sure that they're on the same level playing field and translating it the same way as possible is through something we're calling basically the physical AI audit, right? Okay. Where I'm going to check all the different translations. I have a hundred different tests. I'm going to run over all of your data set to ensure that all your cameras are aligned, that your LIDAR is set up correctly, all your timestamps are corrected, that the car not only thinks it's in the right spot, but it sees it in the right spot.

36:10Every single sensor of the car is working in unison. I would say that probably more than 50 % of the time, large companies that everyone would recognize fail this test, right? And it's just natural. It just happens. Data's messy. Sometimes it's old, as we've been mentioning. And so it's important to know where I need to fill in my gaps in order to get to this latest and greatest world models like Cosmos and such. Because previously, I only cared about the image or the video. I didn't care about everything. Then after I spot the error, well, it's not good enough to just spot the error. I got to be able to help fix the error for you, of course.

36:44So we then head into our next area, which we call enriching. This is a mixture of photogammetry, videogammetry, driving techniques, ML models that will fill in those gaps for you. It's a way to think about it. It's like a nine out of 10 problem. If you give me nine of the things, I can find the 10 out for you, and I'm pretty confident in this case. What we've actually discovered in development is that number doesn't have to be nine. It could be more like four or five or something like that, and I'll fill in the rest. And this actually leads you to get way better data sets than you would if you just got the data off the car and passed it into real reconstructions.

37:16We've done comparisons with certain customers and data sets and such, where if you just pass the raw data set into like Omni versus New Rec, versus if you pass it through audit, then enrich, make it better than it was when it came off the car, then go do Omni versus New Rec. You're actually going to get a better reconstruction the second time because you've added so much more color to it. And then after you've already audited your data set, it's in its best form and possibly be before you go speak and talk to it, work with those world models. That's when you then can work with those tools like internal world models, Cosmos, Omniversity Rec, and any other one, ultimately generating high quality reconstructions or variations, augmentations, which then you can then pass into great tools like Fortelix to then do all these kind of scenario-based analysis.

38:02Right. Yeah, fantastic. We always like to wrap on kind of a forward-looking note, but in this case, it's going to be kind of a very specific question kind of following everything we've been talking about. But if everything you guys have been talking about, this kind of reconstruction, data-centric, reconstruction-driven simulation becomes the standard, how do you think it'll change the way teams operate, team structures, and the skills that individuals need to have inside AV companies over the next, let's say, five years? How do you think this kind of simulation is going to change the game? I think five years is very difficult to predict because as Dan said, five years ago...

38:41I hope I have a job in five years. I know, do you want to go 18 months? Yeah, I thought you were going to say 18 years. I was going to have a heart attack. I know, do you want to go 18 months? Five years ago, we were playing GTA. Now it's a completely different game. Right, GTA, yeah. I think in the very short term, it's going to make life a lot easier for especially verification and validation teams. The neural reconstruction quality is so much better than any photorealistic effort from physics-based rendering solutions. It's going to make trust in simulation a lot easier. That means VNV teams can sign off on new features or new end-to-end driver models without as much extensive real-world testing and reducing costs associated with that.

39:23Again, time is the main currency here. the faster you can develop new models, the faster you can get them out the door, the more data you can collect with those models and iterate on them, improve them. So I think that's going to be the number one benefit from simulation here. Yeah, I agree with it too. I think teams will get a lot more flat, right? As we things like, you know, I'm not overgeneralizing again, it's a podcast, I'm allowed to do that. A lot of things are just going to become profitable, right? I can just ask a model to do something for me and it will generate this reconstruction for me.

39:53right right that's coming down the line it's very obvious so we're trying to remove a lot of these quite outdated skills i want to say skills i'm not trying to say jobs i'm hoping that you know those things are not necessarily going hand in hand where people are just going to be able to do more with the resources they have they're going to be able to do more and fill in more lines in the gap and we're going to remove some of this telephone game you've been playing today comes off the car there's a team for that comes off the car those things they have said that's another team. Then it gets simulated, that's another team.

40:22There's a reconstruction team, there's a synthetic tape team, there's a safety team, there's so many different teams. We can flatten these structures and kind of roll them into one cohesive team that is working together. That sounds like it's much more beneficial for everyone involved. And taking the other side of your question, I think that's a lot of short-term looking forward. For the long-term looking forward, I'm really, really interested in especially the CES announcement from NVIDIA around AlpaSyn and Alpameo. I think is really, really interesting where now we're thinking about even doing simulation inside of world models and using world models to generate these variations for us.

40:58I think it's just really intuitive. I think we are a little bit grounded in the GTA world even still today, right? Where some simulation experts will not say neural reconstructions are good enough unless they can drive the car in the engine themselves and see and feel it. I argue why do you care? right like you're not driving the self-driving car that's the whole point right and so i think uh in some ways we'll just lean on world models even more such as alpa sim i think it's really incredible in the opportunities where the model should know where it performs bad so then go find these simulations yourself and then generate these traffic patterns whatever it is and do all these variations without me having to go in and manually move the cars around and stuff myself right that's probably more like two to three years down the line if we're speaking honestly about when that's probably going to land but i think it's a no-brainer that we're heading in that direction and it's really exciting to see how the field is going to evolve i mean this is less than a month older we're only a month away from cus at this point of recording uh right so it's all we're all still learning right now it's very exciting it's all brand new yeah amazing dan rohan for listeners who want more who want to learn more about voxel 51 for telex want to learn more about just the state of AV simulation and everything we've been talking about.

42:15Best places for listeners to go online. Company websites, social media handles. Where should they take a look? Dan? Yeah, two main places I can point you to. First and foremost, check out our Voxel 51 docs. 51, the main platform that we've been talking about today, is actually open source, right? So a lot of this you can use yourself. We have great tutorials and guides on how to get started with self-driving examples, including some around causing splatting and things like that. So I highly encourage if you've been interested about the topics that we've been talking today about go head over to our docs that's voxel51.com and you're going to be able to easily find our docs and all of our information about this client there.

42:55Additionally, my name's going to be attached to this podcast, Daniel Grau. I appear everywhere exactly the same. So if you ever have any follow-up questions if you're interested in chatting more or just interested in connecting I recommend you reaching out to me on LinkedIn. that's where I'm most available. And I'm always happy to chat about physical AI. It's on my mind all the time. So I'm always happy to share some of these thoughts. Awesome. Great. Rohan? Yeah, definitely reach out to me on LinkedIn. I'm happy to answer anything sensor simulation related. And I can always point you to the right person at Fortelix if you have other questions.

43:27Fortelix.com also has great resources for what we can provide, what services we provide. And yeah, if there's anything we can do to help, definitely let us know. We can also check out our announcements at CES. and see if there's anything that interests you there. Fantastic. Well, guys, thank you again for taking the time to come talk about AV simulation. And just, it's all moving so fast. I say that every episode, but it's never not true. It wouldn't be fun if it wasn't, you know? Exactly. Well said. Well, appreciate the conversation and best of luck to you both on all the work you're doing. And, you know, look forward to recording the next one from the backseat of an autonomous vehicle as it guides us down the road.

44:09Yeah, absolutely. Yeah, you can do podcasts, Getting Coffee in Cars. There you go. Right, exactly. Right, exactly. Hope to see everyone at GTC as well. I'm sure I think both Nia and Rohan will be there as well. So if you're being in person there, feel free to stop by. We'll be in San Jose. Excited to do so. Yeah, look forward to seeing everyone at GTC as well. Great to be here. Thank you.

44:43Thank you.

45:10Thank you.

From the publisher

How can AV teams stop drowning in petabytes of data and actually ship safer autonomy faster? Fortellix’s Rohan Bhasin and Voxel51’s Dan Gural explain how neural reconstruction, scenario-driven data curation, and NVIDIA-powered pipelines turn ordinary drive logs into high-fidelity simulations that close the last-mile gap in AV performance.

GTC is the premier global AI conference. Learn more at nvidia.com/gtc

More from NVIDIA AI Podcast

All 115 episodes
Driving Safer AVs Faster with Smart Simulation, Neural Reconstruction, and Data-Centric Tools - Ep. 289NVIDIA AI Podcast · 45 min
Listen in VO