In short
NVIDIA AI Podcast Episode Summary
Episode Title
AI, Spatial Intelligence, and 3D Content Creation With Sanja Fidler of NVIDIA - Ep. 269
Host
Noah Kravitz
Guest
Sanja Fidler, VP of AI Research at NVIDIA
Episode Description In this episode, Sanja Fidler shares her journey in AI and computer vision, leading the Spatial Intelligence Lab in Toronto. She discusses the lab's research on spatial intelligence, advancements in simulation, 3D modeling, and their implications for content creation and robotics.
---
Key Themes and Topics Discussed
- Background and Journey of Sanja Fidler
- Early Influences: Sanja traces her interest in science and AI back to her childhood, influenced by her father’s storytelling about scientists and inventors like Nikola Tesla.
- Education and Career Path: Overcoming challenges in her schooling, particularly in mathematics, with support from her mother. She pursued a PhD focused on computer vision and later joined the University of Toronto.
- Spatial Intelligence Lab
- Founded: 2018, under Sanja's leadership.
- Definition: Spatial intelligence refers to AI's ability to understand and operate in 3D environments.
- Objectives: To develop AI systems that can perceive, model, and interact within 3D spaces, paralleling language models used in natural language processing.
- Importance of Simulation
- Robotic Applications: Robots must understand and navigate the physical world, which is three-dimensional and governed by physics.
- Physical AI: A concept where AI operates in the real world, crucial for robotics and autonomous systems. Sanja posits that physical AI will surpass generative AI in impact.
- Advances in Content Creation
- 3D Content Generation: The lab focuses on making 3D content creation accessible and efficient, enhancing creative workflows and enabling non-experts to participate in design.
- Innovations Presented at SIGGRAPH: Insights into the lab's work on differentiable rendering, generative models, and integrating physics into 3D simulations.
- Future Perspectives
- Integration of Technologies: Combining traditional simulations with emerging AI models to create more accurate and realistic simulations.
- Challenges Ahead: Addressing complexities in achieving true physical accuracy in simulations and the role of visual language models in enhancing physical AI's capabilities.
- Collaboration and Community Engagement
- Interdisciplinary Approach: Emphasizes the importance of collaboration across various industries to leverage shared knowledge and expertise in AI and robotics.
- Call to Action: Inviting researchers and enthusiasts to join efforts in advancing spatial intelligence technologies.
---
Key Takeaways
- Passion and Perseverance: Sanja highlights the significance of having a passion for research and the perseverance to overcome challenges in scientific exploration.
- Democratization of 3D Creation: The evolution of technology allows more creators to engage in 3D design, broadening the field and enhancing creativity.
- Future of Robotics: There is optimism about the development and deployment of autonomous robots that assist in everyday life, fulfilling a childhood dream of Sanja’s.
---
Conclusion The conversation with Sanja Fidler provides valuable insights into the intersection of AI, spatial intelligence, and robotics, revealing the transformative potential of these technologies in shaping future interactions with our environment.
For listeners wanting to delve deeper into the work being done at NVIDIA and the Spatial Intelligence Lab, resources can be found on the [NVIDIA website](https://www.nvidia.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to the NVIDIA AI Podcast. I'm your host, Noah Kravitz. This past August, three of NVIDIA's research leaders gave a special address at SIGGRAPH, the annual International Computer Graphics and Interactive Techniques Conference that's been running since 1974. One of those people is here with us today. Sonia Fidler is VP of AI Research at NVIDIA, where she leads the NVIDIA Spatial Intelligence Lab in Toronto, Ontario, Canada. Sonia is here to tell us about the lab, to talk about the research she's most excited about right now, including what was presented at SIGGRAPH, and to share a little bit about her own journey through the worlds of research and artificial intelligence.
0:49So without further ado, let's get to it. Sonia Fidler, welcome, and thanks for joining the AI podcast. Hi, Noah, and hi, audience. I'm very excited to be on this AI podcast. We are very excited to have you. Thanks for taking the time. There's a lot going on, obviously. And congratulations on the special address and everything else at SIGGRAPH. So we wanted to start with a little bit about your own journey. You followed your passion for computer vision and artificial intelligence across Europe and into North America. Can you tell us a little bit about what first got you interested in the field and how your journey took you to Toronto?
1:29So maybe I started with my youth. And there were actually three important breakpoints that led me to where I am. And the first one actually starts with my dad. So my dad would sit on a chair next to my sister and me and tell us bedtime stories. And surprisingly, she was very good at it. He was a scientist. And so he would tell us stories about scientists. For example, he would tell us about Nikola Tesla, who was born in Croatia. And my mom was also born in Croatia. I was born in Slovenia. And was this in Slovenia? Where did you grow up? Yeah, I grew up in Slovenia. So my dad was born in Slovenia.
2:08My mom was born in Croatia. My dad was born in Slovenia and I was born in Slovenia. Okay. So he would tell us stories about, you know, how at a young age, Nikola jumped from the roof of their house, holding an open umbrella, thinking he would fly, you know. And every night there would be a new episode about his inventions, like the creation of radio, alternating current. Obviously, we didn't understand what he made it sound. And the competition with Thomas Edison, right? It would be almost like a Netflix series. Yes. For a child, this was very exciting. So I just could not wait to hear more until next day.
2:46So kind of like my childhood, heroes were not movie stars or music stars. They were scientists. Awesome. That really kind of shaped me. So perhaps not surprising, you know, one day I appear in front of my parents and proclaim, I want to be an inventor. And there was even a photo of me and my sister, actually my sister dressed as a robot. I paid quite a bit of money for her to put some cardboard boxes around her. And maybe not surprisingly, she became an economist. So this moment pretty much settled my profession. I was going to be a mentor and this was very young age. The second moment was really thanks to my mom.
3:29And this was in primary school. I was very young and at some point I got pretty ill. It was something like COVID almost. I think it was called looping cough or something. So I was home, fever, coughing two to three months. And I basically missed a lot of school. I missed like, you know, fractions. It was a whole big chapter on math. I had no idea about it. So I come back to school and of course I didn't understand anything they were talking about. And I developed some sort of resistance in going to school. And before a math test, I threw a tent from crying hysterically on the floor. I hate math.
4:10I don't want to go back to school. And my mom is actually a teacher. And of course, having a school-hating child, that was not an option. Yeah, it happens. So even though she was an English teacher, she would sit down with me and work with me on the math. And she made it really interesting. So she had this really nice way of teaching me through, giving me puzzles, math puzzles. And I began to like really love it. You know, as kids understand things, they also love it. And I think to this day, what drives me at the core is solving problems. And I think that's still kind of stuck with me. It's pretty much settled, you know, what I would study.
4:53I was determined age, I don't know, 12, 13, I was going to study math. and the third moment was really, you know, kind of thanks to my grandma. So I was already doing my PhD and I decided to work on computer vision. I saw this one talk. It started actually with math, even my PhD. And I saw this talk on someone recognizing cats and dogs and I was very early AI at that point. And it just kind of spoke to me, you know. I was always kind of dreaming of robots and computer vision felt like the first step to do. So, you know, I was there doing my PhD and my grandma, you know, she was a very smart woman.
5:37She was actually one of the first female plastic surgeons in Yugoslavia. Oh, wow. Yeah, she was always telling me these stories, you know, how she graduated in med school and the day they were graduating and they were out having fun and, you know, sirens came out, World War II started And, you know, she had to basically just go to the operating room and that was her next four years. And, you know, like basically fear became alien to her and wasn't for me really. Right. So I studied my PhD in Slovenia really for the fear of leaving, leaving right to the wide open world alone as a woman. And I was somehow not encouraged.
6:19My mom would scare the hell out of me. So towards the end of my PhD, and I was kind of working on this AI, like something similar to deep networks, just my own take on it. I was presenting at a conference and a famous professor at UC Berkeley stopped by the poster, really likes it and invites me to visit his group at Berkeley. And, you know, I was beyond excited, but I still carried this kind of weight of fear and ambition. and I talked to my grandma and she said, you know, Sonia, don't listen to your mom. Just go. Actually, she passed away a few months later. That was January 13, 2009. And the next thing I remember, I'm sitting on a plane.
7:03I look at my plane ticket to California and it was January 13, 2010. It was exactly one year later. Exactly a year. Exactly. I'm not kidding. This was exactly... It was meant to be. Meant to be. I was both scared and excited, but, you know, chapter two of my life was a lot of strange. Was that your first time traveling abroad? No, I would go before, you know, just visit New York with my family. This was the first time I went alone and, like, leaving. Yeah, very different. You know, I got my bags and here it was, you know. Yeah. It was scary, but. And landed in Berkeley, of all places. I spent there a few months, seven, eight months, and came back, graduated, and then I did my postdoc and I was at U of T.
7:51So that's kind of what brought me to Toronto. Right. Amazing. I feel like the graphics and interactive industry owes a big thank you to many members of your family for all the inspiring you and then grandma kind of giving you that nudge and everything. That's amazing. Why Toronto? What was the link that brought you to Toronto? Yeah, actually, the U of T University of Toronto was doing really great stuff in deep learning. And like I said before, that was kind of my PhD. I was really inspired by doing this hierarchical representation to recognize objects. I was reading all this, like, in a neuroscience paper that basically said this is how the brain works, right?
8:35Yeah. And I was in Slovenia. I was isolated. I kind of had my own take on how that would look like. And then I was reading this deep learning paper. it was papers was really appealing to me and was kind of you know like going between berkeley and u of t and i decided to go to u of t to kind of like learn learn from you know jeff hinton and people like that and that's why i landed here yeah amazing choice and so you've been in toronto since um yeah i mean i was there for a postdoc and then i got um like a research assistant professorship in Chicago. So I did a small stop there for a year and a half.
9:14Then this position, faculty position, opened at U of T and then came back 2014. Amazing. And so now you head up the NVIDIA Spatial Intelligence Lab in Toronto. That's right. Yeah, yeah. I joined NVIDIA. That was 2018. So about seven years ago. Seven years ago. I actually met Jensen at a computer vision conference. That was 2017. and we had a really great chat about simulation. I was already working on simulation for robotics at the time and I was telling him about it and I think he was also thinking about it. So it was a great conversation. And then later he gave me a call or we went on a call and he said, you know, come work with me.
9:53And I had other options, but the fact that he said, come work with me and not for me, just told me everything about his joining and that was it. What a great story. That's fantastic. So tell us about the lab. For those listening who might not fully get the term, what does spatial intelligence mean? And what's the charter of your team? What are you doing at the lab? And you may have just said, sorry, I was imagining Jensen, you know, that whole conversation. But when was the lab founded? 2018, May, that was basically with me. And then we slowly grew and also increased scope. So we recently renamed ourselves to spatial intelligence.
10:34I would say it's a new encompassing word. So spatial intelligence essentially denotes intelligence in 3D, right? Intelligence in a 3D world. So the same as we have LLMs representing intelligence in language, you have all these family of visual language models for intelligence into the images. now we need to build the same capabilities but in 3D. And the question is, of course, you know, what that is and why. Maybe I'll motivate with robots because really, you know, that's one of the prime motivations. So at the end of the day, like robots need to operate in the physical world, in our world, and this world is three-dimensional and conforms to the laws of physics.
11:18And there's humans inside, right, that we need to interact with. You know, we typically hear the term such AI that operates in a real physical world as physical AI. So I'll maybe use that term quite a lot, right? Physical AI is really kind of the upcoming big industry, very likely larger than generative and agentic AI. You know, Jensen typically says everything that moves, all devices that move will be autonomous, right? So that's kind of the vision. So a robot to operate in the real world, obviously it needs to understand the world. What am I seeing? What is everything I'm seeing, doing? How is it going to react to my action, right?
12:00So understanding. It needs to act. You know, if I want to drive you from A to B, make you dinner, you know, I need to actually like control that robot to make an action. But then there are two other capabilities needed that are perhaps a bit less obvious. So basically, it's 3D virtual world creation and modeling and simulation. And the reason is that robots need to have like a virtual playground that almost perfectly or like we would like it to mimic the real world as faithfully as possible, where basically they can train their skills and also test their skills before we're going to deploy them in the real world.
12:38This is basically the critical thing we need to solve for deployment of robots. Basically, spatial intelligence kind of comprises these kind of four core capabilities, which is modeling, so creation of virtual world, but then also, you know, modeling it, how it evolves in time based on our action, understanding, and action in 3D world. And obviously, applications are more than robots, you know, architecture, construction, gaming, everyone that kind of has 3D data, 3D world data. We first started with this virtual world creation, so content creation. And then we, as we go, because in order to develop the spatial intelligence, you also need physics, which evolves in time and understand.
13:22A year or so ago, maybe less, there's so much has happened with generative AI in particular in the past few years that it's kind of blurs together sometimes when I talk about it. But I remember when video models started coming out, the first, you know, Sora from OpenAI and some of the other ones, and discussion around, well, these video models are actually also physics simulations. You know, we're discovering, we thought we were making a video model, but now we're realizing that, you know, there are properties of physics happening inside of the videos that are output and all of these things. What makes a good physics model?
13:57And when you're talking about modeling things that are going to happen in the future, I've also heard, you know, that described as, well, what an AI model does is really predicting what's going to happen in the future, right? And if it's a video that's output, it's sort of frame by frame. How do you think about the four things you just described relating to one another? And I don't know, maybe you can talk a little bit about the physical AI in particular and how the evolution of how these models came to be, you know, so accurate that we can now use them in simulations. Yeah, so NVIDIA Cosmos and the models you're describing, right?
14:33Sora, VO3 and so on, learn their capabilities from videos. And especially NVIDIA Cosmos is kind of targeting physical AI, which really means that it's doubling down on modeling physics, not necessarily the creative aspects, but physics, capturing how our world works. So it's forming these world simulation capabilities by learning purely with videos, and we specifically target collecting videos that are real-world recordings. You know, there's no human editing involved, and if there's any graphics data, it's actually all physically simulated. Okay. How we're using physics is mainly for benchmarks, actually.
15:13So you want to create, because you have full control, right? I can have two bouncing balls, three bouncing balls with this material, you know, and more complex wall. And there you can really go like, you know, every single test. How good are you at that? How good are you at that? And that's our test. And you kind of hill climb that performance. Right. Yeah, it's an evolution of models, right? So the first world model came out, I think it was Juergen Schmidhuber, right? 2019. It was almost parallel to us. Our scheme like a few months later, where the idea was really kind of like AI replaces the game engine kind of.
15:52You know, AI creates the world. You have the user interaction. Next frame is not human written code. It's the AI. It's generated, yeah. Obviously, that was early on. It was, I forgot exactly what they were using. We were using GAN. Ours were called Game GAN. So we trained it on Pac-Man, you know, so you could actually play Pac-Man on a game. Right, right. Like the frames for AI. We had an episode of the podcast with somebody who created GAN Theft Auto. So like Grand Theft Auto, but being generated. Oh, that was yours. Okay. Yeah, that was our stuff. Cool, that's cool. Yeah, yeah, great. I, forgive me, I don't remember offhand who the guest was, but yep.
16:31That was so cool. We released the code. So, you know, people just got crazy. And it was amazing to see where it went. Yeah. Yeah, we actually also applied it to driving. That was 2021. It was called DriveGAN. You know, some technology, but just a lot of autonomous driving videos. And it almost kind of became a driving simulator. You know, Cosmos really took to new heights. But at the time, it was kind of like imagining how this could be useful for physical applications. So that was all kind of GAN-based with all kind of known limitations. And, you know, in the meantime, diffusion models came out, and it was clear that, you know, that's also the next big leap in video modeling.
17:12And actually, 2023, we kind of partnered up with some of the students that did the latent diffusion that really was kind of a big breakthrough in images because you didn't model pixels anymore, but these kind of latent codes It made it significantly more efficient. So we kind of applied that and extended that to video. And that led to video LDM, which really became, you know, you could see the future by looking at those results. Obviously, it was not Sora yet or, you know, Cosmos, but like we were on to something. And then, you know, the industry actually kind of switched to this latent diffusion architecture.
17:49And then, you know, then it's about scaling. and obviously the architecture changed a little bit behind the scenes and data and so on. And that basically is creating the modern age models. So I understand that your lab has grown recently. Can you talk a little bit about the new areas that the lab's now encompassing and how that kind of furthers the overall goals, the overall charter of the lab? Yeah, yeah. So when I joined, we joined Rev's organization. and Rev was building Omniverse. Omniverse is this, you know, like state-of-the-art simulation platform where robots can be robots, as Jensen says it.
18:30Right. And talking to Rev at the time, he mentioned, you know, there was a huge team working on it. Obviously, they were able to render really fast. You know, they had this real-time ray tracing and so on. So really kind of the key missing piece was content. And mine, this was like 2018, right? I was like baby rhymes for that. And that's how we started. We said, okay, like, how can we actually make this platform workable, especially for physical AI, where it's really about modeling the world, which is messy, diverse, you know, like it's really like challenging. So we started with content and yeah, we developed, you know, a bunch of techniques for that.
19:15And through, you know, through kind of the period of our lab, we became more and more ambitious. And, you know, we realized that the pipeline for physical AI or this 3D spatial intelligence also needs to change because you need to have, you know, better physics algorithms. Physical algorithms interact with each other. Plastic, whether it's water inside, I can put it on fire. And, you know, there is no cheating like in a game where I can kind of stage it. Like, this needs to be all simulated. It's real. It's real. It needs to feel real, right? I can put my finger on it. Bad things happen, right?
19:49Like the robot, if it's training there, it needs to kind of experience it in this way. So, you know, it was clear kind of that we need the next evolution of physics and then can join the team. And also, you know, perception is obviously important. And Laura joined a team and she was, she's very interested in 3D perception, but going towards open world, meaning, you know, like anything, anything in this room, I should be able to recognize it and understand my affordances with it. And then, you know, that can lead to a better action. So we expanded the team basically like by building blocks that we actually need.
20:26You know, building the full stack for spatial intelligence. And you mentioned Omniverse. Your lab has been very involved with the creation of Omniverse. What are some of the innovations, some of the research breakthroughs you mentioned, you know, physics models improving? What are some of the other innovations that really made Omniverse possible and helped to grow into what it is today? Yeah, I think, you know, first of all, Omniverse is created by many teams at NVIDIA, right? Much, much, much larger than any single team. Really kind of the vision of Jensen and Rev. It has a mountain of technology, you know, for real-time tracing power by this DLSS that makes, you know, AI in the loop, AI-powered physics, you know, solvers, like I was saying.
21:10So that's just scratching the surface. And I really can't take credit for any of that. So I can maybe tell you a little bit about what we were thinking when we started with our 3D content creation work. And I would really say that we doubled down on two directions. We both turned out to be very important in the end. And it's really kind of this perseverance through time that created something of value. So the first one was, okay, you know, clearly there's a graphics pipeline. We know everything and how that works. So why don't we lift images and videos to 3D to be fully compatible with existing graphics pipelines?
21:50And we really doubled down on differentiable rendering as this foundational technology, meaning graphics goes from 3D and renders to images. If this is differentiable, meaning kind of like amenable to AI. So this path led to one of the first image-to-3D models that we're called GANWAR, one of the first generative models of 3D assets, GAS3D. And as the latest achievement, we also made foundational improvements for 3D Gaussian splats. I don't know whether I need to explain that in further detail, but essentially it's a really, you know, like a new neural graphics primitive that you can easily optimize from videos.
22:30And we added retracing capabilities to it. And at Seagraph, we actually announced integration of, we call it 3D, G-R-U-T, 3D Groot, Omniverse. So basically now you can download Omniverse or Isaac, which basically helps you train robots. You can scan, you know, with your phone or whatnot, this environment, and boom, you have it in Isaac. And you can start, you know, training robots just here. Like there is no, you know, you don't take weeks for it. It's amazing. It all makes sense in terms of, you know, looking at the way you're describing the way things have built up and building blocks and adding features.
23:05And, oh, cool, that makes sense. And then I sort of listened to you describe, like, oh, take your phone, wave it around the room, and now the robot can train in the room. And it's still, it's just so exciting. It's so mind-blowing. It's very cool. Yeah, yeah. It's exciting. But that's basically what you want, right? Like scale. I want to just go and take what's only here. And then same, right? And boom, the robot is training. So the second one, the second path is we kind of saw the fundamental, some fundamental limitations of this graphics pipeline because, you know, you need to also model agents and physics.
23:39Like it all kind of, you know, also felt daunting. So we also made this bold approach of AI that is basically the world model, right? That does the whole content creation, world simulation based on user interaction and all it's one. And that was the chain of models that you described earlier, right? So like two different things that all now kind of like came together in like really, I think, useful capabilities. Yeah. So how has the advent of AI and 3D content creation and sort of specifically in workflows changed the way that people get the work done, the way that researchers or designers can create objects and create scenes and kind of manipulate things?
24:17What's the impact of AI been so far on these workflows? Yeah, I think this technology really democratizes access to these tools and basically it gives everyone the chance to become a creator. I have no idea how to use 3D software. I tried a few times, but now I could be reasonable. Reasonable if I wanted to actually use robotics, I can reasonably get this object in a simulated world. The cool thing is that it also gives additional superpowers, you write, to creators that have the talent, you know, so artists, designers, they can actually use this technology to now do many more creative things.
25:00I have seen so much amazing stuff coming out that I wouldn't even think of. I think it's really kind of empowering to the entire population in different ways, which is great to see. We had Danny Wu from Canva on recently. He's the head of AI products there, if I got that right. And he was describing a similar thing, but kind of more on a level I could relate to because I write, I talk, I mostly work with words. I can't draw or paint to save my life, you know. And so that ability, having that superpower, if I want to see how something might look, an idea, it lets me do that now, right? And so I can only imagine the 3D physical world with, you know, 3D design and talking about simulations.
25:40The stuff you've seen must be pretty cool. Yeah, I think so. So talking about this 3D world and physical AI, and you spoke to it a little bit earlier, But how are all of these advances with the technology and computer vision included enabling robotics, autonomous vehicles? You talked about it a little bit, but maybe you can kind of put a point on, you know, how physical AI has really started to take off. Yeah. If it's anything that people take away from this talk is that physical AI can scale through real world trial and error. Yeah. It's simply not possible to put my car out there or a robot out there and it's going to mess up my kitchen here by bumping everything and so on.
26:26This is super expensive, unsafe, and it's just going to take us forever to get there. So simulation is really the answer here. And if we do it right, if we are actually able to use computer vision and other techniques to basically go somehow create these virtual worlds that feel real, then it's possible to train this kind of parallel virtual universe and safely. Essentially, basically, you know, accelerating time before we can deploy robots and also bring in the overall cost down, right? because now we are doing it in the cloud as opposed to having this very... To remodel your kitchen after every test, yeah.
Read the full transcript
27:10So what are some of the methods that are key to making simulations physically accurate or more physically accurate as we go? Yeah, I think like the jury is still out. How exactly to achieve physically accurate, like something I can completely trust, simulation at a scale, diversity, and realism of the real world. It's hard in a traditional way with this like different physics sellers. You know, that's hard. For the world models, it's also kind of hard. There's still hallucinations and, you know, sometimes objects disappear, go one in another. And obviously that's going to keep improving. So the likely success is going to come in some sort of a combination of both.
27:54And obviously we're going to keep pushing on each direction, as good as is possible. and maybe in between until we reach a point. There is a combination of that, right? Using these world models with this more traditional approach that really makes sure that physics and simulation is correct. Yeah, and the other very important message is that in computer vision and robotics, the really big breakthrough is a VLM, so visual language model that is able to reason. This is basically how humans navigate the long tail of very diverse and rare scenarios of the physical world. So we're kind of bringing that knowledge from language into the physical world.
28:38We're encountering completely new situations that we have never seen before in training. And now the VLM could come to our rescue, basically like release all this long tail. And that is really kind of the discontinuity from before. That is a tool that we have now that before was missing. So that's probably the most bold statement I can make right now. Fair enough. Do the traditional methods get baked into these models, or how do you go about combining them, the two approaches? Yeah, what you could do, for example, is you can use kind of the traditional way, which is also not full of AI everywhere, right, to kind of have a coarse simulation in 3D with solvers that we know.
29:22We know how to model certain effects, right? So you can make that simulation, you can render it out, and that becomes a guidance to our world model. You know, I take that as input. It's kind of telling me, oh, I should really roughly here and here. And then it becomes much more feasible to create pretty pixels out of that, both in time and space. So that's kind of like what we're thinking right now. Yeah. Cool. I'm speaking with Sonia Fidler. Sonia is a vice president of AI research at NVIDIA. And we're talking about the work that her spatial intelligence lab in Toronto has been doing, along with the evolution of AI and models and solvers in the mix and all the things that go into making these models more accurate so we can rely on them and trust them, Sonia, as you were saying.
30:07And we've also talked a little bit about SIGGRAPH. And I mentioned at the top that you gave the keynote a special research address alongside some other NVIDians at this year's SIGGRAPH. What are some of the notable things from NVIDIA's presence at the show that maybe we can impart to listeners here? What are some of the things that they should take away from what NVIDIA did at SIGGRAPH? Yeah, I mean, this year at SIGGRAPH, we really tried to send a message on physical AI in the keynote. And the reason is because this is a really important area with big impact. And the SIGGRAPH community has a lot to give, a lot to give.
30:43A lot of expertise is already there. the Gaussians, flats, nerves. I mean, that all comes out of that, right? We discussed simulation is key. Like I literally mean simulation is key. So everyone in the audience should feel empowered to help us in this quest of robotics, you know? And the cool thing is that it feels also early stage. Like I said before, it's open-ended. We don't kind of know yet, you know? We are hypothesizing. So we really hope the audience kind of connects with, here's a new challenge for you. Yeah, yeah. What do I do next, right? And graphics is very mature, but here is a new challenge for you that maybe needs to think outside the box.
31:25So I think the key to success, you know, we suspect will be the combination of Cosmos. So this is NVIDIA's World Foundation platform. And this is both video generation, so simulation of the world, as well as reasoning. reasoning about the laws of physics, reasoning about all the agents in the scene, the scenarios, and so on, and physics simulation. So all these three pieces interacting together, you know, that's our bet of creating really physically, but also semantically accurate simulations of the real world in the future. So I think that's really kind of like what I hope the SIGGRAPH audience takes away from the keynote.
32:05Yeah, that spirit of, I mean, it even goes back to what you said about when you met Jensen and he invited you to come work with him, right? That spirit of collaboration, you know, open source being what it is, conferences, obviously, but we've had guests across all different industries come on the show and talk about how important it is to, you know, share research and trade notes with other people who, on the other side of the world, working in other industries, et cetera. As AI continues to touch and evolve and change, you know, virtually every industry you can think of. How important is that?
32:43Or where are you getting from this experience of working across so many industries? And does it feel like AI is kind of bringing industries together? Or does it feel like, you know, different industries are kind of hunkering down and siloing in their own approach to how they use these immersion technologies? Yeah, it's definitely bringing them together, right? Because the workflows essentially are very similar. At some point, they're very similar. And the difference is the data and the expertise, domain expertise that's different. And actually, there is even sharing, you know, how I do autonomous driving versus humanoid robots versus factory simulation architecture.
33:21There's some commonalities between things that could be shared. And, you know, having this kind of like data-driven approach to simulation could really bring industries together and benefit from one another and build tech that can essentially make all of us, all of these industries better. And open source, you mentioned open source. I am a believer in open source. And it's great to see NVIDIA is also a big believer in open source. Like I said, you know, in a lot of areas, we're also still early on. And that's the only way to keep progress going, you know, and really build up these capabilities.
33:56Absolutely. Even though in a lot of ways, the timeframe that we've been talking about the last, you know, seven, eight years in particular, isn't all that long in AI terms and particularly in this recent, you know, kind of generative AI revolution, it's a long time. So you've been doing this, you know, almost since the beginning of early object recognition and AI kind of now through to everything we're talking about today. What's next on the horizon? Is there a breakthrough that you're either sort of waiting for or, you know, maybe more secretly kind of thinking like, I think this is going to happen soon.
34:27Whether it's something that's particular to your own work that you're doing or kind of more broadly, what's the next big breakthrough in AI that you're looking forward to or maybe just kind of hoping to see? Well, I think it's going to be robots, right? So I started the story with my sister in a car. Your sister, of course. And me dreaming about a robot taking the dog out in the morning and my parents made me do and I really like to sleep in the morning. That was kind of the early dream. My grandma lived 30 years alone. My grandpa died quite early. So the first talks I gave as a faculty were all started with a grandma and a Wally cute-like robot in a kitchen talking to each other and the robot helping her.
35:17It's just kind of a common thread of let's build this technology because it can be really powerful and useful. And I now believe, after many years in this field, that we're likely going to see that in our lifetime. Robots in some form, autonomous cars are already out there to some extent, right? And more is coming. So that's the breakthrough I'm looking forward to, yes. Have you ever had a robot in your home with you for a period of time? You ever lived with a robot? I have a robot that kind of wipes the floor. I would look to have something that does more than that. So, Sonia, as we look to wrap up here, and this has been fantastic again, thank you for taking the time.
36:05What advice would you give to researchers out there who are interested in the work the Spatial Intelligence Lab is doing, might be interested in collaborating, working with you in some way, joining the lab, collaborating from afar? And then in particular, what are some of the skills and research areas that you think are becoming increasingly important now and, you know, will continue to be at least for the next few years? Yeah, actually, the bar that we have is both low and high. So I'll explain what I mean. So I think what we are looking for is people with immense passion. You know, I feel like I still haven't lost the passion of the first day.
36:45you know i wake up and i am excited so i think the passion is what drives us forward the energy this is only a podcast but the passion comes through in your voice i'd say you're doing all right yeah i mean that's what drives it because you know as a researcher life is not easy most of the time things don't work you know that's basically what we do it could be six months a year where you don't get that result so you're really that kind of like you know passion and energy that makes you keep going. I think kind of wanting and having the ability to go technically deep is very important. You know, not jump from one thing to another as things go hard, but like let's learn really the fundamentals to the level we need so we can innovate.
37:29And I guess like maybe to my first point, the high level of perseverance, right? Like we want to keep going and there is no wall thick enough, right? the rest I think we can teach people like if you have these basic things a lot of the other stuff comes along yeah so in terms of like you know it's mostly also interest right so we are very interested in this 3d world modeling and understanding 3d worlds so people that are interested and shared passion for the same topic you know please contact us we would be very happy to work with you. Fantastic. For listeners who want to learn more about the lab, about the work we've been talking about, where are the best places to go online?
38:16Is there a homepage for the lab? Is it on the NVIDIA site? Any social media handles to follow? Where would you send them? It's all on the NVIDIA website. Probably if you search for any spatial intelligence lab NVIDIA or Toronto NVIDIA, which is our old name, it should pop out. Yeah. Fabulous. Sonia, again, this has been great. I think that the story, just imagining, you know, your dad telling you those stories as a kid and with your sister and then the advice your grandma gave you. Overcome your fear, get out there. Just absolutely fantastic. Congratulations to you and all the team, everyone you work with for all the work you've been doing and SIGGRAPH, of course.
38:57And we really look forward to following your progress in the future. Best of luck. Yeah, thanks. It was really fun talking to you.
39:06Thank you.
39:37Thank you.
From the publisher
Sanja Fidler, VP of AI Research at NVIDIA, joins the AI Podcast to share her journey from early curiosity to leading the Spatial Intelligence Lab in Toronto. Sanja discusses her path through research and what drew her to the world of AI and computer vision. She explains her team’s work on spatial intelligence—teaching AI to understand and create in 3D—and how this research is helping make content creation and simulation more accessible for everyone. She also discusses how breakthroughs in simulation, 3D modeling, and vision language models are powering the future of robotics and autonomous systems. Learn more at ai-podcast.nvidia.com.




