In short
NVIDIA AI Podcast Episode 226 Summary
Episode Details
- Title: Michael Rubloff Explains How Neural Radiance Fields Turn 2D Images Into 3D Models
- Recorded at: GTC 2024, San Jose, California
- Host: Noah Kravitz
- Guest: Michael Rubloff, Founder and Managing Editor of radiancefields.com
Episode Overview This episode delves into the innovative technology of Neural Radiance Fields (NeRFs), a transformative tool allowing users to create hyper-realistic 3D models from a series of 2D images or videos. The discussion covers the technology's implications for various fields, including creative and commercial applications, and its potential to revolutionize how people capture and experience the world.
Key Concepts
What are Neural Radiance Fields (NeRFs)?
- Definition: A technology that enables the creation of 3D models from 2D images or video.
- Functionality: Produces a lifelike 3D representation that can be viewed from any angle, eliminating the limitations of traditional photography.
- User Input: Typically requires 40 to 100 images for optimal results; methods exist that can reconstruct models from as few as three images.
Technical Process Behind NeRFs
- Image Acquisition: Collecting a series of images or video frames.
- Structure from Motion: Aligning the images spatially to determine overlaps and convergence.
- Rendering Techniques:
- NeRF: Utilizes a neural network to generate 3D models.
- Gaussian Splatting: Employs rasterization for a more efficient rendering pipeline.
Applications of NeRFs
- Personal Use: Documenting life events in 3D for personal memories.
- Commercial Use:
- Media & Entertainment: Streamlining film production by creating 3D environments without physical constraints (e.g., filming in locations like Grand Central Station without disrupting traffic).
- Gaming: Integrating generative radiance fields to create assets rapidly for game development.
Creative and Artistic Implications
- Expanded Artistic Freedom: Filmmakers and artists can explore new storytelling possibilities without being limited by traditional filming constraints.
- Dynamic Content Creation: Ability to create 3D models that can be animated or interacted with, offering new avenues for storytelling.
Future Considerations
- VR and AR Potential: NeRFs maintain view-dependent effects in virtual reality, allowing for immersive experiences that mimic real-world interactions with light and space.
- Technological Evolution: As computing power increases, the quality and accessibility of NeRF-related technologies are expected to improve, making 3D imaging more mainstream.
Notable Examples
- Sports & Entertainment:
- The Phoenix Suns utilized NeRFs for a promotional video.
- Music Videos:
- Various artists including Zayn Malik and Usher incorporated NeRF technology in their recent works.
Conclusion Michael Rubloff emphasizes that the advent of NeRF technology represents a paradigm shift in how we document and experience visual content. He encourages listeners to explore available tools and platforms for creating their own NeRFs, suggesting that if one can take a picture, they can create a NeRF.
Additional Resources
- Website: [radiancefields.com](https://radiancefields.com)
- Tools for Creating NeRFs:
- Luma AI
- Polycam
- PostShot (Windows)
- Nerf Studio
Listeners are invited to experiment with these tools and engage with the evolving landscape of 3D imaging technology.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:10Hello, and welcome to the NVIDIA AI podcast. I'm your host, Noah Kravitz. We're coming to you from GTC 2024 in San Jose, California, and we're here to talk about nerfs. No, not foam footballs and dart guns, but neural radiance fields. What is this kind of nerf? It's a technology that might just be changing the nature of images forever. Here to explain more is Michael Rubloff. Michael is the founder and managing editor of radiancefields.com, a new site covering the progression of radiance fields-based technologies, including neural radiance fields, aka NERFs, and something called 3D Gaussian splatting that I'll leave to Michael to explain.
0:50Michael, thanks so much for taking time out of GTC to join the AI podcast. Of course. Thank you so much for having me. So first things first, goofy football jokes aside, what is a NERF? What does that mean? Yeah. So essentially, you can think of a NERF as they allow you to take a series of 2D images or video. And what you can do from that is you can actually create a hyper-realistic 3D model. And what that allows for is once you have it created, it's like a photograph, but it's perfect from any imaginable angle. Composition is no longer a bottleneck. You can do whatever it is that you would like with that file and it will look lifelike.
1:27So if I were to take, I don't know how many, two, three, five pictures of the two of us sitting here right now in this podcast room, if you will, I could put those together into a Nerf and then I would have a file that I can look at from different perspectives. I can sort of move through. How does that work from the user experience side? Yeah. So typically the recommended amount is somewhere between like, I'd say 40 to 100 images. It's really easy from like a video because then you can just slice up individual frames from that video. But there are some methods, actually, that are going all the way down to three images and it's actually able to reconstruct, which is just mind-blowing.
2:13I'm too used to, you know, few-shot and zero-shot learning and things like that. So I'm like, one image? Yes, they are getting there. They are getting there. And there's been a ton of amazing work, one called Reconfusion from Google, which is just shocking. It can go down as far as three images, and it's still Very compelling. But yeah, so once you actually have gone ahead and taken your images or video, you would run it through a Radiance Field pipeline, whether that's a Nerf-based one or a Gaussian Splatting-based one. There are several cloud-based options where it's just drag and drop your images and it does all the work for you.
2:47And the resulting image or resulting file, yeah, you have autonomy over it and you can kind of experience that, whatever you've captured from whatever angle that you would like or whatever your use case might be. for it. How are they created? Like without getting too technical about it, can you kind of give an overview of kind of what's going on behind the scenes to put these together? Sure. So once you have your initial images, the first step on both NERFs and Gaussian splatting is running it through something called structure from motion, where essentially you're taking all the images and kind of aligning them in a space with one another.
3:24So it's kind of taking a look, saying like if image X is over here and image Y is over here. You know, here's how they overlap and converge with one another. And so that's kind of the baseline approach for both methods. And from there, they each have their own training methods where NERFs have a neural network involved in the training of them, whereas Gaussian splatting uses rasterization. Okay. And I guess that kind of begs the question of beyond what you just said, or maybe that's it, but what's the difference between a NERF and a Gaussian splatting? Yeah. So NERFs essentially were created first.
3:55They were found as a joint effort between, you know, the University of Berkeley and Google. Okay. So, Nerfs have an implicit representation for them, and, you know, they're trained through a neural network. So, there are some, there's a lot of work being done to get higher and higher frame rates associated with that, whereas Gaussian splatting uses just direct rasterization, and so you're able to have a much more efficient rendering pipeline where you can really easily get 100 plus FPS and you can use them with a lot of different methods as well. So they're very compatible with 3.js and React 3 Fiber where, you know, you can see them being used in website design now and being on platforms like Spline.
4:36Very cool. And so that kind of gets into the next question, which is how are they being used now? I saw some examples on, I don't know if it's the NVIDIA developer blog, and they used a kind of an obscure song that my wife and I like a lot. So it was perfect, right? But it was like a, it was a Nerf, I believe, of a couple walking down memory lane. They were walking outside and, you know, and the foliage around them. And you were able to, I was able to sort of zoom around from different points of view, see their front, see the back, look at the trees, that kind of thing. But beyond sort of a demo scene like that, how are Nerfs being used out in the world?
5:11Yeah. So that specific one actually is of my parents that I took. Oh, it's your parents. Yeah. Very cool. And so I took, because that's one of the major use cases for me. I want to be able to document my life not only in two dimensions, but I want to have a hyper-realistic three-dimensional moment in time frozen. And so for me, that's one of the personal use cases. But on a more commercial basis, where you're seeing a lot of the early adoption is in the media and entertainment world as well as the gaming world too. So, for instance, Shutterstock has been putting together a library of radiance fields where essentially what you're able to do – so say, hypothetically, you are wanting to film in Grand Central Terminal.
5:52But you cannot afford to shut down all the traffic and all the trains and all the traffic through that to film. What you can do is using a radiance field, if you capture it once, you're able to then bring that file into Unreal Engine and into a virtual production environment. and then you can film infinitely. And there's no more rush outside of the actual rental rate and you can go and get the shot that you actually need. And that's where it's starting to get adopted pretty early on. And similarly, in the gaming side of things, through generative radiance fields, because you are able to create these from text and images and newly video as well.
6:31Now you're able to drop these assets that take, I think the fastest method currently takes about half a second to create a full 3D model. Wow. And you can put that straight into, you know, one of the game engines. Right. Yeah. And so let me sort of play this back and see if I'm grasping it correctly. So if I were to go to Grand Central and do my very short, relatively speed, my very quick shoot, and come away with enough images to create a nerf, I would then be able in a virtual production environment to kind of create scenes or put elements into scenes from all of these different points of view, not just from a single perspective?
7:11Yes. Is that kind of the big? Yes. Yeah, that's correct, where essentially you're able to sync up the Nerf to the actual camera, and then you can use the full virtual production pipeline to go ahead and create. Right, right, right. Wow. This is something I probably should have asked you at the top. So listeners, forgive me for not having a more scientific inquiry kind of way of organizing my own thoughts. But a radiance field, what does that term mean? That's a great question. So essentially, you could think about a radiance field as a – well, if you just break it down into the two simple words, the radiance is just what that individual color would look like based upon your viewing direction.
7:53So say if you're looking at like a, you know, say a glass or something, and you can see that there's an actual reflection there. Depending on how you look at it, you know, what radiance fields offer is something called view-dependent effects. So as you move your head around, just as you do in real life, light changes, light shifts and reflects. And just like that, radiance fields are able to model that effect. So you could think of radiance fields as the actual shift in colors at a given space. So knowing that radiance at a specific point, you could take a look at that as being radiance, whereas it's being contained inside of a field.
8:25Got it. Okay. So you said something a minute ago about generative AI, using generative AI to create radiance fields. And you may have used a term that slipped my mind, forgive me. Is it the same basic principle as doing a text-to-image, using a text-to-image model, chat GPT or DALI, stable diffusion, whatever it is? Is it that same principle that I enter a text prompt and then the system can create a radiance field? Or is it creating a series of images? Or can you get more, you know, kind of more control than that over it? Yeah, so exactly. What it does is it will create a series of images of a singular object.
9:06And actually, in some cases, they're starting to release some papers where they're creating multiple objects. But each object itself is either a Nerf or, say, a Gaussian splatting file. And from that, that's what's actually being used to train the actual resulting 3D model. But they're able to do that in a fraction of a second. So earlier this week, you hosted a session at GTC. I unfortunately wasn't able to make it. I wasn't in town just yet. We were talking about it before we hit record. And you spoke to, if I've got this right, some of the artistic implications, possibilities around using Nerfs, and then also some more kind of business enterprise-oriented applications.
9:46How did that go? What kinds of things did you talk about? And then I'm kind of curious what the audience reaction was, either that night or kind of more generally, you know, what, as people learn about NERFs and Gaussian Splatting, what kind of the reaction is and does it spark imagination and sort of what are some of the implications? Yeah, I was surprised by how many people actually came out to attend. And so it seemed like there was an extreme amount of interest in terms of just like visualizing itself. So I had a roughly 20-minute video of just looping different examples of radiance fields that I've created and some of the people in the community have as well.
10:26And so I think there was a lot of interest across a wide variety of industries where, you know, I spoke to professors, I spoke to people working on offshore drilling sites, I spoke to, you know, physicians, people of really diverse backgrounds and use cases, but I think all of them will be affected by radiance fields. And is the interest in some of those, because I want to ask you about sort of artistic creative implications as well, but we'll put a pin in that for a second. Is the interest from, say, a physician or the offshore drilling site makes me think of use cases of robots and drones, you know, to be able to go to places more safely than sending a human to inspect something?
11:06And is it along those lines of being able to create, you know, create a radiance field and then from a, you know, quote, safe environment, be able to inspect different aspects of the offshore site from different angles? Is it that kind of thing or is it something totally different? Yeah, no, it actually is quite similar where, you know, if you have a predetermined camera path or you, you know, give the necessary information to the model, it will be able to create a hyper-realistic view of what it sees. And from that, you can then, you know, flag for a human if they need to go and make a visit for actual maintenance or repairs.
11:44Right. And so you're able to really give a hyper-realistic look for that specific use case for asynchronous maintenance. Right, right. Are there implications with, you know, VR and extended reality and augmented reality? Yes, yes. And so that was actually one of the demos that we were showing. It's just like VR applications because, you know, they still retain their view-dependent effects when you're in VR. Right. And so, you know, as you move around a scene, we as humans expect light to behave in a certain way. Sure, yeah. And with this, that continues to hold true. And, you know, with Radiance Fields, you can actually walk through the entire scene.
12:22And so it is, I think, the closest thing to actually stepping back into a moment in time that we have. Yeah, amazing. I'm speaking with Michael Rubloff. Michael is the founder and managing editor of RadianceFields.com, a website that's covering the progression of Radiance Fields-based technologies. And we've been talking about them, about neural radiance fields, NERFs, and 3D Gaussian splatting, these techniques that allow us to stitch together 2D images and create a hyper-realistic 3D model, 3D environment that we can do all these different things with. I mentioned, wanted to ask you about some of the artistic implications, and this might be a weird leaping off point, so redirect me if it is.
12:59But recently, we recorded a podcast with a woman from a company called Cubrix, and I'm going to get this wrong, forgive me. But it's basically kind of like a digital soundstage, sort of an advanced digital green screen type of thing that you can use in filmmaking, video making, and, you know, powered by generative AI, kind of similar things. And I remember asking her, what advice do you have for burgeoning filmmakers who, you know, are interested in creating, but wondering how to go about it in this age with all these AI tools now becoming available and advancing so quickly. And her answer was not what I expected, but it was really interesting.
13:40She said, well, the first thing you should do when you're thinking about using generative AI in filmmaking is really delve into your own subconscious. And if I had understood her correctly, I think she was talking about, you know, the capabilities of what types of images and moving images you can create with generative AI tools at your disposal go well beyond what you could create without them. And you're not limited to capturing reality, so to speak. You know, you can create reality, which people have been able to do with, you know, technology for a while now, but easier, faster, you know, perhaps better.
14:17What do radiance fields do or what do you think that they are doing and could do for creative applications? Yeah, I see a very large creative opportunity for radiance fields going forwards. You know, I think that they allow people to take larger risks or be able to actually I actually wouldn't even call them risks because what you can do with them is if you get captured, you can film in post. You don't actually need to film up front. You just need to capture it. And I think that that allows, you know, a lot more thinking about, you know, where you're not being constrained for time to think like, what are the actual camera movements that we want or what's the best way to actually tell this story?
15:00And you have the ability to align and say, if you are a director, you can go to your director of photography and say, you know, here's the exact camera movement I'm trying to convey in this, or here's the exact thing I'm trying to show. And I think that that's going to really supercharge, you know, productions. And in the same way, I think that it's also going to allow more stories to be told in places that we've never been before because we can be transported to these places and be exposed to an actual lifelike interpretation of wherever you want to take an audience. And in the same way, I think for students and for independent filmmakers, it represents such a massive opportunity because you will have the ability to go to locations and tell stories in locations that you may have always dreamed of, you know, shutting down the Las Vegas Strip, for instance.
15:51But now you're actually going to be able to do that. Right, right. Where are we at in terms of the technology when it comes to the resolution of images and particularly in the backgrounds of the images? Are we just constrained by, you know, the quality of your camera and the available compute? Or is it more intricate than that? No, I think that where we are right now, it's fascinating. There's a Meta Reality Labs paper that released late last year called VR Nerf. And essentially what they did was they created this camera rig, effectively known as the Eiffel Tower, where it has 22, I think, Sony A9 cameras strapped onto it, all facing different directions.
16:33And then from that, what they would do is they'd go into a room and then they'd push it through the room and each camera would take nine bracketed exposure shots across different exposures. And then each one of those images would be compiled into a single HDR image and then that image would be trained in a Nerf. Okay. And the resulting image quality of that, you know, it approaches the IMAX level quality. Wow. And it's able to be reconstructed. It's not a bottleneck in terms of the visual fidelity. It's able to handle that. It's more of a compute issue right now. But, you know, obviously, as time goes on, we're going to get more and more efficient computers as well.
17:10And so it's more a proof of concept to me than anything is that, you know, when we get to that level, that's the floor. Yeah, right, right, right. And beyond capturing moments in time for your own use, are you doing other things yourself with the technology right now? Yeah. So I've been doing some consulting work for businesses that want to implement the Radiance Field-based technology into their offerings. And I can't talk too much about the work. Sure. But one of the ones I'm really excited to be working on is Shutterstock, where it's just how do we create these assets that are available for people to actually use today?
17:49Right. Amazing. So kind of out in, you know, in the mainstream world, so to speak, right, in the pop culture world, I would imagine that we've probably seen nerfs in action and just didn't realize it or didn't know how to name it. Is that off base or are there some examples of things still out there that listeners might have seen? Yeah, there have actually been some pretty high-profile uses of Radiance Fields where earlier this year, the Phoenix Suns, actually, the entire team was nerfed. And so there are nerfs of Kevin Durant, Devin Booker, the entire team. And they actually are using it as part of their introductory video for this season.
18:26Oh, like during the games and starting lineups? Yes. Yeah, cool. Yes, and it really, like, showcases a way that, you know, if you as a business can help bring fans closer to the action because, you know, you have these lifelike interpretations that are doing the most insane camera movements that, you know. And so is it like Katie's going up for a shot and the camera sort of seems to, and I don't know if the motion stops or not, but the camera sort of stops and then tracks sort of around him from a different angle, like that kind of thing? Yeah, that's actually extremely close to one of the examples where he's about to dunk and it kind of flies up around him and circles around, which would be very difficult to go ahead and create.
19:05But, you know, Riddensfields make that actually very surprisingly easy. Yeah, amazing. Yeah. And there's also been a lot of other high profile use cases within the music industry. So Zayn Malik, for instance, has a music video for Love Like This, which I think got something like 10 million views in the first 24 hours, which is just insane. There's a music video for R.L. Grimes' Pour Your Heart Out, which is actually comprised of over 700 individual nerfs. Oh, wow. It is just an insane endeavor. And every single shot in that music video is a nerf. There's also Usher for his most recent single. I think it's called Ruin.
19:43It has a few different examples of Gaussian splatting in it. And then Polo G has a song that just released called Saris and Ferraris that also are using Gaussian splatting. And then there's also Chris Brown has one as well. And J. Cole and Drake put out a music video, I think, within the last week or two. And there is a Nerf hidden in there. Nice. And I don't know if this counts or not, but in Jensen's most recent keynote, there actually is a nerf in the background in one of his slides. All right. Should we make it a contest for listeners to spot it or you want to call it out so people can go look?
20:20If you guys want to, or if you guys, you know, want to pause it and see if you guys can spot it from the two hours. This is great. Everybody, you can go queue up your YouTube playlist with all the music videos you just mentioned. Yes. And then, you know, rewind the podcast, listen back, you're spotting the nerfs. Yes. And then when you're done with the pod, go back, rewatch Jensen's keynote and see if you can find it. Yes. But it's in this slide where I think he's taking a look at the different modalities in which NVIDIA operates. And it's under the 3D tab. And it's like a coastal cliff view. And it was actually taken by one of my good friends, Jonathan Stevens.
20:54Nice. And so it was really cool. That was the first time, I think, that we've seen nerfs being featured in the keynote. Right. Excellent. Shout out to Jonathan. Nathan. Michael, closing thoughts. Nerfs, Radiance Fields, 3D Gaussian splatting. For the listener who, you know, never heard of this stuff before, listened to our conversation, you know, ring some bells, spark some ideas in their head. They're thinking about going out and, you know, exploring some of the music videos we talked about. Where do you think this is going to go over the next, whatever the time period is, couple of years, 10 years, generations?
21:28Are all of our photos going to become 3D multi-perspective, you know, sort of models going forward? Are we forever living in a world of 2D images and then we figure out ways to make them more like 3D models? Where's the future of imaging and sort of post-processing headed? Yeah, it's a great question. And, you know, my opinion is that we now have the ability to no longer be constrained to 2D. And, you know, 2D is not actually how we experience our lives. And it's, to me, I feel like, you know, it should not be the final frontier of imaging. And now that we have the technology to do so, I think it's really time to begin exploring how we can actually document our lives in a lifelike way to the way that we actually experience life.
22:19Because not only can you create static nerfs or Gaussian splatting files, but you can also create dynamic versions of them too. And so if you can imagine an analogy of static nerfs to photos, you can also do the same with videos. And so I think that we're really entering into an age where imaging is not the same as it has been since the inception of photography. Obviously, it progressed significantly, but I think that now the technology is there where we can just take a fundamental leap forwards into an entirely new dimension. Come to GTC and you leave in a new dimension. That's how it works. Michael Rubloff, thank you so much for stopping by the podcast.
23:01Your website, again, is called radiancefields.com. Are there other places for people who want to follow your work, learn more about the space, you direct them to go? Are there websites, social media accounts, anything? Yeah. So all my social media handles are just at radiancefields. Okay. And so that's just generally LinkedIn, Twitter are the two big ones that I mainly post on. But yeah, I would just encourage all listeners just to try downloading some of the platforms themselves and some of the really good ones to get started. You can take a look at Luma AI, Polycam. If you're on Windows, you can download PostShot or Nerf Studio.
23:39And, you know, they're all free right now. And it's not as bad as you would imagine to actually capture everything. It's actually quite straightforward and is pretty forgiving. So, yeah, give it a try. If you can take a picture, you can make a Nerf. Exactly. Excellent. Thank you again. Pleasure talking to you. Thank you so much.
24:27Thank you.
From the publisher
Let’s talk about NeRFs — no, not the neon-colored foam dart blasters, but neural radiance fields, a technology that might just change the nature of images forever. In this episode of NVIDIA’s AI Podcast recorded live at GTC, host Noah Kravitz speaks with Michael Rubloff, founder and managing editor of radiancefields.com, about radiance field-based technologies. NeRFs allow users to take a series of 2D images or video to create a hyperrealistic 3D model — something like a photograph of a scene, but that can be looked at from multiple angles. Tune in to learn more about the technology’s creative and commercial applications and how it might transform the way people capture and experience the world.




