In short
Podcast Notes: The Neuron: AI Explained - Episode Overview
Episode Title NVIDIA’s Kari Briski on How to Use NVIDIA Nemotron Open-Source AI
Episode Description In this episode, hosts Grant Harvey and Corey Noles are joined by Kari Briski, VP of Generative AI for Enterprise at NVIDIA, to explore NVIDIA's Nemotron open-source AI models. They discuss the model's features, minimum hardware specifications, the differences between various tiers, deployment patterns for businesses, and the choice between local and cloud AI.
Key Concepts and Discussions
Overview of NVIDIA Nemotron
- Definition: NVIDIA Nemotron is a comprehensive suite of AI resources including foundation models, datasets, training frameworks, and algorithms.
- Purpose: Aimed at empowering developers to create specialized AI solutions without needing to rely solely on cloud-based models.
Open Weight Strategy
- Transparency: NVIDIA has over 500 published models and datasets, making it one of the most transparent platforms in AI today.
- Developer Attraction: By sharing resources openly, NVIDIA aims to attract more developers to its platform, fostering a community that can provide feedback and innovation.
Hardware Specifications and Model Tiers
- Tier Differences:
- Nano: Designed for smaller GPU environments (e.g., laptops).
- Super: Targeted for single data center GPU use.
- Ultra: Needs multiple GPUs but offers enhanced performance.
- Minimum Requirements:
- Nano requires about 18 GB of memory and is efficient for localized applications.
Choosing Between Local and Cloud AI
- Local AI: Preferred for businesses needing control over proprietary data and specialized domain applications.
- Cloud AI: Useful for general applications but may compromise data privacy.
Practical Applications and Use Cases
- Specialization: Businesses can utilize Nemotron to tailor AI applications to specific verticals (e.g., healthcare, finance).
- Efficiency: Smaller models can achieve significant performance while maintaining accuracy through processes like model distillation and reinforcement learning.
Long-term Vision for AI
- Future Trends: There is a growing emphasis on smaller, more specialized models rather than massive, generalized models.
- User Interaction: The vision includes developing applications that allow for natural language interaction and deeper integration into everyday tools and devices.
Key Takeaways
- Empowerment through Open Source: NVIDIA's commitment to open-source models serves to enhance developer capabilities and foster innovation.
- Model Selection: The choice between different model tiers should be based on use cases, computing resources, and accuracy requirements.
- AI's Broader Integration: The goal is to integrate AI into daily applications, making interactions more personalized and efficient.
Resources Mentioned
- NVIDIA Nemotron Models: [NVIDIA Nemotron](https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/)
- Prototyping Tools: [Build NVIDIA](https://build.nvidia.com/explore/discover)
- The Neuron Newsletter: [Subscribe Here](https://theneuron.ai)
- YouTube Channel for More AI Interviews: [The Neuron AI YouTube](https://www.youtube.com/@TheNeuronAI)
Conclusion This episode sheds light on how NVIDIA's Nemotron family of models offers a powerful platform for developers looking to harness AI while maintaining data privacy and specialization. As AI continues to evolve, the importance of these tools in everyday applications will only increase, marking a significant shift in how businesses interact with technology.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:07Welcome, humans, to the Neuron Podcast. I'm Kari Knowles, joined as always by the one and only Grant Harvey.
0:13Kari Briski:Hello, hello. Today, we are unpacking Nematron, NVIDIA's open weights model family built for everything from laptops to data centers. We'll get concrete, what the average person can do, what power users can build, how businesses should decide between local versus cloud, and how to pick between nano, super, and ultra. And for this, we're joined today by Keri Brisky, VP of Generative AI for Enterprise at NVIDIA. Keri, welcome to The Neuron. Hello. Thank you for having me. Happy to be here. That's great. We're sure glad to have you and appreciate you taking time out of what I'm sure is a wild schedule to accommodate us.
0:52Anytime. I love it. I love little breaks. It's fun.
0:56Kari Briski:So, Carrie, let's start with just like a very simple overview. So what is NVIDIA Nemotron in simple terms and why should engineers care? NVIDIA Nemotron is a collection of not just foundation models, but it's also all of the data sets that we put out into the open source. It's recipes and it's also the training framework, algorithms and research. So it's just this all the work that you need in order to build a foundation model. So it's not just the models, it's everything around it. It's not just the model, but it's the ingredients and the recipe, too. That's awesome. It is. And that's one of the things that's really impressed me.
1:35And I'd like to chat just a little bit about NVIDIA's open weight strategy. You know, you've got more than 500 published models, tons of training data sets and much more. And I'd argue it's, if not among the most, maybe the most transparent approach in the AI space today. Can you tell me a little bit about that strategy and why it was the right move for NVIDIA? Yeah, I think, well, it's the right move for us because we are a developer platform. We are an AI developer platform. And the more that you put out the ingredients, it attracts even more developers. And the more developers that you attract, the more feedback that you get on your product and what you're putting out.
2:13And so it was the right move for us because we believe that everyone should be able to do their life's work and make specialized AI. And in order to do that, you need to have all the ingredients. That's right. I even joked with Grant. I said, even the Groot data and model are out there. Yeah.
2:30Kari Briski:If you could make a robot, you have the technology. Yeah. It's like our three family of models, right? You have Nemotron, you have Cosmos, and you have Groot. Those are kind of our foundation model families. Yeah. Wait, Cosmos, is that the world model one? World foundation model. That's right. Yeah. Yeah, that's right. I'm listening to the physical world, yeah. Nice. That's awesome. So, like, what sets Nemotron apart from, you know, like, obviously Cosmos and Group, but then other language models in its class, like, in terms of architecture, training approach, or optimizations? Yeah, I think they're very different.
3:04So, Nemotron is really kind of the brains, the understanding, the reasoning. it started out as a textual text-based model but now it also understands images and and the nano does and then also some video coming soon and we'll get to vision later later down the line but what sets it apart is that it again the transparency and also the reason why we build it I mean we also build it for ourselves to develop our next-gen architectures for like systems at scale and so we also are building it for ourselves. So we have to understand how to build it. So we understand how the systems work. And then when we put it out into the open, I mean, I mentioned how we also released the data sets.
3:46When we started releasing the data sets, we had so many, even enterprises come out and say, Hey, that was, that's fantastic. Can you help me build a model too? Even though we put the ingredients out, like, Hey, we still want to kind of pick your brain and understand like how to use those data sets together. So it really sets it apart in that, again, it's just, it's trust through transparency. That's really awesome. And I think such a refreshing approach, to be honest. Yeah. You said open weights earlier, and I think that it's really open source, right? I think people started to say open weights because people picked apart the fact that things weren't open source.
4:21Yeah. You are so right. I'm so used to -
4:24Kari Briski:This is an actually open source model. Yeah. We've run into that so often where we got to stop saying open source, this is open weights, But in this case, it's really open source. That's right. That's right. Thank you. I appreciate that. Yeah. What can a highly technical person do with the Nemetron suite that's really impressive? Yeah, well, I think the sky's the limit. You know, we have, we use ourselves internally. We have things like deep researchers. You have your own deep researchers if you've ever gone out to Google or perplexity. And you can imagine that we have data that we do not want to upload into an API internally.
5:01And so we have our own deep researchers. We actually put out that blueprint for others to build their own deep researchers locally and for themselves. You can specialize them. We have a lot of customers who are specializing for their domain. So they're able to take their proprietary data, their IP, their personal data, and be able to specialize it for their use case in their domain. Because nobody knows your domain better than you. And you do not want to give away your intelligence, right? So I think that, so that's things you can do with the model. I think what's interesting with some of our recipes is that we put out the recipes and some people have taken our models and distilled them on their own.
5:36And when I say distillation, that means taking a larger model and quantizing it or, you know, reducing the precision of the weights so it's smaller and faster. So we've seen a lot of people take it and pick up our algorithms to do neural architecture search to kind of even change the architecture of the model. So if you're a real techie, you can get into the weeds and really do your own thing, like really change the guts of the model with the tools that we've given you. Okay, so let's talk about a tool that's completely transformed how I personally work. That's Whisperflow. Imagine being able to write full articles, emails, even take complex notes just by talking.
6:18That's what Whisperflow lets me do. It's hands-free writing that's smart, accurate, and ridiculously fast. For me, it solved a decade-long problem. I used to cover baseball as a beat reporter and had to file my story deep in the middle of the night. I always wanted a quality dictation tool that would let me get started on the way home. I wanted to be able to just talk my ideas into a dot. I tried several tools. They all came up short. But with Whisperflow, I can dictate an entire piece, have it cleaned up, and file all while I'm on the go. No more waiting to get back home, fire up a laptop in the middle of the night to rush something out.
6:53I was able to take time that I already had and use it. But it's not just about accessibility. It's about productivity. You'll save hours, get your thoughts down instantly, and stay in flow without breaking a note. So whether you're a writer, a founder, or just someone who needs to capture ideas quickly and accurately on the go, Whisperflow makes it effortless. Seriously, you're going to want to check this one out. You'll wonder how you ever worked without it. It's available on Mac, Windows, and iPhone. visit whisperflow.ai slash neuron today and get started for free. That's whisperflow.ai slash N-E-U-R-O-N.
7:32Tell them the neuron sent you.
7:35Kari Briski:And now obviously you've told us that Nemetron is more than just models. It's the data set, it's research papers, algorithms, libraries, and more. Could you tell us more about like about everything that is available to people? Yeah, I think what's interesting about the data sets is that, well, there's two things. We release the data sets that we've either created or acquired as much as we can. Sometimes we have distribution limitations, but then we try and recreate or do some sort of synthetic derivative that's differential enough to be able to release it. So what we've learned from the data set that we've acquired or paid for.
8:11and so we're actually kind of paying for this these data sets and then also giving them back out to the community the other thing is that the thing about reasoning and reinforcement learning is that you're not necessarily limited by you data you need data but you can now do synthetic data generation and you create and can you can create synthetic environments and so if you think of an environment like a gym where you know the more you exercise your model the better it gets and you give it different variations. So when it sees a problem in the real world, then it can know, say, hey, this is similar to something I've done before and I can go solve this problem.
8:48So if you think of all of the synthetic data, because you're just compute limited at that point. So when we have just GPUs that are doing synthetic data generation and then we package that up and put that out. And so that's a lot of compute savings for developers in the community because now we're not only giving away a synthetic data generation, we're also giving out these gym environments. We're also giving out algorithms for the reinforcement learning. And the research is another example. So we published a paper earlier this year on what's called a hybrid transformer. So it was both a Mamba and a transformer put together, which is really efficient because it reduces some of the attention, memory size of the attention layers.
9:31And so you're able to reduce the size of them all, have a highly efficient model for inference. and some research paper actually was picked up by the Quen model. So the latest Quen model actually released a hybrid Mamba transform architecture. So these are the types of things that we put out into the open for kind of anyone to pick up. And, you know, by the way, we're not just the only one putting out into the open. We love, you know, picking up and reading papers and taking the best of the open source too and putting it back into NemoTron as well. That's the virtual style of the open source.
10:01Kari Briski:That's really awesome. So why would a business want to use a model like Nemetron versus AI over the cloud, for example? Yeah, I kind of alluded to it a little bit earlier when I said that you don't want to give away your intelligence. So you have your own domain expertise and you have your own private IP. You have your own specialty. I think current AI is really great. I mean, it's really great. And I'm not saying it should go away because there's room for everybody. Because when you are doing an AI application, it's systems of models. I actually kind of joked about it with a couple of our team who were joking the other day.
10:44Because it's like a new object-oriented programming. You're sending agents out to do stuff and come back and return an answer. It's just like these objects to do something. So there's room for all types of size and models. But really, enterprises want to specialize. And I think when people are talking about how enterprises really haven't adopted AI fully yet, it's because AI has been very general. And enterprises and vertical industries are highly specialized. If you think about in healthcare, in automated driving, in finance, in chip design, in any name, in industry, you need to be able to take it that last mile.
11:18And so I think we're finally providing the tools. I mean, five years ago, this stuff was really hard, right? But now we're getting to the point where you can reproduce. We have the tools. It's getting easier to train. It's cheaper to run inference. And so now enterprises can take their data and they can apply it to these models and take it the last mile that you need for specialization. And that's what's really interesting.
11:41Kari Briski:that's awesome and also that reminds me too nvidia put out a paper i want to say it was like a couple months ago about how the future is basically small models and that that kind of opened my eyes to like exactly what you're talking about i was like oh okay i get this now yes there will be the giant labs like open ai and anthropic and google trying to make like the giant omni models but also i think like people were kind of thinking like in the future there's no need to create specialized models? And I think we're at the point where it's like, no, there still is. No, there's definitely a need for specialized models.
12:12I think we've been saying it for a couple of years. I just think it's taken, it's the journey of understanding generative AI, because it's just a completely different way of programming, right? It's a completely different way of software engineering, the way that applications are going to be built in the future. And so just as you've had to change out your build systems and your test-driven development in and your QA, like this is just having to redefine how you build software. And I think it's just a journey and it's another transformation journey that enterprises are going to have to make. And that was a really great paper about small models.
12:45It's one view. I think there's also, it's many models. And I think that's why, you know, you kind of introduced us and we have many sized models of Pneumotron because you do need larger models to inform and train the smaller models. And so you do, and larger models are a little bit more robust and they can kind of generalize a little bit more even in a domain and they have a bigger capacity to learn a new domain. So I think that there's room for all kinds of models.
13:11Kari Briski:Yeah, that makes perfect sense. So how does someone, say an individual, use Nemetron? What are the, you know, minimum hardware specs you could run one of the smaller versions or any of them on for that matter, whatever you've got handy. Yeah. So there was a little bit of thought that went into, it seems kind of natural that we did a small, medium, large, but there was a little bit more thought that went into it. For our Nano, we really wanted it to fit into smaller GPU formats, right? So that anyone that's either renting a, like, what I'd say, an older hardware architecture on the cloud or running on a laptop that has a GPU in it.
13:52We want to be able to run that with enough memory and with the efficiency that you need and still have a really great reasoning model. And so that was the Nano. And so it's a 9 billion per meter model. At floating point 16, you really only need like an 18 gig memory requirement, right? And then again, you can use our tools to quantize it down even further and make it even smaller and more efficient. Yeah, so that's kind of what we did for the nano. For the super, we were thinking we wanted to keep it within a single hopper GPU, data center GPU. And you can easily do that. You can run it in FP8 and that's easy.
14:31For ultra, we actually had this debate with our engineering team because a lot of these, again, a lot of the really great large models are larger than a node. And if you, let me just explain what a node is, if you don't know, it's eight GPUs in a box. So most of these really large models are multi-node. So you have to take up more than eight GPUs just to run one model. And so we actually kept the Ultra within a box. So it does take eight GPUs to run inference on or train or just have it within the system. But we did that very specifically. I can't say in the future that we'll stick to those rules for Ultra.
15:15We might expand, but that was our thinking between Nano, Super, and Ultra. That's good thinking. You wind up probably not wasting a lot of GPU space that I imagine you would otherwise if you're needing just over one or a little more than one. Yeah, we did a survey on what were most of the enterprises using, what are their cloud instances, what's available to them. And a lot of people were using Ampere A10s, right, or A10Gs or A100s. And so we really wanted to do that. And then, of course, we work with our GeForce team and understanding, like, what GPUs are going into laptops and what do we need to do to make sure that we have a really great, efficient, small model.
16:00Kari Briski:That's awesome. Yeah, I like that you're even considering people at that level because, like, you know, Corey and I, we would love to be able to run these models on our PCs. and it's just so cool that you're able to do that. Like, yes, you can go to the Frontier Intelligence, but you can also run it locally. It's so awesome. Yeah, and we do have a, I think we announced a little thing called Spark, which is a little development station. So Nano's great for that too. That's awesome. It's going to be great. Yeah, Spark looks neat. It's so cute. It's so cute. It is. Yeah. I remember seeing the footage of it.
16:35Kari Briski:One is like, Was that the video that Jensen announced where he like takes it out of the oven and he pulls it out? Or was that a different computer? No, I think that was a different one. I think that was like our data center rack, right? During COVID, right? Yes. Yes. Those videos are great. We'll include a link to that. Served it up, yeah.
16:58You know how AI coding agents feel fast until you're stuck fixing their code and realize it would have been quicker to just do it yourself? Warp changes that. It's an agentic development environment, a new kind of tool that makes working with coding agents effortless. Warp connects to your code base, understands your contacts, and uses the best models, from GPT-5 to Claude Sonnet 4.5, to generate production-ready code from the jump. And when you do need to step in, Warp's built-in code review and editing tools make it seamless. You can see diffs, reprompt, or make quick changes right alongside the agent.
17:34No tool switching, no wasted time. That's why over 700 ,000 engineers from companies like Netflix, Ramp, and Amazon are already using Warp to save hours each week. So if you're ready to go from prompt to production faster, check it out at warp.dev.
17:52Kari Briski:So the trend with Nemotron, as well as other open source models and open weight models of late, seems to be increasingly better performance out of the same or even smaller model sizes, like you're talking about. You might have alluded to this a little bit, but what exactly is involved in making that happen? Yeah. So I mentioned this idea of distillation, the algorithms to distill models in. By the way, there's lots of different ways to distill models. There's always kind of new algorithms. But you do need the larger model to teach the smaller model. You can also take a larger model, and I mentioned to do neural architecture search.
18:28And so you're basically, it's like pulling apart the Legos and kind of figuring out which pieces really matter to the design and then putting it back together and using 30 % less pieces. And so that's basically what you do when you do neural architecture search. And so you distill it down to a smaller model, but then you retrain it back to the same accuracy. So there is a bit of like retraining back into it. And so the thing is, there's really great success with small models or these SLMs. You do give up a little bit, though, just a little bit. I mentioned earlier, like a little bit of robustness, a little bit of capacity to learn.
19:05But they are great for specializing for particular tasks. So again, I'm talking about systems of models. If you have a model that just needs to do a routing, so someone asks a question and I need to route to whatever the best tool is or whatever the best model that is to answer this question, you don't have to understand the world, right? You just don't have to understand your environment and what its tasks are. And you can be a really great router, but with some reasoning behind this, you're making great decisions based on that routing instead of a rules-based routing. So yeah, SLMs have had that purpose, but you do need that larger model to distill down into it.
19:45That's really, really interesting. And one of the things that's impressed me a lot is like i run 9b and you know while while yes you know there are some minor sacrifices it's really not as much as you would think compared to a 600 trillion 600 billion model billion parameter model excuse me a 600 billion parameter model and uh like this and showing where yeah who's rocking that take me there but it's It's amazing to me that the difference isn't massive, at least in the tasks I'm doing with it specifically. So how do you make sure that smaller and faster doesn't mean less accurate? Is that that second phase of retraining you were talking about?
20:35Yeah, it's a lot of the post-training. It's a lot of this reinforcement learning. Again, for when you do reinforcement learning, you take a question or a query, And then you're actually, what we do is we synthetically generate at least 16 different variations of how you might ask the same question. And for each of those variations, you run that through a gym. And so this gym is a bunch of different environments. And so you might have a math gym of, if it's a math question, how might I answer that? What might, what's the algorithm I might do? If it's a tool calling gyms, what tools do I need? How would I, how would I think about this and how would I go get that?
21:07But you can imagine there's all kinds of different environments. and and if you really kind of look out in the community it's really great because all these gym environments are starting to pop up in the community but um so you do this this post training just to give it uh this breadth of knowledge uh we also have um for lack of a better benchmark but you there's these benchmarks out in the community so you want to make sure that you're at least meeting that minimum benchmark without overfitting to that so you at least want to check yourself how am I doing on some math questions? How am I doing on coding?
Read the full transcript
21:39How am I doing on instruction following? So you're just going to make sure that you're hitting those before you even release the smaller model. Right.
21:46Kari Briski:That makes sense. You mentioned the RL environments. This is something that we've learned about recently. So is it technically like a data set or like actual, like I've heard people spin up like a fake Salesforce for people to use or like, like could you kind of explain that? Because it's really interesting. Yeah, exactly. It's more than a data set. So that's why we call it a gym because you're kind of working it out. And if you really think about it, it's like a little agentic system on the side that you're kind of running the model through and you're exposing it to tools and saying, here, here's these things in front of you, go figure it out.
22:19I've asked you a question. You've put tools in front of you to figure it out. And so you're right. You're absolutely right. People are spinning up ISV or software vendor interfaces to be able to say, hey, here's what this button is. Here's what this function is. How would you do this invoice? How would you write? And so it's very interesting. I think the space is going to get even more interesting. That's awesome.
22:40Kari Briski:Yeah, I love that. So you mentioned the three tiers, right? There's Nano, Super, and Ultra. and I liked how you explained like the thought process behind like which one goes where. I guess like we kind of already know like what the North Star was that governed those splits was like the actual GPU that you would run it on. But what would you think is like what job should each tier own so teams don't just default to the biggest one? Like if you're in the weeds, maybe it's different for everyone. Yeah, no, I think my answer, and you're not going to like it, it depends, because it really does depend on the use case, your compute that you have, your latency requirements, right?
23:26So it's all about throughput latency and accuracy. And so when you triangulate between those, you kind of find the right model size for you and your task. and so we have found that the sweet spot has been the super because it again it fits in a single data for at least for enterprises it fits in a single data center gpu because enterprises want i mean all the accuracy they can get in an itty bitty compute footprint right so they can
23:53Kari Briski:run like as many as possible right yeah right right and and and save on cost right so they want they want the most accuracy that they can and the lowest latency on the smallest compute footprint And that's honestly what we've been trying to deliver with all these efficiencies that we've been doing. And then the super kind of gives that more accurate than the nano, not quite as accurate as the ultra. Actually, our latest super revision just surpassed our ultra. So we've got our roadmap going. There's always innovations happening. But the ultra does have safety and robustness tests and capacity to learn.
24:31So a lot of, like I've talked to healthcare customers who definitely prefer the larger model because of its capacity to learn on their domain and be more robust for a given task. So I hate to say it, but it depends. That's a totally fair answer. And kind of figured that would be the case, but it's really interesting the way you've explained it. When should teams recognize that it's time to lean in to one of the bigger models versus nano? What kind of signs are they looking for when that happens? Normally, they're going for more accuracy or they've acquired more data and what we call the data flywheel.
25:13So the second you put a model out, and actually sometimes it goes in reverse. So you put a more accurate model out, the bigger model, and then once you get your data, then you start to be able to fine tune and then put the smaller model out. So it's almost sometimes goes in reverse. But for people that are just starting out with the nano and want to go larger, it's really about understanding your use case, understanding how to evaluate it. Most businesses actually kind of get stuck in this evaluation phase. It's really hard. You know, when you start a proof of concept, you have to really narrow down the scope.
25:44It's like any other problem, like how will you test it? How will you verify it? And so once you get data, once you have your evaluation metrics and your criteria and your use case, then you can start to expand your use cases and expand your accuracy and start to grow your model needs, especially your specialization needs. We see a lot of customers, if you don't mind me to kind of digress for a second. No, sure. It's kind of like these eras since the big bang of ChatTPT in 2020, right? Enterprises were like, oh, what is generative AI? What do I do with it? Like, how do I use it? Like, I need something similar to that.
26:19And they're like, oh, well, I actually don't have the resources to train my own model. So I'm just going to use the manage APIs. And then everyone kind of went on that proof of concept. And then in 2023 was all about that. And 2024 was about, if you've heard of RAG, retrieval augmented generation. like okay they're like oh you know models hallucinate well now we need rag rag's my silver bullet like i'm gonna all i need is rag and so they started testing with rag and they realize it's it's just you you kind of can you only do so much with it and and then they're like oh you know what i might need to customize my model so they came back to the conversations that we had in 20 the beginning of 2023 we're saying we really need to specialize and like no that's too much So this is what I was talking about earlier is this journey of understanding.
27:06And 2025 has, you know, past RAG, 2025 has been all about agents. And because if you, you know, when you started to do RAG, you're like, well, I really need this to be to route to the right systems. I need it to be able to reason about what data I go pull. It just really needs to be smarter. And so and then RAG became a tool of agents. Instead of everything being about RAG, it's now RAG is just like a small tool that an agent can call to. So that's been the journey.
27:33Kari Briski:I have a question kind of related to that, which is when you work with these companies, are they hiring ML engineers to do this stuff for them? Or are they trying to learn it on their own? I think that's kind of interesting. I think a little bit of both. A lot of the large enterprises are starting their own center of excellence in their own internal platforms. A lot of it is sort of wrappers around managed APIs and then a little bit of open source models, again, doing smaller tasks. And obviously, Llama was a big success last year. So a lot of enterprises started with Llama or Mistral models. I've seen it in production, too.
28:12And so now they're getting to the, I need specialization. We need specialization. Guess what? You need open data. Because we need to know how the models can form, what it was trained on, why it was trained on that. You can't put this into critical applications unless you can truly trust what it's been built on and kind of open up and see inside. That's right.
28:32Kari Briski:Right. So where do you feel like the Nemotron family really shines? And then I guess the flip side of that is where does it have room to grow, in your opinion? Yeah, I think it really shines in our efficiency. I'm really proud of the hardware architecture that we've done. Again, of course, we do. We build it for ourselves. We want to make sure we build the most efficient systems, and it's a spiritualist cycle. And I think it really shines on the reasoning, right? We have really great reasoning data that we've put out. We've even helped translate it to some other languages. And so I'm really proud of the team doing that.
29:06Room to Grow, I think if I had to be really critical on my own product, I think that right now we're working on two things. one is long multi-task complex steps so the more tools and steps that you have to take you can kind of get lost along the way and so you want to make we now we're doing more data collection for those types of scenarios so really complex tool calling is something that that we're that we want to do better on and then maybe this is a little bit visionary but you know we have a really great speech team, world-class speech team. We actually talk about open models. We've got really great speech models at top leaderboards and are open too.
29:50They're called Reva and Canary and Parakeet models. We'll pull them under the NemoTron umbrella and maybe start some more Omni models. So I'm kind of, you know, I'm excited about that. We did a call with someone talking about Parakeet just a couple days ago and how they use it for podcast transcripting and things like that and how great it is. Yeah. Yeah. Again, it's like these little known secrets. You know, people talk about Whisper and OpenAI Whisper, but parakeet is this really amazing model. And we also have things like magpie, these bird names for speech synthesis. So, yeah.
30:24Kari Briski:I like that. Yeah, it's very fitting for speech. I do too. Yeah. And what's great about parakeet and canary is I think both of those, there's versions that you can run locally too, right? That's right. That's right. So, Carrie, tell me, if you don't mind, what has it been like working for NVIDIA through this wild transition over the last few years of, you know, growing to essentially the biggest company in the world? Well, I've been here for a little over nine years, and it has definitely been a wild ride. You wake up every day super excited because you're working with the best people, the smartest people.
31:05honestly every morning is like Christmas morning because you wake up like oh my gosh yeah what's today I'm going to attack it so much to do so much to unwrap so much to build put together and at the same time Jensen is really a great leader a lot of people ask about that he he definitely grounds us make sure that we and the culture is to to work hard to be very transparent and like I said work on the hardest problems and then at the end of the day deliver the platform so which people can do their life's work or fulfill their life's work. So that's kind of the sentiment here. That's the culture.
31:40Work hard, do good. Yeah. We have a, we both regularly talk about, we have a lot of respect for Jensen as well and just kind of the foresight it took to sort of understand what was coming and be so beautifully positioned, just, you know, right in time to become the cornerstone of modern AI, essentially. Yeah, he always jokes. What does he say? It's like it's a 30 year overnight success.
32:09Kari Briski:I love that. I love that. That's really awesome. So what's what's your vision? Where do you where do you want Nemetron, especially Nano even, to be in, let's say, two years, like in everyday apps people are using? Yeah, I think that I like that you said everyday apps people are using. Because right now, again, I alluded to enterprises haven't quite taken generative AI at that last mile. So my vision and my hope is that when I go to the doctor's office or when I get in my car or when I use whatever everyday app I'm using, it becomes more specialized, becomes more AI infused, and it is based on a nematron.
32:55So I think of even things of like call transcripts from from meetings not just the transcript but specialized to me in my role at my company and the people that I work with and my acronyms right there's there's a lot of improvement that can be done there so I think that um for me it's that nematron will like be this base of specialization for enterprises and in every vertical domain yeah that's so cool even think about your banking app right and like I want to be able to talk to my banking app the way and engage with it the way you might engage with chat gpt or perplexity right so i yeah i wonder i want i want that capability in my everyday life you blew my mind i've never thought of ai in my banking app before i just you just want to talk to it right yeah and so and that's that and if and that you know we're talking about gym environments and and being able to verify things and if you think about the most success for for gender of AI right now, it's been coding apps, right?
33:58So these, the way you are able to talk and just type your intent to code, to write lines of code, you don't have to write lines of code anymore. And the reason why you can do that is because it's so easy to test, right? And you're so easy to verify it compiles or not. It's a yes or no. Did the code compile? All right, then there's some good code. Did it not? No. Right. So it's been really easy. That's why these apps are the first ones out there. They're really useful to developers. But if you think about it, I just want to be able to talk to the rest of my apps and my devices and in my applications in the natural, and not just talk, engage, right?
34:38In my own personal way. So if you think about like finance, retail, shopping, like that's definitely going to get there. Healthcare, automotive, of like in every industry energy like my my you know the um my solar panels on my report we like you know just to be able to engage with it instead of you know wonder so yeah and for it to smartly just lower my electric bill while we're at it and that's right hey while you're at it smartly
35:06Kari Briski:um yeah that's so cool like uh cory and i yesterday in a live we tested uh you know the Logan Kilpatrick at Google DeepMind just released like a yap to app is what they're calling it, where you can like talk directly to the app that you're building in in AI studio. And the thing that I was thinking is like, I want it to be able to almost like combine a speech model with computer use where I can say like, OK, open my, you know, let's search for, you know, in my bank statements, for example, like search for all of my deposits for this month. OK, show me that. All right. Now Let's look at my, you know, how many of my expenses were and like to be able to navigate pages with just chatting.
35:44Kari Briski:Yes. That would be awesome. Exactly. We need to get there. Right. Yeah. Yeah. I always joke. You should be able to click a button with talk, I think would be great. I always joke that my test is I want to be able to when I'm driving the car, holler at my phone to order me a pizza and have it at the door by the time I get home, you know, in 45 minutes or something. And it know what kind of pizza I want, where I would order from. Yeah. And my credit card and just, well. Read your mind. That's what I want. I wanted like the step right below reading my mind, you know? Yes. It kind of knows who I'm going to ask, but I still have the agency to decide for myself.
36:22Yes. Exactly. Exactly. So again, specialization, personalization, memory, all these things, systems are being built now. That's what this build out of this infrastructure is so that there's capacity to be able to deliver these types of applications in the way that they need to be built for every enterprise in the future.
36:42Kari Briski:That's awesome. Yeah, that's something people maybe don't, people are missing in the whole like data center debate. They're like, oh, are we overspending for this? It's like, well, if you think about all the things we could do, if we have more compute, you're kind of underspending at a certain level. Yeah, exactly. I just know where we are and we have so much more to do. That's awesome. So I'm excited. Well, we're sure anxious to keep following it. We've really felt for a while like Nemetron's been slipping under the radar a little bit. And when the opportunity came to be able to have a chat about it, we thought that was a great idea because we've been really impressed.
37:16Yes. Well, thank you for having me because I get excited to talk about Nemetron and the exciting things that we're doing here. So I appreciate it.
37:24Kari Briski:Yeah. Thanks so much for joining us today as well. Where can people go to learn more and find some good resources on Nemetron? Well, I think first and foremost, it's all over Hugging Face. So all of our data sets, assets, recipes, you can go to build.nvidia.com and just try it out for yourself. There's an interactive UI that you can use and download the models from there as well. And just at nvidia.com. If you've liked and enjoyed today's video, please take a moment to like and subscribe. We really appreciate the support. Make sure you check out the Neuron AI newsletter at theneuron.ai as well and join some 600 ,000 others who read us every morning.
38:04We really appreciate you listening. And that's it for today. Farewell for now, humans.
From the publisher
Learn how to use NVIDIA's Nemotron open-source AI models with VP Kari Briski. We cover what Nemotron is, minimum hardware specs, the difference between Nano/Super/Ultra tiers, when to choose local vs cloud AI, and practical deployment patterns for businesses. Perfect for anyone wanting to run powerful AI locally with full control and privacy.
Resources mentioned:
NVIDIA Nemotron Models: https://www.nvidia.com/en-us/ai-data-science/foundation-models/nemotron/
Start prototyping for free: https://build.nvidia.com/explore/discover
Subscribe to The Neuron newsletter: https://theneuron.ai
Watch more AI interviews: https://www.youtube.com/@TheNeuronAI
