In short
NVIDIA’s newly released Nemotron 3 Super (120B parameters, 10% active via mixture-of-experts) and why its hybrid Mamba+transformer design, latent MOE, multi-token prediction, and 1M-token context make it well-suited for agentic multi-agent systems (reasoning, tool use, long workflows).
Key claims
efficiency (compute like ~12B), practical 1M context (Mamba linear scaling with transformer retrieval precision), faster structured generation (multi-token prediction enabling up to ~3x wall-clock speedup), and improved multi-step research performance.
Notable examples
powers NVIDIA AIQ Research Agent to #1 on Deep Research Bench/Bench 2; ~450–480 tokens/sec; open weights + permissive commercial license; 10T+ datasets, 15 RL environments, NemoGym.
Guests
none mentioned; only host Jon Krohn and third-party companies/providers (Perplexity, CodeRabbit, Greptile, Siemens, Palantir, Cadence; Hugging Face, Google Vertex AI, Oracle OCI, AWS Bedrock, Azure; Base10, Deep Infra, Fireworks AI, Lightning AI).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Nemotron 3 Super
0:45 to 2:02
Exploration of Nemotron 3 Super's model specifications and efficiency.
“are active at any given time during inference.”
Hybrid Architecture Explained
2:02 to 3:41
Discussion on the hybrid architecture combining transformer and Mamba layers.
“But pure state-space models can struggle with precise retrieval tasks, finding one specific piece of information buried deep in a long context.”
Innovations in Prediction
3:41 to 5:00
Overview of novel techniques like latent MOE and multi-token prediction.
“These architectural choices together deliver impressive throughput numbers.”
Performance and Adoption of Nemotron 3 Super
5:12 to 8:14
Details on the performance benchmarks and adoption of Nemotron 3 Super in various industries.
“As companies move beyond simple chatbot interactions into multi-agent AI applications, they run into two major bottlenecks.”
Accessing and Utilizing Nemotron 3 Super
8:14 to 9:21
Information on how to access and utilize the Nemotron 3 Super model.
“companies like Siemens, Palantir, and Cadence are deploying it for manufacturing, cybersecurity, and semiconductor design workflows.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:This is episode number 976 on NVIDIA's Nemotron 3 Super. Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. Today's topic is NVIDIA's brand new Nemotron 3 Super model, which is a mouthful, and was announced to coincide with this week's big NVIDIA conference GTC. NEMOTRON 3 Super is an openly available model that deserves your attention not only because of its impressive technical specs, but because of what it signals about where the AI industry is headed, specifically toward agentic AI systems that can reason, use tools, and operate autonomously over extended workflows. So let's start with the basics.
0:43Jon Krohn:Nemotron 3 Super is a 120 billion parameter model, but only 12 billion of those parameters, 10%, are active at any given time during inference. This is because the model uses a mixture of experts architecture, or MOE for short, where different subsets of the model's parameters, the so-called experts, are selectively activated depending on the input. So you get the knowledge capacity of a 120 billion parameter model, but with the computational cost closer to a 12 billion parameter one. That's a massive efficiency win. And if you'd like to hear more about mixture of experts, refer back to episode number 778 of this show.
1:21Jon Krohn:But now, NemoTron 3 Super isn't just any mixture of experts model. It's built on a hybrid architecture that combines two fundamentally different approaches to sequence processing, the common transformer-based attention layers that predominate LLMs today, as well as relatively exotic, though increasingly common, Mamba layers. I did a whole episode on Mamba back in episode number 758, but quickly, Mamba is a so-called state-space model that processes sequences in linear time with respect to sequence length, which is way more efficient than the quadratic scaling you get with the traditional transformer self-attention.
1:57Jon Krohn:This is what makes Nemotron 3 Super's 1 million token context window practical rather than theoretical. But pure state-space models can struggle with precise retrieval tasks, finding one specific piece of information buried deep in a long context. So NVIDIA interleaves a small number of transformer-based attention layers at key depths to preserve high-fidelity information retrieval. It's a best-of-both-world design. Mamba for efficiency, transformers for precision. On top of this hybrid backbone, NVIDIA introduced a novel technique called latent MOE, latent mixture of experts. In a standard MOE setup, tokens are routed to experts in their full hidden dimension, which gets expensive as models scale.
2:41Jon Krohn:With latent MOE, tokens are first compressed into a smaller latent space before routing, which dramatically cuts computational overhead. The savings are reinvested to activate four times as many expert specialists for the cost of one in a traditional setup. More experts consulted per token means specialized knowledge being brought to bear on each prediction, and the data show this translates directly into better accuracy. There's one more architectural innovation worth highlighting, multi-token prediction, or TP. Standard language models predict one token at a time. Nemotron 3 Super predicts multiple future tokens simultaneously using specialized prediction heads.
3:22Jon Krohn:At inference time, these heads function as a built-in draft model for speculative decoding. You generate several candidate tokens quickly and verify them in a single forward pass. The result is up to a three times wall clock speedup for structured generation tasks like code or tool calls, and you don't need a separate external draft model to get it. These architectural choices together deliver impressive throughput numbers. On an 8 ,000 input token and 64 ,000 output token benchmark, NemoTron 3 Super achieves up to 2.2 times higher throughput than the comparably sized GPT-OSS-120B and up to 7.5 times higher throughput than QEN 3.5 at 122 billion parameters, while in both cases matching or exceeding on accuracy.
4:11On this podcast, I'm always going on about how Claude Code is mind-blowing, but now Claude Cowork is making my jaw drop as well. For example, I recently wanted to quantify how healthy my sales pipeline is for my AI consulting business. I simply asked Claude to estimate my sales for the coming quarter, and it brought info from relevant Google Sheets and my Gmail to create a professional spreadsheet of clients with estimated revenue for each one. Whoa, this might have taken me a day. Instead, it was done flawlessly with Claude Cowork in minutes. Claude is the AI for minds that don't stop at good enough.
4:42It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. Ah, and you'll appreciate that I can ask Cowork to show me data, such as my sales spreadsheet, and it provides an interactive chart right in the conversation. For problems worth solving, get started with Claude at Claude.ai slash superdata. That's Claude.ai slash SuperData. And check out Claude Pro, which includes access to all of the features mentioned in today's episode.
5:11Claude.ai slash SuperData.
5:15Jon Krohn:The model was also pre-trained natively in NVIDIA's 4-bit NVFP4 precision, which on Blackwell GPUs pushes inference up to four times faster than FP8 on the previous generation Hopper GPUs, and again, with no loss in accuracy. Now, why does all of this matter? As companies move beyond simple chatbot interactions into multi-agent AI applications, they run into two major bottlenecks. The first is what's called context explosion. Multi-agent workflows can generate up to 15 times more tokens than a standard chat because each interaction requires resending full histories, tool outputs, and intermediate reasoning.
5:53Jon Krohn:This ballooning context increases cost and can cause goal drift where agents gradually lose alignment with their original objective. Nemotron 3 Super's million token context window lets agents retain the full state of a workflow in memory without truncation. Cool. The second bottleneck is the thinking tax. Complex agents need to reason at every step, but deploying a large, expensive model for every subtask requires or makes multi-agent pipelines too slow and too costly. Nemotron 3 Super's combination of sparse mixture of experts computation and Mamba-based efficiency is aimed squarely at making step-by-step reasoning affordable at scale.
6:33Jon Krohn:That's its core value proposition, frontier class reasoning at a fraction of the typical compute cost. And did it all work? Yes, it did indeed. The benchmark data bear this out. Nemotron 3 Super currently powers the NVIDIA AIQ Research Agent to the number one position on both the Deep Research Bench and Deep Research Bench 2 leaderboards, which measure multi-step research capability across large document sets. The model has also claimed the top spot on artificial analysis for efficiency and openness in its size class, outputting tokens at around 450 to 480 tokens per second depending on the provider.
7:08Jon Krohn:Speaking of openness, as I briefly mentioned at the top of this episode, NVIDIA is releasing the model with open weights under a permissive commercial license. But they went further than just releasing weights. They're also publishing over 10 trillion tokens of pre - and post-training datasets, 15 reinforcement learning training environments, and their full evaluation recipes. For researchers and practitioners who want to reproduce the training, fine-tune for a specific domain, or build their own hybrid architecture models, these data and recipes are invaluable. The model was post-trained using reinforcement learning across diverse agentic environments via NVIDIA's open-source NemoGym library, which evaluates the model on sequences of real actions, tool calls, functional code generation, verifiable multi-step plans, rather than just optimizing for single-turn responses.
7:54Jon Krohn:And to coincide with the launch, there were companies evidently working in the background to make sure that NVIDIA could announce that adoption is already picking up. Perplexity, for example, is offering Nemotron 3 Super for search. Software development agent companies like CodeRabbit and Greptile are integrating it into their coding assistants. And on the enterprise side, companies like Siemens, Palantir, and Cadence are deploying it for manufacturing, cybersecurity, and semiconductor design workflows. In terms of where you can access the model, weights are on Hugging Face for self-hosting. For cloud deployment, it's available through Google Clouds, Vertex AI and Oracle Cloud Infrastructure with Amazon Bedrock and Azure reportedly coming soon.
8:35Jon Krohn:On the inference side, it's available through providers including Base 10, Deep Infra, Fireworks AI and Lightning AI where full disclosure, I hold a fellowship. So despite not being an objective information source, I can nevertheless provide objective third-party data from artificial analysis showing that Lightning AI at the time of me recording delivers the fastest Nemotron 3 Super output speed of any inference provider, coming in at 480 tokens per second. I've provided a link to this in the show notes, plus anything else I cited in today's episode. So what this means is if you don't want to go through the hassle or expense of setting up Nemotron 3 Super on your own infrastructure, working with an inference provider like Lightning will make your life super easy and you can just get going with this innovative mixture of experts model today.
9:21Jon Krohn:If you're building multi-agent systems, whether autonomous coding assistants, research agents, or enterprise automation workflows, a model like this that combines open weights, frontier class reasoning, and blazing fast throughput at a fraction of typical compute costs is exactly the kind of tool that can take your project from prototype to production. All right, that's it for today's episode. If you enjoyed today's episode or know someone who might, consider sharing this with them. Leave a review of the show on your favorite podcasting platform or on YouTube. If you tag me in a LinkedIn post with your thoughts, I will respond to those.
9:55Jon Krohn:And if you aren't already, of course, subscribe to the show. Most importantly, however, we hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
NVIDIA just dropped Nemotron 3 Super, a 120-billion-parameter open-weight model that only activates 12 billion parameters at a time and it’s built for the agentic AI era. In this Five-Minute Friday, Jon Krohn breaks down the model’s hybrid Mamba-Transformer architecture, its million-token context window, and why its combination of frontier-class reasoning with blazing-fast throughput matters for anyone building multi-agent systems. Find out how Nemotron 3 Super claimed the #1 spot on the DeepResearch Bench leaderboards, which companies are already adopting it, and where you can start using it today.
Additional materials: www.superdatascience.com/976
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.




