1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnell

18 Sep 2026 · 30 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How agentic AI changes inference workloads and why SambaNova’s RDU chip architecture targets faster, higher-throughput token generation by removing GPU decode-phase memory bottlenecks; includes deployment and ROI claims for the SN50.

Guest backgrounds

Anton McGonnell, VP of Product at SambaNova (Bay Area; founded from Stanford research; raised $2B+). Focuses on inference-optimized AI chips and systems.

Key claims

Agentic AI needs larger KV caches and shifts the bottleneck to decode, where GPUs are “kernel-by-kernel” and bandwidth-limited. SambaNova’s data-flow-based RDU overlaps communication and computation, scales linearly across chips, and enables faster payback. SN50 customers can recoup investment in ~6 months (claimed ~10x ROI), then profit for 5+ years.

Notable examples

GPU Blackwell/Rubin-style all-to-all scaling across 72 chips; SambaNova deploying lighter, air-cooled racks into brownfield/retrofit data centers; software entry points via model deployment, BYOM using PyTorch/VLLM, and research tooling for agent-driven optimization.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to AI Inference Challenges

0:00 to 0:15

Learn about the limitations of GPUs in AI inference and the new chip designed to address these.

“GPUs are today, of course, the workhorse of AI inference in production, but they aren't actually optimized for that.”

Overview of SambaNova's Mission

0:54 to 1:50

Anton discusses the mission of SambaNova and its focus on improving AI inference.

“Anton, welcome to the Super Data Science Podcast.”

Understanding Inference Workloads

1:50 to 2:37

Explore the difference between training and inference workloads in AI models.

“And so I suspect our core audience is very familiar with the idea that when you're creating your own large language model, your own AI model, your own foundation model, whatever it does, you do lots of training cycles.”

The Rise of Agentic AI

2:37 to 3:20

Anton shares insights on how Agentic AI has changed inference demands.

“And particularly in the past 12 months, why is that, Anton?”

SambaNova's Focus on Inference

3:20 to 4:49

Discussion on why SambaNova specializes in inference over training in AI.

“because we've got so many problems to solve.”

Speed vs. Throughput in AI Chips

4:49 to 6:04

Examining the balance of speed and throughput in SambaNova chips compared to competitors.

“Yeah, I mean, traditionally we did both training on inference.”

Deep Dive into RDU Architecture

6:04 to 8:38

Anton explains the innovative RDU architecture that enhances AI inference.

“And I somehow had never thought of that myself.”

Overcoming GPU Limitations

8:38 to 11:28

The discussion focuses on how SambaNova chips address the GPU bottlenecks in AI workloads.

“Yeah, I mean, it's a combination of those things.”

Deployment and Infrastructure of RDUs

11:28 to 14:00

Anton discusses the practical aspects of deploying RDU hardware and its advantages.

“So this RDU, reconfigurable data unit, I guess it builds on what you were talking about nine years ago is this approach to having really refined data flow.”

Setting Up SambaNova RDUs

14:00 to 18:18

Learn how SambaNova RDUs are deployed and utilized in data centers.

“Is there some kind of equivalent to a CUDA library?”
Show all 14 chapters

The Economics of the SN50 Chip

18:18 to 26:11

Understand the financial benefits and ROI of the SN50 RDU chip.

“So thanks for giving us that tour of both how you get things set up from a hardware side and how you use Samanova RDUs from a software side.”

Customer Segments and Usage

26:11 to 27:55

Explore the different types of customers utilizing SambaNova's technology.

“and the advantages that that has over other ways that you can be doing inference compute, particularly in this agentic AI era that we're now in.”

Engagement on Social Media

28:03 to 28:39

Learn about Anton McGonnell's presence on social media platforms.

“So if you search for my name on LinkedIn, high probability you'll find me as well as on X where I spend a lot of time reading and sometimes commentating.”

Key Takeaways from the Episode

28:40 to 29:11

Explore the main insights from Anton's discussion on Agentic AI and SambaNova.

“I hope you enjoyed this hardware, AI hardware conversation.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:GPUs are today, of course, the workhorse of AI inference in production, but they aren't actually optimized for that. Today's guest built a new chip that is pushing the frontier of real-time AI speed and bandwidth. Welcome to another episode of the Super Data Science Podcast. I'm your host, Jon Krohn. My guest today is Anton McGonnell, VP of Product at SambaNova, a Bay Area company that has raised over$2 billion to develop a chip optimized for AI inference. In this episode, Anton explains why Agentic AI has changed the shape of inference workloads and how SambaNova's chip sidesteps the memory bottleneck that slows GPUs when generating output tokens, making SambaNova the most profitable way to run AI hardware.

0:49Jon Krohn:Enjoy this informative episode.

0:54Jon Krohn:Anton, welcome to the Super Data Science Podcast. Thank you for joining us today. Where are you calling in from? Thank you, John. I'm calling in from our office in San Jose. And our is SambaNova, right? Do you want to tell us a bit about what SambaNova does? Yeah. So SambaNova is an AI chip and systems company. We've been around for nine years now. It was founded out of Stanford University, many years of research there, in the exploration of better approaches to solving data flow-based computational problems, of which AI has become the premier one of today. What else even is there? Yeah, there used to be other things, believe it or not.

1:35Jon Krohn:All right. Well, so one of the key things that Samba Nova seems to be solving is dealing with inference time compute for AI workloads, right? So let's provide a bit of context to the audience on what that means. And so I suspect our core audience is very familiar with the idea that when you're creating your own large language model, your own AI model, your own foundation model, whatever it does, you do lots of training cycles. So you can have chips that are specialized to training, or there are also chips that are made to do both to kind of handle training workloads as well as inference workloads.

2:14Jon Krohn:and those inference workloads are after you have your model trained up, you want to be using the model in real time. So some input goes in, some output comes out, the model weights don't change. You just get some kind of AI model result. And so probably most of our audience already knew that, but just to get that in their minds, what people might not realize is that the inference market is changing fast. And particularly in the past 12 months, why is that, Anton? Yeah, I mean, I think the big catalyst has been agentic AI. Obviously, for the past three or four years now, there's been a huge uptick in AI workloads.

2:54Slowly, they've transitioned from not just training, but in front of us, people start to use AI for useful stuff. What's really happened in the past, I would say nine months in particular, is the usefulness of these agentic systems, which is really just using AI to automate work or automate tasks or automate discoveries and research. The usefulness of this has accelerated to such a degree that we cannot possibly have enough capacity because we've got so many problems to solve. And as a result of all of these problems to solve, we need more inference to be able to help us solve these problems. And the interesting thing about agentic AI relative to the workloads that came before it is that it is a different computational profile.

3:46It's got larger inputs, much larger inputs, but a high degree of reusability of those inputs. So even though there's lots of input, you're only computing a small amount, but the amount of data that you're then feeding to generate output tokens for the decoding phase of inference, you're dealing now with much larger KV caches. So for a lot of reasons, the systems that were really good at the first iteration of generative AI inference are not necessarily the ones that are very good for agentic AI inference.

4:22Jon Krohn:Right. That makes a lot of sense. And I guess SambaNova's chips specialize in the inference phase, which you probably know this stat better than me. But my understanding is that chips actually rarely get used for training. Chips are mostly used in the AI workload sense for inference. Like 99 % of the time or more, chips are being used for inference. So yeah, let us know about whether I'm right about SambaNovi being specialized for that big chunk of the AI workload market. Yeah, I mean, traditionally we did both training on inference. We started to specialize in inference to, frankly, lower the aperture of work and focus our resources on what is a market that is becoming much larger relative to training, as you mentioned, but also where the barriers to entry are lower.

5:11it's less concentrated market whereas in training there there's just few buyers they buy large amounts of compute but but you know nvidia frankly is that market quite quite locked in so the inference market strategically is is better for non-incumbents but then also just the the profile of our chip and what it's good at lends itself really well to be able to really differentiate on inference, which we've sort of, you know, over time that has become more true. And just on your data point, even in the training phase, there is more and more inference. There's more need for inference compute for training now than there is for actual backpropagation for the actual training part of training.

5:55Right. I had never thought of that. Yeah.

5:58Jon Krohn:So much of the post-training compute is actually just inference compute. Right. Yeah, that's a really good point. And I somehow had never thought of that myself. So what makes SambaNova chips different? You claim that they're the fastest inference chips, but I'm sure a lot of companies claim that. How do you compare yourself with other chips available out there and how do you achieve those differences? Yeah, speed is relative, right? For end users, speed really matters in terms of how quickly they get their output tokens. For service providers that are buying systems, our customers, the speed that matters is their payback.

6:43So they buy a big system or they buy many systems. They want to know with this big capital investment that they're making, how quickly are they going to make a return on their investment? And that's where, from our perspective, our chips are incredibly differentiated because we provide the speed, the token generation speed, but we balance it with really, really good throughput. And this really is the juxtaposition that AI inference finds itself in, where if you want to give users more speed, which they really need, because Agendic AI is all about automating work, the more work you can automate, the faster you can automate it, the more you can differentiate from your competitors and move faster than your competitors.

7:28But the more you do that, the fewer concurrent users you can actually run on each chip. So this is sort of why we see these Pareto curves, where on the Y axis we see throughput per chip, and on the X axis we see speed per user, and those things are at odds with each other. So what we do incredibly well is balance both of these things, give you really, really fast speed, with really, really great throughput, such that as a service provider, you're able to charge more for your tokens because you're serving them to your customers faster, but also serve more and more and more users such that every second you're making more money.

8:10You're making more money and you're making more money per token.

8:15Jon Krohn:Right, so you're getting speed and throughput on a single rack. I'm starting to understand this and the differentiator there for SambaNova being on that part of the Pareto curve where you're maximizing kind of both of those things as much as you can at the same time. But how do you do it? Like what is different about a SambaNova chip? Is it the memory, the compiler? Is it the chip itself? Yeah, I mean, it's a combination of those things. It's hard to disentangle at all because the right compiler is only right for the right chip and the right workload is only right for the right compiler. Really what it comes down to is the pre-fill phase of inference, which is processing the input tokens is highly parallelizable.

9:01It runs well on GPUs. But the bottleneck is the decode phase where you're generating output tokens. And that's where the GPU's architecture is just ill-suited. The reason being it's a kernel-by-kernel execution machine. It breaks the model's computational graphs into these chunks that they call kernels, and they execute the computation of that kernel. And then they feed it back to their external HBM memory. They load the next kernel, and they do it again. And so it's very, very time-consuming, and they're bottlenecked by their memory bandwidth. Our architecture, the RDU's architecture, is a data-flow-based architecture where we just...

9:45Jon Krohn:If you don't mind me quickly interrupting you, I think a key thing here, and I can't believe I didn't mention this earlier, is that instead of calling it a GPU, like an NVIDIA GPU, you call your processor an RDU, or I believe it stands for reconfigurable data unit. And so our listeners would use that in lieu of a GPU, right, at inference time with their AI workload? They could use it alongside it. They could use it in lieu of it, or they could use it alongside it. I see. The GPU being really good at pre-fill at the input token processing, but insufficient for the decode, the output token generation phase.

10:24So yeah, the reconfigurable data flow unit is our chip. And what makes it really different is you're not executing all of these kernels, these slices. You are loading the entire model, rolling it out spatially across the chip and letting data flow through. such that you're never having to do all of this back and forth with your external memory, nor are you having to have all of this overhead of communication overhead between each chip. So really what the RDU, what Sanbenum's architecture enables is the ability to perfectly overlap communication and computation. And this is very, very, very important for the decode stage of inference because time is so finite.

11:15All of these operations are happening at such a granularity that any overhead in computation or communications will just automatically become a big bottleneck and that'll slow down your ability to generate tokens.

11:31Jon Krohn:Right, gotcha. So this RDU, reconfigurable data unit, I guess it builds on what you were talking about nine years ago is this approach to having really refined data flow. And so it's kind of like data moving like an assembly line through the processor that allows you, and it's kind of that specialized use that allows you to get such throughput at such high speed. Yeah, yeah. And really hiding all of that overhead. And the other thing is, it's not just on one chip. It's across many chips. So a lot of the overheads that the GPU will have to deal with for the decoding phase comes from the fact that they need to parallelize over many chips.

12:23NVIDIA created the newer version of their system, the Blackwell system and their upcoming Rubin system, has a single all-to-all communication between 72 chips in a rack. The reason they do that is because they need more chips to be able to run bigger models, longer contacts lengths, with more throughput and more speed. The problem is, as they use more chips to run these models, the overhead of the communication between the chips really exasperates their inefficiencies. So by virtue of our architecture, we are able to scale linearly. So whenever we can be twice as fast on two chips as we can be on one chip.

13:07and four times as fast as four chips. So that ability to scale linearly is really an incredibly important characteristic of the architecture.

13:19Jon Krohn:You should have led with that, Anton. That's brilliant. That's so easy to understand and how that differentiates you against your competitors. So yeah, so that's the whole idea of this assembly line. You can break up loads and have them sent down each of the assembly lines that you have. the more RDUs that you have, the faster the inference time compute, the more throughput that you have. That is pretty cool. Talk us through kind of the, you know, how this works. Like if we have listeners out there today who are like, this sounds awesome. I'm doing lots of inference time workloads. I'd love something that's more efficient than what I can get from the, you know, the well-known incumbents out there.

13:59Jon Krohn:How do they do it in terms of hardware? How do they get it set up? in terms of software? Is there some kind of equivalent to a CUDA library? Or how do they run and process on RDUs? So it starts with how easy is it to deploy the hardware? And another consequence of our architecture is that we don't need to densely pack our racks with 72 chips. We keep fewer chips in a rack such that the rack can be lightweight and air-cooled. So by virtue of doing that, now we can go into a lot of these existing brownfield data centers or retrofitted data centers that were previously telephone exchanges or data centers that ran CPUs, for instance, that don't have the same needs for power delivery, cooling, et cetera.

14:56So we are in a world right now where perhaps the biggest bottleneck is these net new data center build-outs. And that's going to take some number of years to play out and for all of that data center capacity to come online. So we can be deployed into existing brownfield data centers today where other providers can't do this because they're having to build very dense racks. So that's the first big unlock. Just the ability for customers, whether it be enterprises or new clouds and service providers and sovereign clouds, if they're able to secure existing brownfield data center capacity, they can roll in RACs and get it up and running same day as it rolls in.

15:40In terms of then the software stack, we have specificities about our architecture that will see us mapping the model differently. We don't map the model the same way the GPU does because it's a different architecture. Our benefits are different. But all of that is abstracted from the user. And there's really a couple of different entry points that a user can come in at. The first of which is if they just want models, they know which models they want and they want them to run fast, they can click a few buttons and deploy that model and start serving inference. Then you have customers that want to bring their own model.

16:22Maybe it's not a model that we support out of the box, providing them an ability to do that via some abstraction that they already know, like PyTorch and VLLM. We provide that capability as well. Then for researchers that are really trying to figure out how to better utilize this architecture in ways that we haven't thought about, for instance, in Sam & Nova, because we need to be focused on our customers and our immediate-term problems, being able to democratize RDU programming so that anybody can do it, that's what we're now enabling, making it easier to interact with the lower levels of our stack.

17:02And the reason that's become so important is because now AI agents can do this. You can just point AI agents at our software and let them iterate and try to bring these models up and map them and make them run really fast and performantly.

17:18Jon Krohn:Regular listeners will already be aware that I'm obsessed with Anthropik's Fable 5 model and it has taken over my working life. I'm writing a technical book that includes LaTeX files, mathematical notation, Python code examples, and Fable 5 and Cloud Code handles requests I make across whole chapters with accompanying Jupyter Notebooks end-to-end, work that a few short months ago would have been dozens of separate requests with way more manual fiddling required. With Fable 5, it just works, essentially like magic, first time. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you.

17:55Jon Krohn:Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai slash superdata. That's claude.ai slash superdata. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash superdata. Nice, all right. So thanks for giving us that tour of both how you get things set up from a hardware side and how you use Samanova RDUs from a software side. My understanding is that the latest generation of your SambaNova RDU is called the SN50.

18:39Jon Krohn:It's your fifth generation chip. And I guess you're shipping them to customers this quarter. Tell us about the economics of that chip and how it compares to your competitors. How long would it take for somebody who buys an SN50 RDU to make their money back? Yeah, so SN50, if a customer buys a cluster of SN50s and runs inference on these systems, even on the latest and greatest open source models, will get their money back within six months. And that is really unprecedented within this industry for something that, as a substantial capital investment, to then get that money back within six months and to be able to keep that system in production for five years, at least five years, and in some cases longer.

19:40your minus, you know, your, your small amount of OPEX, everything after that six months is pure profit. So this is an incredibly profitable machine, uh, for, for any service provider. It's like a 10 X ROI where it takes you six months to pay off the purchase.

19:58Jon Krohn:And then the remaining, yeah, you have nine times more time, you know, assuming that they're keeping it online for five years to be bearing the fruits. Basically, as you say, a small amount of OPEX, but almost cost-free. Exactly, exactly. So typically within this industry, a two to three year payback period is the expectation. To be able to collapse that down into six months really is just a huge unlock for all of these service providers, who, by the way, are already making money because they're running inference at scale. we're in a capacity constrained world because inferences there's we just have such a need for inference and to be able to get net new inference capacity but also to be able to make such an incredibly attractive return on that investment is why our business is just just really really really accelerated over the past nine months it's pretty remarkable can you tell us a bit about what your typical buyer is like who's your typical customer yeah so we have a number of different segments who have similar needs, but their own specifics.

21:10The NeoClouds, this emerging class of customers called the NeoClouds are... Sure, like Lightning AI? Like Lightning AI. There's a broad scope, I would say, of NeoClouds as a category. There are those that are focused on software and optimizing software and leasing hardware capacity from others. And there are those that are building capacity directly, building data centers, making capital purchases of systems, and then they're leasing that either directly to enterprises or consumers or leasing that capacity out to the software-focused neoclouts. For better or worse, the industry has coalesced around calling all of these neoclides.

21:58But I would say more specifically, when I refer to neoclides as direct customers, there are those that are building data centers, building racks, managing the data center operations, because they're the ones that actually make the capital purchases. The other class of customers is sovereign clouds. They have differences, and in some cases are neoclouds themselves, but some differences than the more general neoclouds. And then we have enterprises. And enterprises have emerged as a surprisingly latecomer to this category because so much of AI inference offtake has come from these AI natives and startups.

22:46but any enterprise in the world that is not wholly embracing AI inference today is going to be left behind. So the risk for them is company defining and they are now wholly and readily embracing AI inference, not just consuming cloud APIs, but starting to think about their AI sovereignty, like owning their AI and not being reliant on third parties to control their destiny. So those really are the categories. Really, they're all service providers. Even the enterprises themselves are service providers. They're just service providers for internal users. But loads of similarities, but then lots of nuanced differences too.

23:32Jon Krohn:Yeah, it makes perfect sense. My last technical question for you, and this is something that actually we're kind of getting right into perfectly before I started asking you about the different kinds of customers you have, is that inference compute is scarce. You know, the more that inference becomes available, the more usefulness that we find, and the more demand there is for even more inference. And so because of that, compute is scarce. What is the cost of running the wrong AI inference workload on general purpose capacity? So probably most of our listeners are using general purpose GPUs for their inference time.

24:13Jon Krohn:What's the cost to them of doing that? I mean, ultimately, the cost is time and money. We are inhibited by a couple of different things today. Number one is we have finite amount of energy that we can deliver to AI systems to run inference. You want to get as much out of every watt as possible. You want to be able to generate as many tokens for each watt of power that you deliver. Then you have space, which is a constraint and all of the infrastructure needed to support running these systems. Then you have, and that's sort of the infrastructure and service provider layer. They're the big constraints they have.

24:57Capital is another big constraint. More and more that's being solved through financing products, but still a real constraint. As you move up the stack, the constraints become, what tokens am I consuming? Because not all tokens are created equal. You want your tokens to be high quality from a really, really good model, but you also want them to be very fast. And we recently ran a survey with AI infrastructure leaders. Four out of five of them said they would be willing to pay a premium for faster inference. The reason they're willing to pay that premium is because they recognize that faster tokens are better tokens.

25:38It allows them to out-compete their competitors. If you are spending time and money on tokens that are low quality and low speed, then that's opportunity cost. So the answer really depends on the layer of the stack, but those are the constraints across the market today. And frankly, we think we are very, very well positioned across all of those constraints.

Read the full transcript

26:03Jon Krohn:Nice. Well, thank you for the introduction to SambaNova chips, to the SN50 RDU in particular, and the advantages that that has over other ways that you can be doing inference compute, particularly in this agentic AI era that we're now in. Before I let you go, Anton, I think you've been prepared for this, although I forgot to tell you myself, which is that we usually ask for a book recommendation at the end of episodes. Do you have anything for us? I unfortunately exclusively read classical fiction. So I don't know how relevant it is for this podcast. Our audience loves all the recommendations.

26:42Jon Krohn:Hit us up. Good, good, good. I read a lot of books. The latest book that I finished was a book called Lonesome Dove, which is about a very interesting time in American history after the Texas Rangers had spent a lot of time fighting the bandits and trying to find purpose. The interesting takeaway of that book actually is that sometimes maybe you've solved a big problem in your life and you're very proud of yourself, but we have this strange insatiable desire to just go and solve more problems, sometimes without understanding what problem we're trying to solve. And I think it's an important lesson for all of us working in AI these days that it's incredible the amount of progress we've made.

27:35But we need to keep sight of what the end goal is and make sure that it's all for the right reasons.

27:44Jon Krohn:Nice. Love it, Anton. Thanks for that great recommendation. Nice takeaway from the book as well. Thank you. My final question for you is how people can keep up with your thoughts or on Sambanova after this episode. Yeah, I'm lucky enough to be the only Anton McGonnell in the world. So if you search for my name on LinkedIn, high probability you'll find me as well as on X where I spend a lot of time reading and sometimes commentating. Nice. And SambaNova more broadly is very active on LinkedIn, X, and all of the social media platforms. Perfect. Well, yeah, we'll have links to all of those in the show notes for our listeners.

28:31Jon Krohn:Anton, thank you so much for taking the time with us today. And I can't wait to see what comes out of SambaNova next. Thanks so much, John. Great episode today. in it, Anton McGonnell detailed why Agentic AI has changed the computational profile of inference, the trade-off between speed per user and throughput per chip, and how SambaNova's RDU aims to deliver both by laying the whole model out spatially across the chip. He talked about how that architecture scales linearly, so two chips are twice as fast as one and four are four times as fast, while the SN50 rack stays lightweight and air-cooled enough to be rolled into retrofitted data centers the day it arrives.

29:11Jon Krohn:I hope you enjoyed this hardware, AI hardware conversation. To be sure not to miss any of our exciting upcoming episodes, subscribe to this podcast if you haven't already. But most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

29:38Thank you.

From the publisher

In Episode #1028, Anton McGonnell (VP of Product at SambaNova) joins Jon Krohn to explain why the chips running most AI inference today were never designed for the job. Agentic AI has changed the computational profile of inference, with much larger inputs and far heavier caches feeding the token generation that follows, and that shift has exposed where GPU architecture struggles. SambaNova has raised over $2 billion to build an alternative, the reconfigurable dataflow unit, which lays a whole model out spatially across the chip rather than executing it kernel by kernel. In this episode, Anton discusses why the speed that matters is payback, and how speed and concurrency are what turn a fixed hardware cost into a six-month payback. He also walks through the trade-off every inference provider faces between speed per user and throughput per chip, what the RDU architecture changes about scaling and data center deployment, the economics of the new SN50, and why four out of five AI infrastructure leaders say they would pay a premium for faster tokens.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1028⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(00:02:41) Why agentic AI is reshaping inference workloads

(00:08:39) How SambaNova's RDU differs from a GPU

(00:17:54) The economics of the SN50

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnellSuper Data Science: ML & AI Podcast with Jon Krohn · 30 min
Listen in VO