939: Mixture-of-Experts and State-Space Models on Edge Devices, with Tyler Cox and Shirish Gupta

11 Nov 2025 · 1 h 6 min · 24 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How Dell’s on-device AI stack (rebranded from “Dell Pro AI Studio” into the “Dell AI Factory”) makes running modern LLMs locally/at the edge easier, including model catalogs, on-device frameworks, model management, and enterprise fleet deployment via Intune/Dell Management Portal. It also explains why hybrid LLM architectures like IBM Granite 4.0 (state-space + transformers, plus MoE variants) matter for long-context, memory-efficient inference on AI PCs.

Guests (Dell)

  • Tyler Cox: Distinguished Member of Technical Staff / Distinguished Engineer at Dell; works on on-device AI for Dell client solutions (AI PCs and workstations).
  • Shirish Gupta: Dell colleague; co-led the earlier “Dell Pro AI Studio” effort and discusses enterprise deployment and performance.

Key claims

  • On-device AI provides hybrid optionality, lower latency, privacy (queries processed locally), and resilience to cloud/network outages.
  • Dell’s solution abstracts hardware/runtime diversity (multiple backend runtimes) behind consistent APIs and enterprise-grade lifecycle management.
  • IBM Granite 4.0 models are Apache 2.0 open-weight, hybrid state-space/transformer, and include MoE variants for better cost/throughput.

Notable examples

  • Quality-of-tokens: run smaller models for code autocomplete on PCs, larger models for full refactors.
  • Benchmarks: Dell Pro Plus (Intel Core Ultra 200 “Lunar Lake”) vs non-AI PC claims include 88% more Teams battery runtime, 4.8x graphics performance, and “>6x” Mistral 7B/Llama 3.1 textgen plus “up to 5.6x” Stable Diffusion 1.5 imagegen.
  • Granite 4H memory example: ~15GB vs ~80GB memory for 8 sessions at 128K context (micro model).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Welcome and Guest Introductions

0:32 to 0:45

Hosts welcome Tyler Cox and Sharish Gupta to the podcast.

“This episode of Super Data Science is made possible by Anthropic, AWS, and Garobi.”

Background on Guests and Their Work

0:45 to 2:06

Discussion about the backgrounds of Tyler and Sharish at Dell.

“Tyler, where are you calling in from today?”

Dell Pro AI Studio Overview

2:06 to 4:20

Sharish explains the development and rebranding of Dell Pro AI Studio.

“And so we are going to be going into a really deep technical episode because you're here.”

Challenges in On-Device AI

4:20 to 7:25

Overview of challenges faced in deploying AI on device and how Dell addresses them.

“And why that's important is in the construct of the Dell AI factory, the toolkit, formerly known as Dell Pro AI Studio, it really extends the capabilities of the Dell AI factory to our entire Edge and PC portfolio.”

AI PC Solution Components

7:25 to 11:15

Discussion on the components of the new AI PC solution from Dell.

“There wasn't at that time a really good solution that would abstract away all of that engineering backend complexity for developers.”

Enterprise Deployment Needs

14:00 to 14:32

Learn about the essentials for deploying AI workloads in enterprises.

“that enterprises would deploy through their existing app management consoles.”

Enterprise Deployment Needs

14:40 to 15:19

Learn about the essentials for deploying AI workloads in enterprises.

“AWS Tranium 2 instances deliver 20.8 petaflops of compute, while the new Tranium 2 Ultra Servers combine 64 chips to achieve over 83 petaflops in a single node, purpose-built for today's largest AI models.”

Importance of Local Workloads

15:19 to 16:06

Explore reasons for running AI workloads locally and their advantages.

“Now, so we talked about this a lot in episode number 921 with you, Sharish, and Ish.”

AI Workloads and Cost Management

16:06 to 22:37

Understand the dynamics of AI inference costs and model selection.

“But just really quickly in like a minute, if one of you wouldn't mind kind of providing an overview of why we'd want to do that.”

Privacy and Reliability in AI

22:37 to 23:25

Examine how local AI can enhance data privacy and reliability.

“Even though there is a, you know, a database retrieval request going through the network, the input and the output from the model stays completely private to the user on the box, right?”
Show all 24 chapters

AI PC Solutions from Dell

23:25 to 24:24

Learn about Dell's innovative AI PC solutions and their implications.

“And I had to go around on foot everywhere for a couple of days, which was weird for me because of that AWS outage.”

Integrating AI Technologies

24:24 to 28:00

Discover how integrating various technologies simplifies AI application.

“Tyler, distinguished engineer, how is it all possible?”

Optimizing AI Applications for Diverse Platforms

28:00 to 28:54

Learn how to streamline AI model deployment across various platforms using best practices.

“and pick and choose different versions of the application to land on those different platforms.”

Dell and IBM's Granite Model Partnership

28:54 to 30:28

Discover the collaboration between Dell and IBM to enhance AI model capabilities.

“So in addition to all those kinds of features, let's talk about particular models that might be interesting.”

Deep Dive into Granite 4 Model Variants

30:28 to 35:34

Explore the different variants of the Granite 4 model family and their applications.

“Because yes, one, they're open source under the Apache 2.0 standard license.”

Understanding Mixture of Experts Models

35:34 to 36:24

Learn about the mixture of experts approach and its advantages in AI models.

“Yeah, the mixture of experts approach is certainly something that all listeners should be familiar with.”

Introduction to State Space Models

36:24 to 42:00

Gain insights into state space models and their importance in AI and deep learning.

“you know, you mentioned earlier on in this episode about the Dell Enterprise Hub on Hugging Face.”

Granite 4H Models and Mamba Architecture

42:00 to 48:11

Explore the innovative design and advantages of the Granite 4H models incorporating Mamba architecture.

“That kind of accumulated into the Mamba and Mamba 2 language Bush model blocks that IBM pulled in to the Granite 4H series models.”

Challenges in Enterprise AI Projects

48:11 to 51:43

Understand common pitfalls in enterprise AI projects and how to avoid them for success.

“So changing topics here a bit from all this really cool technical stuff to kind of the real world implications of this.”

Performance Benchmarks for Dell AI PCs

51:43 to 56:00

Learn about the performance improvements and benchmarks of the latest Dell AI PCs compared to previous generations.

“And, you know, I'll defer to Tyler there to sort of talk about what are some of the best practices that, you know, customers can take advantage of with the solutions that we're providing.”

AI Performance and Device Optimization

56:00 to 57:04

Learn about the advantages of new chipset architectures for AI workloads.

“Just showing you the breadth of the AI performance and the battery and, you know, runtime and power consumption improvements for these devices relative to some of the non-AI PCs.”

Accelerating AI Transformation

57:04 to 58:26

Discover how the Dell AI factory speeds up AI model deployment timelines.

“So final kind of technical question for you.”

Book Recommendations and Follow-up

58:26 to 59:37

Get book recommendations and learn how to follow the guests for more insights.

“If you want to go do an update to your application, give your user five or 10 or five, 10 % performance or two X performance, right?”

Technical Insights on AI Models

59:37 to 1:02:01

Understand key concepts in state-space models and mixture-of-experts architectures for edge devices.

“So make sure you read it before the theater.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:Mixture of experts models and statespace models are powerful new LLM architectures, but they are cumbersome to work with, particularly if you want to use them locally or on edge devices, right? Well, not anymore. Welcome to the Super Data Science Podcast. I'm your host, Jon Krohn. today. I'm joined by Tyler Cox and Sharish Gupta, deep experts from Dell, who not only detail what mixture of experts and states-based models are, they reveal solutions that make using these state-of-the-art capabilities locally a piece of cake. Enjoy. This episode of Super Data Science is made possible by Anthropic, AWS, and Garobi.

0:39Jon Krohn:Tyler and Sharish, welcome and welcome back to the Super Data Science Podcast, respectively. It's great to have you on the show. Tyler, where are you calling in from today? Hey, John. I'm calling in from Austin, Texas. All right. And then, Sharish, I assume you're also calling in from the Austin area. Is that right? I am. Perfect. And the reason why I could assume that you're also in the Austin area is because, one, you've of course been on the show before, Sharish. In fact, episode number 921 of this podcast from September is one of the most watched videos that this podcast has ever had on YouTube, over 100 ,000 views, which is very cool to see because we kind of just started making a push into the YouTube format.

1:24Jon Krohn:And it's nice to see that that took off. Very fun, very informative episode. In today's episode, you're accompanied by a different colleague from Dell. We'll see if we can get the banter up to the same level with Tyler. Tyler, your title is Distinguished Member of Technical Staff at Dell. What does that mean? It sounds distinguished. Yeah. So I actually got a recent promotion. I'm a distinguished engineer at Dell, but I work on all things on-device AI for our client solutions group. Lots of AI PCs, all the way from our mainstream AI PCs, all the way up to our great workstation products. Very nice.

2:08Jon Krohn:Congrats on the promotion. And so we are going to be going into a really deep technical episode because you're here. I mean, Sharish obviously can get big into technical details as well. But Sharish, I've got some opening questions for you kind of around framing the particular products and services that we're going to be talking about in today's episode. And then I've got some deep technical questions for Tyler to use his new distinguished engineer title with. So, Sharish, when you were last on the show, we talked a fair bit about something called the Dell Pro AI Studio. And I understand that that was created on a whiteboard with you and Tyler.

2:56Jon Krohn:But there's also been some reframing internally around that Dell Pro AI Studio. and so it sounds like it's something that's in development hot off the press. Do you want to tell us about this? Absolutely and you know it's a nod to the whole team at Dell for something like something as complex as Dell Pro AI Studio doesn't just come into being unless there's a whole village behind it right. Special mention to a couple of folks by the name of Spencer Bull and Jacob Mink, who worked on a very key part of Delpro AI Studio before Tyler and I got to the whiteboard. So, you know, just want to call that out, those two special gentlemen out and their leadership.

3:41All right. So to get to your question, you're absolutely right, John. And if you recall, when we spoke last, Delpro AI Studio is really, you know, it's a toolkit for our customers. It's not a product unto itself, right? And which is one of the main reasons why we are rebranding it. We're in the process of rebranding it. And so don't know what it's going to be called, but just know that by, you know, in short duration from here, it will have a different name. And why that's important is in the construct of the Dell AI factory, the toolkit, formerly known as Dell Pro AI Studio, it really extends the capabilities of the Dell AI factory to our entire Edge and PC portfolio.

4:41And that was a critical part of the reason why we actually brought it into being. Nice.

4:49Jon Krohn:And so, yeah, so this thing that was formerly known as the Dell Pro AI Studio in the previous episodes that you were in, that's what we called it. It's now, it's being kind of, it's going to be part of now a bigger offering called the Dell AI Factory. And what exactly it's called, we'll probably know by the time this episode comes out. And so I will be sure to make that really obvious in the show notes on whatever platform you're watching or listening to this episode on. And so keep it out there for, you know, so that you can easily look it up. We'll have links to it and you'll be able to understand everything about it.

5:24Jon Krohn:So regardless of whatever it's called, tell us about the origin story and what problem you were trying to solve. Yep. I mean, this journey started, I'd say, arguably early last year, right? And everyone in the business was on an AI journey, right? It was the new buzz, was still developing very rapidly. So not everyone really knew what they were doing, but everyone knew that it was a very urgent journey to be on, right? And we wanted to hear from our customers and hear what problems they're trying to solve and how they intend to use AI, especially on our edge and PC devices. So a lot of research, a lot of customer conversations led to the clear realization that we needed to make it real for them.

6:26What was on device AI? That was the big question when we started out a year ago. And since then, and we've covered this briefly in one of the earlier podcast episodes, it ultimately comes to the PC fleet in the form of software that's managed on any enterprise software catalog, which customers can either buy software directly from independent software vendors like Microsoft, Adobe, CrowdStrike, McAfee, you know, next thing, SysTrack, et cetera. Or they can build AI solutions that are embedded into their workflows and solve problems that use their own data in a very, you know, private and proprietary or sovereign way, right?

7:16So the long-winded, the short form of this long-winded context is the software, the AI PC solution that is formerly known as Dell Pro AI Studio, was formed in order to help customers solve problems that was making it exceedingly difficult for them to bring their own custom AI workloads and run them on device, on their PC fleet, right? And some of the problems that we learned that we needed to solve was one, you know, This inherent complexity with bringing models and running them on PCs, diversity of silicon, diversity of runtimes and execution providers. There wasn't at that time a really good solution that would abstract away all of that engineering backend complexity for developers.

8:12And we started seeing that it was slowing down adoption, not only for customers, but also for ISVs. So we had to solve for that. The second thing we saw was there were a lot of tools that were starting to crop up that did allow you to bring AI and run models and run them on device. But there wasn't anything that truly was enterprise grade. And what I mean by that, just as an example, is it's not enough to just bring a model and run it locally on one PC. You need to have the ability to deploy and manage that through the entire lifecycle of that AI workload or app across a whole fleet of PCs. And it needs to come with the manageability and security frameworks that enterprises are familiar and expect.

9:08So that was the second problem. That just didn't exist for on-device AI. So we knew we had to solve that. And the third thing we had to make sure we did was that we had to make it as easy as deploying AI in the cloud, right? Because everyone was racing to value creation. And that means what was the easiest path of realizing AI value and outcomes. And a lot of those workloads started in the cloud and customers still start their journey there for the most part. If we had to have a chance of bringing workloads reliably on client, we had to make it as easy as deploying on the cloud. Right.

9:50Jon Krohn:So lots of our listeners will be familiar with how easy it is to call an API and use an LLM from a third-party proprietary provider like a Cohere or an OpenAI or an Anthropic. Very, very simple to write an API call and have this model kind of magically. You provide some context, you provide some instructions, and magically you have results. One step more complicated than that, typically, historically, has been being able to deploy your own model to the cloud. So regardless of kind of which cloud provider you use, there's lots of functionality there that makes it easy, relatively easy, not quite as easy as calling an API, but pretty easy to have your own model running up in the cloud in kind of like a serverless kind of feeling, or you could have it running on a server depending on your particular workload constraints.

10:48Jon Krohn:But historically, having the models run on the edge or on local devices like PCs, which are the most common kind of computer in the world and which could benefit from it being easy to run models locally on those devices, it sounds like the solution that you came up with, this AI PC solution, now allows that to happen. So it makes it as easy as deploying to cloud to deploy locally, getting around all the kinds of concerns you mentioned, like security, as well as just all of the different kinds of devices that people can be using, CPUs, GPUs, neural processing units that we've talked about in previous episodes with you, Sharish, on the show.

11:33Jon Krohn:Did I kind of summarize that back to you accurately? That was pretty good, right? And so let me tell you a little bit about the components of what used to be Dell Pro AI Studio, right? This AI PC solution. It's really three parts. One, it's a model catalog, right? And it's hosted on Hugging Face inside Dell Enterprise Hub, which is where we also have models for enterprise, just like you said, serverless, right? So the same exact place where you can get models, larger models for on-prem deployments, you can now find models for the PC that service a variety of use cases, right? Text, image, voice.

12:21the second key part which is where those esteemed gentlemen i referred to earlier come into you know the equation is every that everything that happens on device there is a ai framework that provides apis and that actually ends up you know making it super easy for existing cloud workloads to just be pointed to the PC as a local host using existing API standards. So that is the, again, goes back to my point of making it really easy to deploy workloads. And then there is also all of the model management that happens on the device. That's another model management service. That is where the secret sauce was being built already within the organization when Tyler and I kind of brought it all together.

13:20And then the third part, again, resides in the cloud, which is our partner portal, which is accessible through Intune. It's called Dell Management Portal. And this is where enterprises have that lifecycle manageability capability for all of their enterprise AI deployments that were integrated with the Dell AI PC solution, formerly known as Dell Pro AI Studio, right? So I think the other cool thing that we've done, which no one else has picked up on yet, is that we've made every component modular and behave very much like any other app that enterprises would deploy through their existing app management consoles.

14:09So both the models, the frameworks, and the installers, the core service, they're all bundled into a single installer that they can just push scripted or click to publish through Dell Management Portal, which is, I think, really what enterprises need, IT teams need.

14:32Jon Krohn:This episode of Super Data Science is brought to you by AWS Tranium 2, the latest generation AI chip from AWS. AWS Tranium 2 instances deliver 20.8 petaflops of compute, while the new Tranium 2 Ultra Servers combine 64 chips to achieve over 83 petaflops in a single node, purpose-built for today's largest AI models. These instances offer 30-40 % better price performance relative to GPU alternatives. That's why companies across the spectrum from giants like Anthropic and Databricks to cutting-edge startups like Poolside are choosing Tranium2 to power their next generation of AI workloads. Learn how AWS Tranium2 can transform your AI workloads through the links in our show notes.

15:19Jon Krohn:All right, now back to the show. fantastic so this sounds like a massive initiative and as you mentioned there are lots of people involved that have allowed this initiative to happen quickly before i get to you tyler and start asking you about you know this architecture and how it how it all works either of you i suppose just really quickly to remind our listeners why why should people be considering running workloads locally. Now, so we talked about this a lot in episode number 921 with you, Sharish, and Ish. So if people kind of want to go over all of the detail, kind of like an hour of, you know, why you should be considering deploying workloads locally, we'll do that.

16:06Jon Krohn:But just really quickly in like a minute, if one of you wouldn't mind kind of providing an overview of why we'd want to do that. Well, that's a great question, John. I think if you think about, you know, from the lens of a customer, if they're on their AI journey already, which most customers are today, unless they've explicitly made the decision not to embark on that Gen AI path, right? But those customers that are on the path firmly, what on-device AI does is it gives them optionality, right? It's become very clear to me that most workloads or most experiments are starting in the cloud. Customers are doing POCs.

16:49They don't want constraints when they're doing POCs. And when they start deploying to production, they want to get the value and outcomes and not be constrained by capabilities of models or compute, etc. So those customers that have actually gone through and deployed use cases and started scaling them out to their user base or doing multiple use cases, what they're seeing and they're signaling to us is that they're starting to realize that not every inference is equal. right and they're going to run out of uh either you know depending on their model you know their business model or how they've implemented ai they're either going to have costs that are going to skyrocket right as they scale out through their organization or if they're doing it on prem they're going to run out of of compute space right um we know that the bulk of workloads are going to be inferencing in the next few years, right?

17:58Almost, you know, the prediction is 90 % of all of AI workloads will be inferencing by the turn of this decade. So we believe, I firmly believe that the future of AI is hybrid, right? And workloads will run seamlessly across the cloud, the edge, and the PC fleet, right? So going back to my point about all inferences not being equal. That's a great point for AI leaders to start thinking about what is the right compute engine for the inference that they have at any given time. So it brings me to the concept of quality of tokens. And what that really means is the quality of the model needed to get a quality response for a particular task can be vastly different, right?

18:54And if I take one scenario of code generation or code assistance, that's a use case that is really prevalent out there. And we hear about it all the time. So take that use case. For that scenario, if you're doing autocomplete, you can probably do it on a smaller model, right? Then why burden a GB300 super chip with that? Why not do that on the PC? And then, yeah, if you're doing a full-scale code refactor, then sure, you may need a frontier model, right? So what I think will happen is there will be a lot of intelligence that goes into deciding where workloads will run. But just that, you know, if you extend that thinking alone and play it out, I see AI PCs playing a pretty important role in the future of enterprise AI strategy.

19:51Jon Krohn:Nice, okay. And then so it's kind of like concrete examples, you know, just kind of reeling off some off the top of my head. There'd be kind of situations where, you know, yeah, you gave examples there where you just don't need a gigantic model running, so why not run it locally? You could save money, you could save time on bandwidth. There also could be lots of situations where you're on a factory floor and there isn't any internet connectivity or there isn't high quality internet connectivity or where performance in real time just matters so much. You don't want any latency. And so, yeah, I don't know, just kind of reeling through some reasons why you might want to have something.

20:33Yeah, and beyond cost, right? Just like you said, if you recall, I introduced the AIPC mnemonic back in episode 877, right? Right. That's that's still one of my favorite contributions. You know, that's my mic drop moment. Just walk away, Shereesh, your contribution. But jokes aside. Right. It's still it's still a great way of recalling the benefits of on device AI inferencing. So, like you said, you know, there's plenty of use cases where latency is the most important, especially when it involves voice or video for human consumption. You know, humans can pick up, you know, those lags very quickly and it creates a poor experience.

21:14So doing some of that compute locally makes a ton of sense, right, on the PC, is you take away that 200 millisecond cloud latency. The other thing, of course, privacy of data. This is probably understated and poorly understood today. But to me, even if your data is not completely on the device, right? Like very few of us have data only on the device. A lot of our repositories are inside our firewalls, but they're in the cloud. Even in that scenario, if your LLM or your model is processing locally, your query remains completely private. And this is a big deal because if you're using a cloud model, public cloud model, your query is completely indexable by search engines.

22:09And people can know everything that was asked and everything that was put into the context, input, output, context, everything. So think of any sensitive or data that inherently needs to remain private, contractual obligations, contracts, IP, or private sensitive queries to an HR database. You know, there's just a few examples. Even though there is a, you know, a database retrieval request going through the network, the input and the output from the model stays completely private to the user on the box, right? So that's a pretty big deal in my book. And there's a new one that's not in the AI PC mnemonic, but it made this recent AWS issue made me realize that this AI on your bots is always available.

23:05You're not reliant on the wireless networks or browsers for cloud AI. Right, right. So wherever you are, whatever's going on with the network, you always have access to your AI companion.

23:15Jon Krohn:At the time of recording, we've just gone through a big AWS outage where lots of services I, for example, in New York, I couldn't use the shared bike scheme here. And I had to go around on foot everywhere for a couple of days, which was weird for me because of that AWS outage. If only they'd been running. And no, that wouldn't work in that scenario. Having local at each bike dock station, having a model running doesn't help because you need to know which bikes are available and which aren't. But yeah, OK. So we have the context now around the value of doing local compute. And both of you, Trish, Tyler, as well as lots of others at Dell, have come up with this brilliant solution, again, formerly known as Dell Pro AI Studio.

24:00Jon Krohn:And now it's this AI PC solution that will have a name for you in the show notes when this episode comes out. And so it sounds like it's a huge initiative in order to have this work across all different kinds of hardware and have the security people expect and be easy to do like people have gotten used to with cloud deployments or with just calling an API. Tyler, distinguished engineer, how is it all possible? Yeah, so we've got a great ecosystem of partners, right, from the silicon layer, model partnerships, application partnerships. And what we saw as we mapped out the journey is that we were in the right place to make a difference.

24:44Right. So we're pulling together a lot of great technologies that cover a lot of the different pieces and trying to simplify it for our end customers and developers. Right. So so a lot of this is is glue. Right. It's making sure that when you're on a Dell Pro Max system versus a Dell Pro Plus, that we're solving for the different hardware ecosystems that are inside of that and presenting that consistent, simple API across those so that you don't have to know. the different details of the quantization parameters and the tool chains that were used to prepare the models for those different platforms and the various runtime components that are in there.

25:31We've got that, right? Right now we're at eight different backend runtimes and counting. Those are not things that we built, right? Those are things that we're integrating and solving for you so your app doesn't have to carry the difference between an AMD NPU and an NVIDIA RTX Pro GPU and the different models that they need, right? So that's how, is we don't operate on our own. We're bringing together a rich and robust and innovative ecosystem and trying to simplify on top of that.

26:09Jon Krohn:Nice. Well, that sounds great. I guess it takes a whole village to raise a little software child. and so yeah there's other kinds of features that sound complicated multi-chip support model management command line tools how do those kinds of things how do those kinds of features support people like our listeners who are hands-on practitioners data scientists ai engineers software developers how do features like that allow them support them in their workflows yeah so um just take them one at a time here. So multi-chip support, we have systems that have four or five different flavors of accelerators in them, right?

26:49For an application to be able to route traffic appropriately across those, to be able to say, hey, I want this model running on the NPU for this session, but on the GPU for the next session, that's quite a bit to manage, right? On a system that's going to cover a couple percentage of your deployment base, right? But for the end user, being able to take advantage of all those different silicon types on the system they buy, that's really important to them, right? So what we've done is we've tried to make it simple for applications of all kinds to take advantage of the best local resources that are available and kind of mask some of that complexity.

27:31Now, on top of that, we've built a robust model management framework so that we're making sure that the right versions of models and all of their dependencies are getting to the right platforms. Sharish talked about our Dell management portal integration. That's making it easy for the IT admin to deploy, right? So you're loading up bundles of our solution. You're attaching them to applications. and then when you go and deploy it to 5 ,000 users, you're not having to hunt and peck and pick and choose different versions of the application to land on those different platforms. We unroll that complexity for you.

28:08We make sure that the right versions of the models and all of their dependent packages are landing on those systems and showing up underneath a consistent application. Command line tools, that's really part of our developer and management experience so that you can get granular. That's something that we're not reinventing. we're making sure that we're leveraging some of the best practices that are out there so that those of you who are familiar with local AI development, you're going to see very similar patterns in using Delper AI Studio or the formerly known solution in your development experiences.

28:46So we're trying to make it easy for you and also easy for enterprises to deploy.

28:53Jon Krohn:Excellent. All right. So in addition to all those kinds of features, let's talk about particular models that might be interesting. So this is something that I think people like to hear a lot about on the show or in general is kind of latest models, latest capabilities. And so Dell recently announced a partnership with IBM around their granite models. So granite, like the stone. And so can one of you give me the big picture around what that partnership is all about? Absolutely. We're pretty excited about that partnership with IBM, which we recently announced at IBM Tech Exchange. And what I'm most excited about is through that collaboration, we're actually bringing really best of breed models into the AI factory that then customers can use to build, deploy, and scale rapidly across not only Dell's, you know, entire hardware ecosystem, not only from the infrastructure, right, servers to edge devices to PCs.

Read the full transcript

29:58What I really like about these models, and we're talking specifically about IBM's Granite 4.0 family of models, is that they're covered under, you know, the standard Apache 2.0 license. They are a hybrid architecture model combining state space models with transformers. And that really changes the game in terms of long context performance and memory efficiency. And these are enterprise grade. And why do I say that? Because yes, one, they're open source under the Apache 2.0 standard license. They're trusted. They can be trusted. Open weights, open training data for full customizability and deployment flexibility.

30:42And to top it off, they're performant, right? They're near the top of the leaderboard of the standard Helm IFE valve for open weights. And actually, bookended by 100 billion parameter plus models on both sides, right? So these models really punch well above their weight. And in my opinion, they're ideal starting point for fine-tuning and customizability for enterprise deployments. Tyler, why don't you share some more about the models themselves?

31:16Jon Krohn:Yeah, let's talk about, so that was a really helpful overview, and it's great to hear that it's doing so well on the Helm. I'll have a link to Helm in the show notes for people who want to understand more about that. And so this is specifically at this time, this is the fourth version of this model family. So it's the Granite 4.0 model family. And I understand. Fourth major version. Fourth major version. Right. Yes. Thank you for getting, yeah. It's good to be a software people on this call to get me right on my model versionings. And so within the Granite 4 model family, there is coincidentally four different variants.

31:53Jon Krohn:And so maybe you can fill us in on the details of those four variants and why you might choose one of these models for a particular use case, Tyler. Yeah, yeah. And there's a reason I interjected there. And it was that what IBM's been doing with Granite is really exciting for enterprises. They're delivering some pretty consistent results. And I'll start with the bottom of the family and highlight that. Right. So Granite, the three series had a 3.0, a 3.1, a 3.2, a 3.3 version, all improving on the same model. The first model in the Granite 4 family extends and improves on that again. Right. So that's the same familiar pure transformers model architecture with new knowledge and improved performance.

32:45So it's a great 3 billion parameter dense transformer model. Then from there, they get even more interesting. So the Granite 4 model family has a, the other three all have a H tag. So they are Granite 4 H models. That H there stands for hybrid. And specifically, it's a hybrid state space model. So what the IBM team has done is they've blended state space models, which I think we'll talk about here in a minute. So I'm going to leave that definition hanging and transformers to get great long context scaling specifically. Right. One of the big things with attention mechanisms built into transformer models, one of the big limitations is memory scaling.

33:34There is a quadratic scaling law that means if you double the input context, you are quadrupling the memory required to run. For example, your KV cache in most scenarios, what a state space model, one of the important properties it has is you get linear scaling. So with the Granite 4H family of models, you have sub-quadratic scaling there. So the first of the Granite 4Hs is a Granite 4H micro that accompanies the Granite 4 Micro that I just talked about. So that is also a 3 billion parameter dense model. But now we have hybrid state space. Then the last two models in the family, Granite 4H Tiny and Granite 4H Small, are both hybrid state space models like the Micro, but now they are also a mixture of experts.

34:27So these are sparse models. They have a total parameter count of 7 billion for the tiny model and 32 billion parameters for the small model. They have activated parameter counts that are lower than that. So it's a 1 billion. So on tiny, you've got 1 billion active parameters. On small, you've got 9 billion active parameters. What that means, right, for those of you who a sparse or mixture of experts designation is a new concept, that means that they have a knowledge and accuracy profile closer to their total parameter count, but they have a throughput performance profile that is closer to their activated parameter count.

35:08So what they're doing is they're picking portions of the model dynamically for each request and really at a token level to answer and kind of retain only the most important pieces of the model in answering a given request. So really, really interesting blend of technologies in that Granite 4 family.

35:34Jon Krohn:Yeah, the mixture of experts approach is certainly something that all listeners should be familiar with. If you've been listening to this show for years, you've already heard about it a number of times, but if you haven't, you definitely need to know it in our space. It's important because it provides exactly that kind of power that Tyler was talking about there, where you get the power, the capabilities of very large models, or yeah, like as you said, all of the available parameters in the model that you choose, but you get the speed and cost performance of a much smaller model because you're not using all of the parameters in the model on any given call.

36:12Jon Krohn:You're using a subset like, you know, it might be common to use like an eighth of the parameters in a given mixture of experts model when you're actually running at inference time. So yeah, really cool approach. And so my understanding is that, you know, you mentioned earlier on in this episode about the Dell Enterprise Hub on Hugging Face. And so all of these granite models are available easily through that infrastructure, right? And so, you know, you can give me an answer to that while also telling me how Dell optimized these granite models to run on AI PCs and workstations. Yeah, so Dell Enterprise Hub is our partnership with Hugging Face.

36:55If you go to dell.huggingface.co, you will see our model catalog there. That's what we're talking about. What we've done in Dell Enterprise Hub is we've, number one, all of the Dell Technologies portfolio devices from AI servers through workstations through PCs, they all show up there, right? So if you say I have a Dell Pro 14 premium, what can I run? Dell Enterprise Hub is a great place to go. So inside of there, yeah, we have the Granite 4 family of models. We had day zero support from AI servers to AI PCs on there. So great place to go to get started. For optimization, so like I mentioned before, we go in and we run all the models on our devices.

37:51We figure out what the best set of inference time parameters are to fit different footprints of different devices, whether that's context length or some of the other things. And we'll bake that in, right? So we're presenting an API interface to applications. we're picking some of the best parameters for different devices so that you don't have to worry about whether you're on a 16 gigabyte AIPC or you've got access to something with an RTX Pro 6000 Blackwell and 96 gigs of VRAM. We'll make sure that you're getting best performance out of that model.

38:30Jon Krohn:Nice. That sounds like a useful feature. Now, let's dig more into what you're talking about there. We dug a little bit now into a mixture of experts. Now I want to dig into state space models and hybrid architectures, which you touched on, Tyler. So, you know, the Granite 4.0 release introduced a hybrid architecture that combines state space models with transformers. And that sounds important too, but maybe, maybe I've gone too far. Maybe we should start with digging a bit more into what state space models are first. Yeah. So state space models are a really important family of models, even pre-AI usage.

39:15So if you go back to the 1960s, these are being used in spaceflight control. They've been used in population studies and economic modeling for a long time. The core construct is that you have an input signal that you map onto a hidden or latent state space with one set of equations. And then you have a second set of equations that translate that state into an output that's observable. So there's lots of different ways that you can construct a state space model to represent different problems. um what we what uh what's what's happened over the last um 10 years or so is that state space models uh were investigated is part of the deep learning um kind of revolution uh is is how do you construct uh your your matrices for state space models to better model um different tasks without as much um kind of classical uh feature engineering approach to it right um so that that was one important thing.

40:23But then there's a couple of researchers at Carnegie Mellon and Princeton who really kind of drove this home over the last five years or so, right? So there's a great set of papers. I invite everybody listening who wants to learn more to look up the work here on Mamba and structured state space models and things like that. There's a series of papers from 2021 all the way up to 2025, still working on it, that introduced some really great optimizations inside of state-based models to make them appropriate for sequence transformation and language modeling tasks. So you get this evolution of structured state-based models, your S4 paper, and then you go into structured state-based models with selection and computation by scanning your S6, which turns into, hey, that's a lot of S's, sounds like a snake.

41:27Now we've got Mamba. And the Mamba architecture really optimizes states-based models for the kind of compute profile that's needed to be relevant for for the language model tasks that we're applying here. And so some great optimization work that happened there, some great mathematical insights into the matrix properties of those. And I won't be able to do those full justice here, but really, really interesting work. That kind of accumulated into the Mamba and Mamba 2 language Bush model blocks that IBM pulled in to the Granite 4H series models. Right. So composition here, you've got a nine to one ratio of Mamba layers to attention layers in the Granite 4H family.

42:29So quite a bit of state space model in that hybrid. Those are, like I said earlier, those are linear context scaling. They are the Granite 4H models are no positional embedding. So they have in the data set out to 512K context represented in the samples. They're validated out to 128K. The IBM team says, theoretically, you should be able to push it past that, right? So some really great long context performance. I think one of the key measurement points in the release notes are if you take eight sessions at 128K context on a micro, so a 3 billion parameter model, you get about 15 gigabyte of memory usage versus about 80 on a pure transformer architecture, right?

43:27So some really great context reduction, which means that on more constrained devices, on the edge, you can make use of more useful context in RAG workflows, in multi-term workflows, and things of that nature. Just to close here, Shereesh had mentioned IF eval. That's a really great benchmark for instruction following. right, structured output tasks. The other thing that Granite scores really well on is the Berkeley for function calling leaderboard, specifically BFCLV3. It shows up in the top five as we sit here recording, among a bunch of other frontier models and hundreds of billion parameter models even at, for the small, 32 billion parameter footprints.

44:26So really, really punching above weight class there.

44:29Jon Krohn:Really cool. Thanks for all those stats in there. You really packed them in. And so it's kind of like a key takeaway, basically by having these post-transformer architectures as part of these granite models. And so it sounds like specifically they've gone with Mamba as their particular state space model. It allows us to have better performance over long context windows, like you're talking about half a million input tokens, being able to handle that. And when we handle that kind of large amount of context, that large number of input tokens, we start to become worried about memory and about compute performance.

45:09Jon Krohn:And so over these very long context windows with these state-based architectures like Mamba, we're able to have scaling that isn't the quadratic scaling that we were used to with transformers. And what I mean by quadratic really quickly for listeners is with the transformer architecture, which is still the predominant model architecture within LLMs today, as you increase the number of input tokens, the compute and the memory increases quadratically. So by a squared factor, basically. And that starts to add up really quickly. And so, yeah, so we're getting sub-quadratic memory and compute performance, which matters a lot in these very large input token situations.

46:06Jon Krohn:So really cool. So you get a lot of power, a lot of flexibility with these Granite 4.0 models because they incorporate the state-based models like Mamba into it. And so I guess my final question for you on this, and it was actually kind of where I opened around this, was I was talking about how the Granite 4.0 models are actually a hybrid that include, I think you mentioned there, a nine-to-one ratio of Mamba versus traditional transformer attention heads. Why, what's the advantage of having both? Yeah, it's a great question. So one of the things we do see with Mamba is it has great global context performance, but it loses a little bit of the sensitivity of local context that attention players are really known for.

46:52Right. Now, one of the other interesting things that happened right after IBM launched the Granite 4 models is there was a great meta fair research paper that compared some of the performance of a pure transformers architecture, a pure Mamba architecture, an intralayer hybrid and interlayer hybrid. Right. And so the intralayer versus interlayer, that's looking at running through attention layers and Mamba layers in parallel versus interlayer, which is what the Granite 4H models are, which is running through them sequentially. Right. So I think one of the recommendations from that team was build more more of these hybrid models that the the training efficiency, which we didn't talk about, and the inference scaling efficiency is really great without a lot of tradeoffs on accuracy, given some of the the Mamba improvements.

47:54So I would expect to see more of these type of hybrid models in the future.

48:01Jon Krohn:No doubt. No doubt. It is the future indeed. Some, yeah, a post-transformer architecture will be the dominant architecture in the future. All right. So changing topics here a bit from all this really cool technical stuff to kind of the real world implications of this. If you are an enterprise and you're trying to take advantage of AI, there was a recent study, I talked about this on an episode in the past, probably several times, but I did a whole episode, episode 924, on this MIT researcher claim that 95 % of enterprise AI projects fail, where fail means that it does not deliver a return on investment in production.

48:44Jon Krohn:So either it never makes it to production or it's never profitable in production. It's never successful. It never meets your success criteria in production. And so can you speak to or provide examples of what typically goes wrong in enterprise AI projects and how our listeners perhaps leveraging things that we've already talked about in this episode can prevent those kinds of failures? That's a great question, John. And honestly, you have done a stellar job covering it in that podcast episode that you referred to. But, you know, my take on it, I think some of the most important things that everyone that's on their AI journey will do well to sort of ground themselves to is it starts in a very cliched manner.

49:35It does start with setting clear goals for what outcomes are desired, right? And that includes very specifically what is the desired type of response and accuracy of said responses. And I think we're still very reliant on humans for ensuring that outcome. In other words, you just can't proceed unless you have humans in the loop and expect to have, at least during the training phase, and expect to have good results. In fact, I think having humans in the loop even during production as feedback baked in to the solution itself is vitally important because you want to account for drift. Right. But another cliche I will share is don't apply Gen AI for the sake of it.

50:35You know, use the right tool for the job. I think I've seen plenty of misuse where everyone just wants to apply Gen AI to a problem that is very well solved with traditional. I call it traditional. It's funny, but machine learning. Right. Solid, solid machine learning deterministic outcomes. So those are the basics, right? I think beyond that, you just want to really be very deliberate about the problem you're trying to solve, right? So again, problem, goal, tool for the job. Sounds pretty basic. You could actually take that response and apply it to literally any, you know, even a classroom discussion on how to do a project or how to manage projects.

51:22So I think those are the basics that you need to ensure. And then as you go from there, there are obviously going to be very use case specific, task specific best practices, which need to be uncovered and adhered to as you go along the journey.

51:43You know, one other thing I will mention here is that it goes back to some of the pain points that we anchored on earlier on in the conversation about the problems that we were solving for with the Dell AI factory and the entity formerly known as Dell Pro Ahas Studio. And, you know, I'll defer to Tyler there to sort of talk about what are some of the best practices that, you know, customers can take advantage of with the solutions that we're providing. Yeah, so I think taking advantage of the device through our solutions means that you've always got a enterprise-ready model at your disposal, right?

52:31There's 2 million models on Hugging Face. They all do different things really well. We're trying to bring the ones that we think are really great for a lot of key enterprise use cases ready to go. We also make sure that they're performant on our devices of all types. So when we pick models, we try to make sure that they've got great coverage across different silicon types. So they're really good building blocks. A lot of times you'll start down a pathway with one model and you might hit a dead end when you try to scale up or scale across different device categories. It can be a little bit challenging to figure out what the right pathway is to getting everything hooked up the right way.

53:25Jon Krohn:All right, that all sounds great, Tyler. Sharish, you mentioned to me before we started recording, and you mentioned actually a little bit in today's episode as well, how Dell PCs perform on benchmarks, how Dell AI PCs perform compared to previous generations, particularly around compute, battery life, efficiency. Do you want to fill us in a bit on those key benchmarks? Yeah, absolutely. I think it's important for customers to understand what they're actually getting when they invest in the newer hardware, right? So, for example, the Dell Pro Plus with Intel Core Ultra 200 V-series chips, there's just tremendous benefits, advantages over, say, an N-2 non-AIPC chipset like the Core Ultra 14th gen, right?

54:27So I'll just give you some examples. 88 % more battery runtime running Microsoft Teams meetings compared to the non-AIPC on the, you know, the Dell Pro Plus with Intel Lunar Lake. 4.8X higher graphics performance. So very important to think about the GPU as well, the iGPU, which is tremendously, you know, is shown tremendous gains gen over gen. And even within the gen of Intel's latest chipsets, Lunar Lake, with its system-on-chip architecture, is just tremendous in terms of its gains for the iGPU, not to be ignored. We've talked so much about NPUs, but I think if you look at the overall CPU, GPU, NPU as a whole, it's just a tremendous value proposition.

55:24Almost 10x on-device AI CPU performance and 4.2x AI GPU performance on that same device relative to the non-AI PC. And then if you really want to talk about specific models, because that's what your users are going to be most interested in, the same device with Lunar Lake provides greater than six times higher performance on Mistral 7B and Llama 3.1 for TextGen and up to 5.6 times higher performance for ImageGen with Stable Diffusion 1.5. Just showing you the breadth of the AI performance and the battery and, you know, runtime and power consumption improvements for these devices relative to some of the non-AI PCs.

56:17So I really encourage people to look at these new chipset architectures and the new devices from a holistic lens and not just focus on the NPU itself.

56:30Jon Krohn:Yeah, so it sounds like we've gotten to a point where companies like Intel, with the way that they make their CPUs, as well as obviously GPU providers, neural processing unit providers, and then even companies like Dell, the way that they package all of that up with something like an AI PC, you are optimizing everything for AI workloads, for training AI models, for real-time inference on the edge, on AI PCs. And so that's how you're getting the big multiples that you just went through in dozens of examples of, Sharish. Yep. Nice. All right. So final kind of technical question for you. Historically, AI transformation was a multi-year, huge undertaking for enterprises.

57:16Jon Krohn:But that kind of timeline is irrelevant these days because, you know, if you think about a multi-year timeline, we have no idea what the capabilities are going to be like a few years from now. We can't be thinking about a solution today that's going to be ready in multiple years. How does the Dell AI factory accelerate the timeline so that we're talking about weeks or months instead of years getting to deployed AI models? Yeah, that's a great question. I think part of it is you've got to simplify how you think about it, right? If you can't get, if you're down in the details on everything, then you're slowing down.

57:55So picking the right technology partners matters. It will help you go faster. The second thing is you need to build on technologies that will scale and help you, not just at deployment time, but six months, a year, two years down the road. We all know AI is not sitting still. The model you deploy today is not going to be the same one that you deploy in six months or a year. Having manageability built in means out and control that, right? If you want to go do an update to your application, give your user five or 10 or five, 10 % performance or two X performance, right? We don't know what tomorrow brings.

58:40Then you need to build on top of the right tool sets.

58:44Jon Krohn:Awesome. Really enjoyed this interview with both of you today, Sharish and Tyler. Sharish, we've had enough book recommendations from you already. Tyler, what have you got for us? Yeah, so I read a bunch of different things. I use it as kind of escape, right? Whether that's fantasy, science fiction. One of the things I think is something that, a story that a lot of people will love is one that is being adapted, right? So Project Hail Mary from Andy Weir is my recommendation. It's a great, great fun story. Touches on a lot of interesting environments and excited to see the adaptation. I think next year is when that will come out to film as well.

59:37So make sure you read it before the theater.

59:40Jon Krohn:Yeah, yeah, exactly. All right, thanks for that recommendation, Tyler. And we might as well just stick with you for a second. How should people follow you after this episode So to get more insights on, I mean, yeah, you went into huge technical detail on models, on capabilities. Where can they get more insights from you after the show? Yeah, I'm on LinkedIn. We'll put the profile in there. I will say I'm not a heavy poster, so don't expect to see a ton from me. But you can be sure that you'll get the latest updates on the Dell AI Factory from my feed. Excellent. And Sharish? Same. I'm also on LinkedIn.

1:00:19My link will be in the show notes. And that's, you know, I'm maybe a little more active compared to Tyler, but still not one of your prolific, you know, everyday prolific posters. So, but yeah, LinkedIn is a great place to follow and I appreciate the connect.

1:00:39Jon Krohn:Fantastic. All right. Well, thanks for that. And thanks for this whole episode again, Sharish and Tyler. and yeah, I feel like it's probably not going to be long. We don't have it planned, but I wouldn't be surprised if we're welcoming one or both of you on the show again in the future for more. These are the only kinds of episodes where we dig in detail on how people can be getting inference or model training on the edge. And so I always learn a lot in them. I really particularly enjoyed all the conversation today around state-based models and mixture of experts and the granite release, all really cool stuff.

1:01:19Jon Krohn:Thanks, guys. Thank you, John. Thanks for having us again, John. Yeah, it's always a pleasure to be here. And I was delighted to have Tyler join me for this episode because his technical depth is just awesome. It's awe-inspiring. I feel humbled every day when I enter the innovation lab where he sits and just can't get rid of my imposter syndrome. Like, you know, I don't deserve to be here. This is the place where genius, you know, way beyond my years thrives. So I just enjoy working with him and his team every day. interesting episode for sure in it Tyler Cox and Sharish Gupta covered state space models like Mamba and how they use selective mechanisms to process information more efficiently than transformers making them ideal for edge devices where memory bandwidth is the bottleneck they also talked about mixture of experts architectures that activate only specific subsets of parameters for each task allowing large models to run on resource constrained devices by keeping most parameters dormant.

1:02:25Jon Krohn:And they talked about how the Dell AI factory simplifies deploying and managing custom AI workloads across enterprise PC fleets by abstracting away silicon diversity, simplifying deployment, and providing enterprise-grade manageability, accelerating AI transformation from multi-year timelines down to weeks or months. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Tyler and Sharish's social media profiles, as well as my own at superdatascience.com slash 939. All right, that's it. Thanks to everyone on the Super Data Science podcast team, our podcast manager, Sonja Brejevic, media editor, Mario Pombo, partnerships manager, Natalie Zajski, researcher, Serge Massis, writer, Dr.

1:03:09Jon Krohn:Zara Karche, and our founder, Kirill Aramanko. Thanks to all of them for producing another excellent episode for us today for enabling that super team to create this free podcast for you. We're deeply grateful to our sponsors. If you're ever interested in sponsoring the show yourself, you can find out how to do that at johnkrone.com slash podcast. Otherwise, you can support us by sharing the show with people who would like to listen to it or watch it, review the show on your favorite podcasting app or on YouTube. Subscribe if you're not already a subscriber, but most importantly, just keep on tuning in.

1:03:41Jon Krohn:I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

From the publisher

State space models (SSMs), granite models, and Mamba: Dell’s Tyler Cox and Shirish Gupta discuss with Jon Krohn why state space models can process information so efficiently, and how Dell’s AI factory helps enterprises manage custom AI workloads. Hear the latest on the Dell Pro AI Studio and Dell’s partnerships with IBM and Hugging Face in this episode. 

This episode is brought to you by the Trainium2, the latest AI chip from AWS and by Gurobi.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/939⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(02:58) Dell Pro AI Studio news

(23:17) How Dell manages interoperability

(28:08) About the Dell/IBM granite models

(47:38) How to troubleshoot AI tools

(52:36) How Dell performs against benchmarks

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
939: Mixture-of-Experts and State-Space Models on Edge Devices, with Tyler Cox and Shirish GuptaSuper Data Science: ML & AI Podcast with Jon Krohn · 1 h 6 min
Listen in VO