921: NPUs vs GPUs vs CPUs for Local AI Workloads, with Dell’s Ish Shah and Shirish Gupta

9 Sep 2025 · 1 h 12 min · 25 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How to choose between CPUs, GPUs, and NPUs (neural processing units) for local AI on client devices, plus when to run workloads locally vs in the cloud; includes Windows vs Linux setup and Dell’s approach to simplifying on-device deployment.

Guests

Ish Shah, Dell Head of CTO Pursuits (client devices); MBA from MIT; previously a consultant at BCG. Shirish Gupta, Dell Director of Product Management; calls from Dell HQ in Round Rock, Texas; has been at Dell ~22 years; previously appeared on episode 877.

Key claims

Windows remains dominant for software dev (~64%); Linux best matches production for large deployments (~96% on Linux servers), with WSL2 bridging the gap. NPUs are purpose-built for AI matrix/vector math and deliver better performance per watt (important for battery). GPUs scale better and support much larger models; NPUs are evolving. For future-proofing, Dell frames devices into AI PC tiers: entry NPUs (10–15 TOPS; ~1–3B params), higher NPU PCs (~40–50 TOPS; up to ~9–10B), and high-performance PCs with discrete GPUs/NPUs.

Notable examples

Dell Pro Max Plus workstation with discrete NPU for offline inference (demo: live colonoscopy video inference for Northwestern Medicine’s ARIES model; no cloud uplink). Claim of running a 109B Llama Scout speculative decoding model at FP16 locally. Dell Pro AI Studio (built on OpenVINO) reduces model/toolchain conversion complexity; POC with Deloitte: ~3 months to deploy with OpenVINO vs ~4 days with Dell Pro AI Studio. Cloud vs local reasons: privacy/IP, connectivity, speed, and token/cost control; future is hybrid/optional.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Meeting the Guests

0:42 to 1:08

Hosts introduce Ish and Shirish, discussing their backgrounds.

“This episode of Super Data Science is made possible by AWS and the Open Data Science Conference.”

Guest Backgrounds and Roles

1:09 to 3:52

Ish and Shirish share their experiences and roles at Dell.

“Shurish described Ish as his more flamboyant sounding version.”

Choosing Between Operating Systems

3:53 to 8:16

Discussion on why to consider Windows over Unix-based systems for data science.

“now that our listeners are hopefully familiar with the voices of Ish and Sharish, we can get into the content of today's episode.”

Advantages of WSL and Flexibility

8:17 to 9:38

Exploration of Windows Subsystem for Linux and the benefits of using both OS.

“It seems kind of like there was, I went to Oxford University for my PhD and they have something called the Oxford Union there, which is a debate club.”

Understanding NPUs and AI Workloads

9:39 to 14:02

In-depth overview of neural processing units and their role in AI.

“I'm going to butcher the Spanish that you so well said there, but porqué...”

Understanding AI Workloads and Processor Types

14:02 to 14:59

Explore the differences in running AI workloads across CPUs, GPUs, and NPUs.

“And the way that that runs through those transistors is like different than the rest of the stuff.”

The Introduction of Discrete NPUs in Laptops

14:59 to 16:27

Learn about Dell's new laptop featuring a discrete NPU and its implications.

“And it is a big honking thing that we put into a laptop chassis just because we could.”

Use Cases for Discrete NPUs

16:31 to 17:24

Discuss who benefits from discrete NPUs and their real-world applications.

“Okay, so this sounds pretty exciting, what you're talking about there, with having a discrete NPU.”

Use Cases for Discrete NPUs

17:29 to 19:13

Discuss who benefits from discrete NPUs and their real-world applications.

“Like everyone listening to this podcast, well, actually, maybe this is a pretty biased sample.”

Comparing NPUs and GPUs

19:13 to 24:10

Understand the performance differences and applications of NPUs and GPUs.

“And that's, and the, sorry, just really quickly, Sharish, the name was Dell Pro Max something workstation.”
Show all 25 chapters

Future-Proofing Your Device Purchase

24:10 to 28:00

Identify key parameters to consider for future-proofing your tech investments.

“And the differentiator for NPUs is performance per watt.”

Understanding AI Workloads and Device Categorization

28:00 to 32:20

Explore the classification of AI devices and their suitability for different workloads.

“The data scientist persona is something that like an IT decision maker is constantly thinking about.”

Commercial Considerations in AI Hardware

32:20 to 34:50

Discuss the commercial aspects of selecting AI hardware and its implications.

“So that's the three-pronged categorization today for AI-BCGs.”

Utilizing NPUs, GPUs, and CPUs for AI Development

34:50 to 42:01

Learn how to leverage different processors for AI workloads and the associated complexities.

“And so I think we're used to, probably most listeners are aware that CPUs are doing the most kind of general work, running your operating system.”

Introduction to Dell Pro AI Studio Features

42:01 to 43:34

Learn about the key features and functionalities of Dell Pro AI Studio.

“So, you know, and that's not like a linear, you know, simplification of time.”

User Stories for Local AI Workloads

43:35 to 46:36

Explore how users can benefit from running AI workloads locally.

“One thing that I think might be helpful to me and to our listeners is to understand, I realize that, as you mentioned, there are lots of kinds of personas out there.”

Understanding CPU Architecture: Lunar Lake

46:37 to 47:46

Gain insights into the significance of Intel's Lunar Lake CPU architecture.

“Am I forcing this down someone's throat because I want to?”

The Importance of Local Processing in AI

47:47 to 49:54

Discuss the role of local processing in the context of generative AI.

“So it was quite a different architecture in which you had memory on chip.”

Choosing Between Local and Cloud AI Workloads

49:55 to 55:16

Understand the factors influencing the decision to use local versus cloud AI workloads.

“So what are the kinds of things that gives us the capability to be taking these open source models.”

Future of AI Workloads: Hybrid Models

55:17 to 56:01

Explore the future potential of hybrid AI workload models and their implications.

“The future is hybrid because it is not practical or feasible to move all workloads to the PC.”

The Future of Distributed Computing

56:01 to 57:54

Explore the emerging trends in distributed computing and local AI experimentation.

“Whether it's the PC fleet, a GPU cluster on a data center or wherever.”

Upcoming Windows Refresh Insights

57:55 to 58:31

Discussion on the significant upcoming Windows refresh and its implications for users.

“there's a uh there's a big windows refresh coming up uh in a couple of months in october well, I guess at the time of this episode release in a month.”

Dell Pro Max: New Hardware Capabilities

58:32 to 1:02:52

Details on Dell's new line of PCs equipped with advanced GPUs for AI workloads.

“So let me go first and then I shall, you know, we'll get your take.”

The Importance of OS Refresh

1:02:53 to 1:04:49

Understanding the importance of timely OS updates and hardware upgrades.

“So that's, uh, I think the only thing I can add to Sharish is like, Hey, like October is here.”

Book Recommendations for Data Enthusiasts

1:04:50 to 1:07:26

Recommendations for books that provide insights into AI, data science, and business strategy.

“Like, for me to start now, mid-career, trying to learn this stuff, why even bother, right?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:Welcome to another episode of the Super Data Science Podcast. I'm your host, Jon Krohn. Today, I've got not one, but two guests for you named Ish and Shirish. I am not making that up. Both Ish and Shirish are senior leaders at Dell. And not only are they extremely knowledgeable, this episode is packed with fascinating, actionable facts about AI hardware considerations, including cloud versus local, and what workloads are best suited to CPUs versus GPUs versus NPUs, neural processing units. In addition to all that knowledge that they provide, they are also both very entertaining and play off each other in funny ways.

0:36Jon Krohn:This episode is critical listening for anyone working on training or deploying AI. Enjoy. This episode of Super Data Science is made possible by AWS and the Open Data Science Conference. Welcome to the Super Data Science Podcast. I'm joined by two people today. So I've got Ish and Shurish. So a lot of Ishiness happening today. I'm going to take that, John. That's mine now. You can't have that. Yeah, and that's Ish speaking right there. Shurish described Ish as his more flamboyant sounding version. So if that helps you distinguish their voices through this episode, we can try that. I'm also going to have them introduce themselves briefly to give you a bit of a sense of who they are.

1:26Jon Krohn:So some listeners may actually already be familiar with Sharish. Sharish was in episode 877 of this podcast. Sharish, welcome back to the show. Where are you calling in from today? Great to be back, John. I am calling from our headquarters in Round Rock, Texas. Yeah, Dell headquarters in Round Rock, Texas. And so you're a director of product management at Dell. You've been there for like 22 years or something like that? Wow. And basically in Round Rock that whole time? Yes, I've been in the Austin area for the entire duration. Wow, wow, wow, wow. Nice stability there. Ish, you have not quite been there as long.

2:07Jon Krohn:You've been at Dell for four years. You have a title that honestly, I don't understand what it means. You know, most days I don't either. It's okay. Head of CTO Pursuits. Yes, sir. Our client devices. You know, it's a really fun way of saying I get to work with people who, very fortunately for me, are way smarter than I am. This is a really good example. I like to joke that Sherish is the evolved Pokemon version of Ish. And I think that that's true for pretty much everyone that I work with out of the CTO Pursuits team. Because it's about getting in the field, forward deployed engineering style, roll up your sleeves, find a problem, help someone solve it.

2:47kind of going beyond the whole, you know, we're going to sell you some laptops vibes that maybe some people know Dell for. We've got this whole new function that we're stood up and there's some really, really talented people helping solve some really, really complicated problems. So it's been a lot of fun for me.

3:02Jon Krohn:Nice. Well, I think you're downplaying it a little bit. You have achieved some things in the past that make me think that there's some intellect happening around on your side there ish. So you have an MBA from MIT and you were a consultant at BCG before being at Dell. You know, those are, these are all my dark, dirty laundry things. I don't bring up here. Here's the consultant for you all. I just saw the viewership number crashed as here. It's okay. I love saying like, you know, I, I adored my time at Sloan, but very much, you know, kind of our fake MIT. Sometimes I, I was schooled frequently by some of our visitors from the other courses as majors are called at MIT.

3:44So yeah, it was great. I hope my mom listens to this episode. I think I'm going to get some kudos I very surely deserve.

3:51Jon Krohn:Nice. All right. Well, so let's get into now that our listeners are hopefully familiar with the voices of Ish and Sharish, we can get into the content of today's episode. So Dell famously makes PCs. It's one of the things, if not the thing that Dell is probably best known for. And so the first question that I would like to kick things off with is for our listeners who are interested in data science, AI applications, a lot of open source, the first thing that might come to a lot of listeners' mind is Unix-based systems. And so why should somebody be considering a PC instead, in particularly maybe even a Windows machine?

4:35Yeah, well, let me go first and Ish will have his take on it as well. I would say that you have to start with what are your objectives, right? If you look at the data science and, well, the software development community at large, let's just look at the facts. Windows is still the most popular OS for software devs, right? It's about 64 % the last time I checked of software devs use Windows for development, which is still the leading number out of all of the OS platforms. It is, yes, the standard in the enterprise and for consumers on PCs. So it's familiar, it's friendly, user-friendly, great for beginners in the data science field.

5:24It's compatible with popular apps, right? Who doesn't want to use productivity apps? You know, there's other data manipulation and visualization apps that run better on Windows. And then at the end of the day, as I said, it's a standard in the enterprise. So if I'm thinking from an IT management perspective, you get the enterprise security integrations are much more easy, right, with Windows. But that's not it, right? Because let's not kid ourselves. I think if you look at large ML and data science deployments, I think 96 % of them are still running on Linux-based or Unix-based servers. So if you are doing large deployments, Linux is probably still the platform that you want to be developing in.

6:14Because best practices, you want to be developing on the platform that is being used for production environments. But I would say at the same time, there's an alternative. There's Windows subsystem for Linux. And WSL2 specifically, I think, starts closing the gap in all of those scenarios. Right. Because you can run a Linux kernel directly on Windows. So if you have Ubuntu running natively, guess what? You now have all of your Linux command line tools that you can run right there on Windows. And so you have best of both worlds. You have all your productivity and other ease of use benefits on Windows.

6:55And you don't lose all of the benefits you get working with Linux for data science solutions. So I think that is something that's worth considering. I think WSL's use is still pretty small out there today, but that's an alternative for someone who's on Linux, used to Linux and wanting to bridge the gap. Right. Then last but not least, this is something I touched on briefly in our last episode, John. If you're building for PCs, if you're building apps for PCs, software developers, it goes back to the same best practice. you want to be building in the environment that you're going to have your apps in production.

7:37You want to be building in Windows, right? Let's look at the number. 16 million PCs with Linux, 1.6 billion PCs with Windows. That's less than 3 % Linux, right? And you get all your IDEs, VS Code, Visual Studio, PyCharm, IntelliJ IDEA, et cetera. They're all run natively on Windows. So, you know, it sounds like I'm making a case for Windows here. I'm not a Windows, I mean, not a Microsoft employee, but I have to say at the end of the day, it comes down to what you want to do, right? There's a place for Linux for data scientists and there's a place for Windows.

8:16Jon Krohn:Yeah, that was really well argued, Sharish. Thank you. It seems kind of like there was, I went to Oxford University for my PhD and they have something called the Oxford Union there, which is a debate club. And I went and watched some of their debates. and that sounded like kind of opening remarks. You have all the stats. Very compelling argument. Ish, what do you have to say? Are you on the other side of the argument? I think mine is probably a lot less articulate because it's basically porqué no los dos. It's why pick one thing when you don't have to pick one thing, right? With WSL, it's like, what do you need?

8:52Just kick into it. There is no dual boot. It's just sort of in the same space. And it's a very fluid transition back and forth. You can bop around in WSL and Windows as you need. And I'm using very technical terms here. But I think the biggest thing is you can hit a nail with a wrench, but why would you if you have a hammer, right? You've got the right tool for the right job at the right time. And I think that's going to be a theme in this conversation, John. Dell's whole point of value in this whole brave new world we live in. is choice. It's like you get to decide what you want to do. You want to throw Linux on your device?

9:31Do it. Want to do Windows on your device? Do it. You want both? Do that too. Whatever works for you.

9:37Jon Krohn:Awesome. Speaking of that idea of being able to have your cake and eat it too, I'm going to butcher the Spanish that you so well said there, but porqué... Porqué no los dos. Why not both? Porqué no los dos, yes. And so on the theme of that, I'd like to talk about how your machines, how Dell machines are so supportive of all different kinds of hardware. So not just CPUs, not just GPUs, but NPUs as well, neural processing units, and allowing all three of these to work together. So, Sharish, in your previous episode on my show, in episode 877, you began with a really brilliant explanation of what NPUs, neural processing units are, since that's something that maybe listeners, if they haven't listened to that episode, they should probably get at least a short intro to that.

10:33Jon Krohn:I'd love to hear a bit about MPUs, and then maybe we can also talk about CPUs, GPUs, and the different kinds of situations where you would use those. We'll get to that next. Let's just start with what MPUs are. Yeah. So I do encourage everyone to look at or listen to episode 877, where I do cover that in detail, as John said. But for those who are coming in to this cold, an NPU stands for a neural processing unit. And it's integrated into... Actually, I'll step back. There's two ways you can have access to an NPU. You can have it either integrated into the chipset, or it can be a discrete card that you either put into a slot in the PC or built into a mobile device, but it's still a discrete card.

11:21So, you know, you have both flavors of NPUs. At the end of the day, what sets them apart from the architectures that we've had for, you know, maybe more than a decade at least, is that they are purpose-built to handle AI and ML workloads, right? So that vector math, that is such a foundational aspect of of neural networks, the NPO architectures are almost hardcoded to handle those workloads. So what you get as a result of that is a much more efficient, you know, processing engine. You don't, especially important for a mobile device where battery life is pretty important, right? The last thing you want is these new workloads that are running on the PC and they tank your battery life.

12:18If your battery only lasts half of what it normally does, that's a no-go, no bueno, right? So NPUs fulfill that. Its performance per watt is what you get from the NPU for AIML processing.

12:33Jon Krohn:Yeah, so it's basically, it's optimized for the kinds of matrix multiplication operations, neural network architectures that are so common in everything, transformer architectures, And so therefore, all of the large language models, most of the AI capabilities that we have today can run more efficiently. And especially, as you said there, in a more conservational power way, a more efficient way on mobile devices. But even on a laptop, even on a desktop, a server, I'm sure that those improved efficiencies make everything run along more smoothly. I think if you go back to the CompSci 101, think about the logic gates, think about the levels of abstraction, and think about how a computer does what it does.

13:29How does a computer compute? we've been having these debates around x86 and ARM and even RISC, right? The whole RISC-V project. Like the way that these things flow through the transistors on the device, how are you organizing those transistors on the device and on the chip? And that's really, this is the natural evolution of, hey, I've got a CPU. It's really good at this. I've got a GPU. it's really good at this. Now there's this new thing and this new thing is an AI workload. And the way that that runs through those transistors is like different than the rest of the stuff. Right. And I know there's going to be people in the comments being like, oh, it's just a current through a thing.

14:11Like, yes, we know. But you understand the point, which is that there is a correct tool for a particular job at a particular time. That's not to say that GPUs and CPUs can't run AI workloads. Of course they do. And they do it tremendously well. And there are entire organizations and companies and startups dedicated to how can I get an AI workload inference really, really well on the CPU? How can I get it to run more power efficiently on the GPU? And then there are people who will skip that process entirely and say, well, if I've got a purpose-built chip. And the reason Sharish paused when he was describing what an NPU was is because, John, the last time he talked to you, he couldn't tell you something that is now true, which is that Dell has announced a device that has a discrete NPU in it.

14:58So it's not just on the SOC. And it is a big honking thing that we put into a laptop chassis just because we could. But I think that right now, what used to be a very easy matrix and a very easy matrix for data scientists to understand, IT decision makers to understand, IT buyers to understand. Two by two, I've got this many companies making this many things and I can pick from this matrix. That matrix has gone from two by two to like eight by eight. And now there's this like, who needs what tool for what purpose? And it's a really complicated question. Whereas two years ago, it was like, well, what's the newest thing on the menu?

15:39I'll take that. It's not so simple anymore.

15:42Jon Krohn:Curious about Tranium 2, the latest AI chip purpose-built by AWS for large-scale training and inference? Each Tranium 2 instance packs a punch with 20.8 petaflops of compute power, but here's where things get really exciting. The new Tranium 2 Ultra Servers combine 64 chips to deliver a massive 83 petaflops in a single node. These Tranium 2 instances deliver 30 to 40 % better price performance relative to GPU alternatives. Major players in AI like Anthropic and Databricks, along with innovative startups like Poolside, have teamed up with AWS to power their next-gen AI projects on Tranium 2. Want to see what Tranium 2 can do for your AI workloads?

16:25Jon Krohn:Check out the links in the show notes. All right, now back to the show. Right. Okay, so this sounds pretty exciting, what you're talking about there, with having a discrete NPU. Correct me if I completely butcher the way I'm explaining this, but so a discrete NPU on a laptop. And so if somebody wants to get that, how do they find that? Like, how do they, how do they search for that? I can have it in the show notes. This is something that like Dell sells, they could just buy. And why specifically, you know, you talk about this eight by eight matrix. Why should a listener be interested in getting a laptop with a standalone with a discrete NPU.

17:02Yeah, I'll tell you that the most requested device out of our CTO team right now for like their next device refresh within Dell, everyone is on our boss about, I want one of those. And he's like, yeah, we really got to think about how many of these we can get for you guys. But it's the Dell Pro Max mobile workstation and it has a discrete NPU device coming to it. And the CPU is also powered by an Intel chipset. So it's a very interesting machine. And it's not for everyone. Like everyone listening to this podcast, well, actually, maybe this is a pretty biased sample. Maybe everybody listening to this podcast is going to want one and will find a use for it.

17:43But think of these specialty devices, right? Think about the GB10 and GB300 coming out from NVIDIA not too long in the future. When we launched this Dell Pro Max, it's just another beast. And it's like the right person will want that versus something else. A very cool demo that we ran on this device at Del Tech World not too long ago. We took a model that Northwestern Medicine had fine-tuned and trained. It's their ARIES model, which does imaging, CT scans, x-rays, all done in-house, by the way. Dr. Mozzie and his team over there, shout out to them because they are the true mad scientists. And, you know, it was a demo that we worked on with them.

18:23And we ran inference on a colonoscopy video, which we blurred out, but a colonoscopy video that inferenced live, like in front of you on that device, no internet uplink needed, no queuing, no waiting at a server somewhere to be processed and to be shipped back. So the question of like, who would want one of these things? Hey, if you've got this sort of specialized use case, or say you live in a part of the world where there are like restrictions on what you can and can't send to the cloud and in which circumstances you can, this kind of device is going to change the game because it has that level of performance and speed that you just, six months ago, 12 months ago, the last time Sharish was on this podcast, you didn't have that option, right?

19:06So it's coming out soon, not out yet, but yes, if you look for the Dell Pro Max with discrete NPU, you will see the press about it.

19:13Jon Krohn:And that's, and the, sorry, just really quickly, Sharish, the name was Dell Pro Max something workstation. What was it? So it'll be on the Dell Pro Max Plus devices. It will be, as I said, it will be available in the second half of, well, I should say Q4 now at this point. I mean, I guess both are true. More to come. Okay, okay, that's cool. All right, so I might not exactly be able to include a link in the show notes. I'm not sure. I'll do my best, but regardless, people listening in the future will be able to either find it by looking it up or they will be able to find it soon. Cool. And sorry, Sharish, I interrupted you, please.

19:52I was going to say that, you know, and Aish can correct me here, but what's incredible about it is that we could also run 109 billion parameter Lama Scout speculative decoding at FP16 on this card, which is insane. So think of that running locally and the things you can do with it. This really opens up a lot of use cases that can become viable for on-device inferencing.

20:23Jon Krohn:Does this mean, are you able to kind of give an estimate of like the number of parameters of an LLM you'd be able to fit on there? Yeah, yeah. So 109 billion parameter LLM4 scout, right? But that's a speculative decoding model at FP16 cloud native resolution. Nice. And Ish, it sounded like you had something to add. No, only that, yeah, it's a pretty big honking model to put on the laptop. Oh, yeah. Again, me taking a very sophisticated comment from Shurish and distilling it down to base components here. This is bigger than anything you've played around with in anything LLM or LM Studio or OLAMA to date.

20:57Like, this is a, today you just call it, you know, kind of a cloud quality model, right?

21:04Jon Krohn:Not long ago, I remember sitting in meetings with data science teams that I manage where I would be trying to encourage them to be thinking about using 13 billion parameter LAMA models, for example, because that was about as large as we could fit on a GPU. And so that's pretty wild to be thinking about something that's almost 10x that size running on a laptop. Yeah. And more to come, you know, like this is this is sort of the progress that's been made in 2025 alone. And it's just staggering to see both the models getting smaller and the hardware getting more capable. Like those two curves, those two things are converging on each other.

21:44And it's not going to be one or the other. They're both going to happen together.

21:48Jon Krohn:Very exciting. So you've mentioned, we've now talked about NPUs in some detail now. You've mentioned GPUs. What is the key difference between a GPU and an NPU? It sounds like today, a lot of people would be using a GPU for all the workloads that you've been describing you should actually be using an NPU for. GPUs integrated and discrete, right, Trish? Yeah. So, you know, I'd say that GPUs are far more versatile today, right? Because they're ultimately the best at parallel processing. And so, by definition, they're very performant at AI and ML workloads, right? And we all know that because that's before the advent of the NPUs.

22:34That is the accelerator of primary value when it came to AI and ML. So the point I was going to make is what really sets GPUs apart is their ability to scale and performance relative to NPUs today. If you go back to what Ish was talking about earlier, also the GB10 and the GB300 that have the NVIDIA GPUs that are also coming out later this year. there's some pretty staggering capabilities within those boxes, right? So like the GB10, for example, is going to be capable of up to a 200 billion parameter model that can be fine-tuned locally, which is game-changing for your data science community. Now you don't have to set up a whole cloud instance.

23:24You have this appliance and your work group in a lab and you're on business, right? You can really play with some of the frontier models. So that's pretty game-changing. And that's fairly accessible from a price standpoint. I won't go into the details, but it's definitely more affordable than some of the alternatives. And then if you really want to scale that up, you have the GB300, which gets you up to 500 billion parameter models at 20 petaflops. But that's at a much different price point. So again, it comes down to what you want to do. But GPUs today have far more scale in terms of performance than NPUs.

Read the full transcript

24:03NPUs are still evolving. I think we're talking about first couple of generations. So as Ish said, they're going to get better. And the differentiator for NPUs is performance per watt. GPUs are not that conscious when it comes to power consumption, let me say. And that's a really nuanced difference, right? For GPUs right now, and we're talking about client devices here, right? There is a whole aspect of Dell that is in the server land. It lives up in the cloud, on-prem, hyperscale, whatever you want to do. But on client devices, the highest ceiling right now, GPU. On client devices, the best efficiency right now, CPU or NPU, depending on your use case.

24:54Speed also depends. What kind of model are you running also depends. And this is that matrix I was talking about, John. It went from fairly evident what you needed to do to, hey, I'm now thinking about buying a device here in August of 2025 or September of 2025. And I'm trying to figure out what the commitment of this device is going to be over four or five years, right? I mean, you can really see that lacking something with an NPU in it or lacking something with a discrete GPU in it, you may come to realize that you're going to need to upgrade sooner rather than later if you don't make that investment now.

25:39And, you know, Sharish and I are not part of sales. So like you buy one laptop, you buy a million, it doesn't really impact us, but that's the truth. And that's what I'd think about if I were making a purchase right now. Right.

25:50Jon Krohn:So let's talk about that next in terms of the kinds of things that people should be looking for if they want to be future proofing for the next five years. What are the things, what are the kinds of parameters? Let's go over an eight by eight matrix. in an audio-only podcast. But just kind of generally, let's talk about the kinds of things that people should be looking for in hardware that they're buying today. And I guess, as you said, this is specifically about what you described as client devices. And so I'm assuming that that isn't a term that I use in my kind of day-to-day language, but it seems to me like that's distinguishing against servers.

26:26Jon Krohn:It's like laptops, desktops. Yeah, most normal people aren't running around saying client devices. That is a very Dell kind of term. When you think about what to buy right now, like if I were starting college or if I were doing something, wow, God, that was a while back. If I was starting college today, thinking about what kind of thing do I need, right? And there are different brands and different price points and different pursuits that you would have with this device. What's it going to be used for? Yeah. I think an NPU makes a lot of sense for a lot of knowledge work type work, right? And if for no other reason than to get the most out of your operating system.

27:05We know from our friends in Redmond that Windows is going to start baking AI features into itself that it intends to run on the device, right? This stuff is expensive to ship to the cloud and back every single time. So some of the stuff like background blur on a Microsoft Teams call, speech to text, all of this stuff is going to look for a home somewhere on your device. And guess what? CPUs, the workload hasn't gone anywhere. That CPU is still going to have to do all of it. It's the workhorse. It's still going to have to do all the things it's always done. And now if you don't have an NPU or a GPU, it's also going to have to support this new kind of workload.

27:45So that's one thing to keep in mind where if you decide no NPU, no GPU, well, gosh, your CPU better have some slack in it. It better have some bandwidth. GPUs, I like to talk about the birth of a new persona. And persona, again, being a word that people in our world think a lot about, right? The data scientist persona is something that like an IT decision maker is constantly thinking about. Like, what does that persona need? And you really have the birth of a new persona with all this AI stuff. Because you have people like myself who are not formally trained in that way as engineers, but who know enough to be dangerous.

28:24And now with the right kind of device, I get supercharged. And with the wrong kind of device, I get throttled. So this is very much a productivity gains question. And that is sometimes really hard to quantify. So knowledge workers, NPU makes a lot of sense. Knowledge worker plus, maybe like these new persona at the edge of a dev and a kind of regular knowledge worker. That's me. and I would ask for something like a discrete GPU because I know that's going to last me. And also if you want a device that you can use to train AI workloads during the day and give your kid to play Fortnite later, like GPU is probably the way to go.

29:04So there's a dual use argument to be made there. Yeah. And just to add to what Ish is saying, I would classify them today. And again, you have to keep in mind, this is rapidly evolving. But today I could classify devices into three categories. You have the essential AI PCs, which have what I, for lack of another moniker, call them entry-level NPUs. Think 10 to 15 tops or trillions of operations per second. And those are great for basic workloads coming from, as I alluded to earlier, your background blur, your voice correction, and other optimizations. Offloading that from the CPU so you have a much better experience.

29:50And they can accommodate smaller models like up to maybe one to three billion parameters. But once you get there, now you're bringing workloads back to the CPU if you go beyond it. So that's probably the limit there. Then the second category is maybe slightly more advanced AI PCs with more performant NPUs or state-of-the-art NPUs. Today, that's about 40 to 50 tops. And that really brings on-device AI into focus. Right now, you can actually bring custom workloads, perhaps run up to 9 to 10 billion parameter models for custom in-workflow embedded use cases across a variety of verticals, in addition to the Copilot Plus features, which run locally on your PC that Ish talked about.

30:39So this is, again, a very nuanced difference here. Microsoft's co-pilot branding refers to everything that runs in M365 in Azure, right? And that's so that's all cloud-based, subscription-based, largely. That's their co-pilot brand. Co-pilot plus is everything that runs locally as part of the OS itself. It's part of Windows, no extra charge. and they can, you know, as Ish said, is going to continue and Microsoft is going to continue to add more and more capabilities that run locally on the PC. So for you to harness those capabilities and not lock yourself out of those capabilities in the future, you definitely want a PC, an AI PC with at least 40 tops on the NPU today, right?

31:33That is my recommendation for the knowledge workers and the most common use cases. And then the third one, it's kind of self-explanatory now. It's your high-performance PCs. Those have your high-end CPUs from the CPU suppliers, which are capable of much more performance single and multi-threaded processes. And then you have those augmented with discrete GPUs and discrete NPUs. Now you're talking about the persona that is talked about is you're starting to create that separation between your power users, your AI and ML and data scientists that can really now do data crunching and, you know, work with models right there on the device itself.

32:20So that's the three-pronged categorization today for AI-BCGs. I was the recovering consultant, and here is Sharish with his three buckets, right? Like, BCG would be proud. One thing, John, I want to add to that is like, we're not a walking infomercial here. And I know there's like a big corner of the internet, like, let's be real for a second. That's like, hey, like, I watched the NPU advertisement in the Super Bowl. Like, I watched the Copilot Plus PC ad with the zebras and the scientists in the forest, right? But like, really, what does it mean to me? And again, it's about this temporal mismatch.

33:01How long are you going to use this device, right? Oh, I'm skeptical of the features that this particular company is building. I'm never going to use any of those. Again, think about the future. Think about the things that are happening at breakneck speed, breakneck pace. That's what you have to be thinking about, that temporal mismatch. So even if it's not up to your tastes in this moment, there's something bigger to consider. And again, it's not about an infomercial. These are just the things that I would be thinking about if I were buying one device or if I was buying a million.

33:34Jon Krohn:Yeah, for sure. It's an interesting situation here because you guys obviously do both work at Dell. And so it's easy to feel like it's infomercially. But simultaneously, everything that you have been saying so far is useful to me as somebody thinking about. I mean, like, you know, we on the show, we have a lot of different kinds of episodes on different kinds of topics. You know, a lot of the time we're covering open source software. but you need your software to run on something and there's no open source hardware. You can't just download free hardware and have that going. Exactly. So by its nature, it has to be in some way kind of a bit of a commercial conversation but hardware is something that you need for any of the work that we're doing, any computing.

34:25Jon Krohn:um so speaking of software and kind of the experience of using these devices so if we're talking about having npus on our machine having gpus on our machine having cpus and i actually i want to dig into cpus in a bit uh briefly as well so i'm going to get your subconscious thinking about you know about cpus in general but uh when we have you know i think it's pretty easy at least in terms of utilization, the CPU is the most general of all these devices. And so I think we're used to, probably most listeners are aware that CPUs are doing the most kind of general work, running your operating system.

35:05Jon Krohn:And what I'm trying to get to here is it sounded pretty obvious from what you're describing that in the future, or maybe even today, Microsoft will have automatic support when you're using, say, Windows for NPUs. So you're talking about Copilot Plus features, text to speech, any of these kinds of things. They'll look for the best device available for running those. And if there's an NPU on the machine, it gets sent there. So I get that for the kind of click and point operating system experience. But if I'm a data scientist, if I'm a software developer, if I'm an AI engineer, then what do I need to do to engage those resources?

35:45Jon Krohn:Like, what is my experience like? What software do I use to get access to NPUs, GPUs, or be making the decision? How does all that work? Yeah, multi-part answer, I think, here. The first part answer is it's nothing you've not done before in some ways. GPU acceleration has been around for some time, right? You mentioned all the uses for the CPU. Like, yeah, I know of all of you and your 100 tabs that you refuse to clean and close. We know about you at Dell. We know you require some CPU juice. But GPU acceleration, it's been around. Solvers like Garobi have GPU support. CAD software, GPU accelerated.

36:29So the sort of tried and true methods of acceleration, the abstractions are sort of similar, right? And then you're going to start to hear about all kinds of funny name stuff. And it's really going to strike you as like, well, what is all this, right? You've got Llama CPP. You've got Ollama. You've got sort of to a lesser extent, CUDA has been around, but you've got this CUDA X layer now that sits in the middle. Like all of these things combine to have this plethora of stuff. and maybe not all of it is deployable at scale in like an enterprise environment. Like if you are a data scientist, sure, go pull this particular, go pull Lama CPP and run a GGuff model.

37:16Fine. You know how to do this. You will know how to do this. If all of these words are alien to you six months from now, maybe John will invite us back and they won't be as alien to people anymore. Right. But like the work that the folks at Unsloth are doing, the work that folks like the Bloke are doing. And these are all names you see on the hugging face boards of like these models getting rolled out. To land those on particular pieces of silicon requires that software layer. And if it's not a software layer, it requires sort of this abstracted API. Dell, and John, you opened this with, you are a devices company.

37:53That is what you are known for. And we know that that's our identity. To use the device to make the most of it, we've created something called Dell Pro AI Studio. It's a thing that lets you not have to worry about this as much if you are the developer of an application. I'll let Shurish talk a little bit more about Studio. Absolutely. So I think that was very insightful-ish. I learned something from that too. So that was a great monologue there. I wanted to add one more thing.

38:25Jon Krohn:Oh, God. Sorry. We got that. Shakespearean. We got to cut that one out. We are not cutting that. That was perfect. Absolutely. That is in the episode. Mario, do not listen to the guests. So, yeah, I just wanted to add that, that, you know, the complexity of being able to bring your AI workloads and land them on the PC accelerators is not a trivial task, right? Ish painted a really vivid picture there. and it's not for the faint of heart today, right? So if you think about someone who's beginning their journey as an AI engineer or who's trying to run AI locally on a PC, whether it's the GPU, the CPU, the NPU, they have to make sure that they have the right format of the model.

39:23And if it's not the right format, they have to have the right tool chain to convert it to the right format And that conversion process is different for every target CPU architecture, right? And IHV. So if it's Intel or AMD or Qualcomm or NVIDIA, you're looking at a completely different tool chain to go make that conversion. And then it proliferates from there because you're talking about one model, right? Now, that's not how models are typically. There are model families. So if you want to make your entire model family of, say, Lama available and you want to quantize them or do Lora's or something, now you're really creating some tremendous, you're creating a huge tree now, right?

40:10Massive branches. So that's a task that is daunting even for experienced practitioners. So with that said, that's why Dell Pro AI Studio became almost an imperative for us to create. And we focused on abstracting away all of that complexity. And the focus being, how can we democratize access to AI on the device for the broader development community? So you don't have to be working directly with Intel OpenVINO and be six months into your journey there before you can actually land something on the NPU, because that's what it takes today, right? And more on Dell Pro AI Studio,

40:56that's the paradigm, right? Reduce complexity, go faster. I'll talk about some of the speeds that we've accomplished, right? So we did a POC with a partner of ours, Deloitte, and there's a white paper coming out on this pretty soon. We did a comparison of a use case that was built for on-device AI, and they'd already worked on it before they used Dell Pro AI Studio. And then we normalized it to a team that was just beginning their AI journey, right? And so we kind of looked at it three ways. And what we learned was for the team that was just starting their journey with, say, Intel OpenVINO, it took them about three months to actually go from scratch to have an app that is fully integrated with a model that's actually running on the Intel Silicon.

41:56and it took the same team with Dell Pro-Ace Studio about four days to do that. So, you know, and that's not like a linear, you know, simplification of time. That temporal concept is really, we just eliminated that engineering complexity for them. So that's the second reason why we built it, right? We wanted to democratize access for the developers at large to bring AI workloads to our PCs. And we wanted to make it quick and easy for them to land it on a variety of silicon. And last but not the least, we also put in an API that actually is following the OpenAI de facto API standard for hosts. And so if you are running a web app that is pointed to either an on-prem workload or even a cloud hyperscaler workload today, but you're using that open AI de facto spec, you can just make a quick one-line configuration change, point it to the local server on a Dell machine, and you're done.

43:20It is now running on a local model that's running on your device accelerator. So that's a long-winded explanation for what Dell Pro AI Studio is, but hopefully that resonates with, you know.

43:34Jon Krohn:It all sounds great. One thing that I think might be helpful to me and to our listeners is to understand, I realize that, as you mentioned, there are lots of kinds of personas out there. But is there kind of like a user story that you could describe for us to kind of so that we can visualize, so that we can imagine what it's like for us to use Dell Pro AI Studio? Like, I think I get the functionality. It allows us to take advantage of all these powerful devices for, you know, whatever kind of AI workload we have. But what is it? What does it look and feel like? In terms of real life application, and this, John, your question is also super salient because it's also about the bigger conversation around why would you ever run a workload locally?

44:18Like, why would you ever do that to begin with? Like, as a data scientist, as a knowledge worker, as anybody, why would you ever do that, right? Offline mode comes to mind, right? If you want your AI-based application to continue working, even when the person's on an airplane and has a lousy connection, guess what? You're going to want to run that workload locally. Reasons related to speed, cost, security, connectivity, all of these signals are the reasons why one might run a workload locally. And examples of those, John, models are not apps and apps are not models, right? The data scientist universe knows this, right?

44:58Like credit card fraud detection is the objective. Within that objective, there are many models doing different things. It's the same for local apps. Today, there are a lot of them like LM Studio, anything LLM that make it easy and put a GUI on top of, I'm going to go in and I'm going to do AI work locally on my machine. Look through the app, abstract away the app and start to think about you as a data scientist. what is the purpose of the thing you are building? That's what the GUI is going to look like. On the back end, the user doesn't care. Is it going up to the cloud? Is it going down to my silicon?

45:38The user literally doesn't care. They're going to care about the performance and they're going to notice that there's a difference, yes. But that's why there's sort of this routing objective around when you run workloads locally. And that's why Dell Pro AI Studio makes it as simple as a one-line config change instead of sending the workload up. you send it down. You just change where it's pointed at. And that's the whole point of what we've made. Sharish mentioned OpenVINO a couple of times. Dell Pro AI Studio is an extension of OpenVINO. OpenVINO is in there. We worked with Intel to bring this thing to life and to make it easy to use.

46:16And as are the OpenVINO twins from other people who make pieces of silicon. right? So the whole purpose of defining the use case, it's not about running something locally just to say you did. Although I enjoy that and I spend a lot of weekends doing it, right? It's also about when and why does this make sense? Am I forcing this down someone's throat because I want to? Or is there genuinely a reason this data-related workload, this AI-related workload should happen on the device?

46:48Jon Krohn:You talked about OpenVINO there and working on Intel with it, which is a great segue into a topic that I wanted to make sure we covered in this episode. So the Dell AI PCs that we've been talking about, the hardware we've been talking about in this episode, we've talked about NPUs, GPUs that are available on them. And I already alluded to this. I said that you should get your subconscious going. We've talked about the purpose of CPUs as being relatively general purpose compared to NPUs and GPUs. But I wanted to specifically, because we have experts like you on the show, I don't actually know very much about CPUs or considerations about them.

47:25Jon Krohn:So for example, Intel Core ultra processors that were released in 2024 last year, there's this lunar lake architecture. What is the significance of that architecture, which I think is the standard now on AI PCs? What is the significance of that architecture for my listeners, for the data science persona? Yeah, so I'll take that one. A little bit about Lunar Lake. So it was quite a different architecture in which you had memory on chip. So quite unique from Intel. And so with that integrated memory, you actually had very fast transfer rates, you know, from processor to memory, because it's, you know, it's MOC.

48:13They are the first Intel architecture that supports 40 plus tops on the NPU. So, you know, if you're looking for a Copilot plus PC from Intel, you know, with an Intel processor, you're looking for Lunar Lake. That's kind of the highest level, right? Apart from everything else, well, it's still x86 based, so was nothing different there, nothing unique or different. It's not ARM, right? That was the biggest difference. The MOC is what allows for tremendous performance gains. And I will mention that their iGPU, their integrated GPU also achieved some step function performance gains, wherein they're actually rivaling or I would say exceeding the performance of some of the entry-level discrete graphics cards.

49:07Jon Krohn:Very nice. That was a clear definition as we've come to expect from you, Sharish. And now I do feel like I understand the significance of lunar-like architectures in particular, the kind of step gain efficiency improvements that you were describing there, like memory on chip. So I want to kind of bring all this together. We've talked about CPUs now most recently, obviously GPUs, NPUs earlier in the show. And we've also been focused on what you call client devices, kind of local processing, laptops, desktops. Let's talk about why all of this is significant in the Gen AI era. So for example, at the time of recording, not too long ago, a couple of weeks ago, OpenAI released open source LLMs for the first time in a while.

49:53Jon Krohn:Earlier in this episode, we talked about llama models. So what are the kinds of things that gives us the capability to be taking these open source models. You were talking earlier in this episode about laptops that can have 109 billion parameter models on them. You're talking about desktop appliances like the GB10 that could have something even double that size running locally plugged into your machine. And so in this brave new world where these kinds of gigantic, highly capable open source models are available, where we can be downloading them onto local advices, what are the kinds of circumstances where you'd recommend that versus doing something in the cloud?

50:39Yeah, use case dependent. And I know that's a frustrating answer because it's the same question that a lot of Dell customers ask us. And to come in and to tell them that here's the formula you're going to use, I wouldn't be being honest with them because when I decide to run something locally, it's because I've analyzed the situation, determined what it is I wanted to do and picked local as the thing that makes most sense. Let's say I'm working on something that contains a lot of personal information that I don't feel comfortable putting into a cloud-based AI tool. Or let's say that I am subscribed to a cloud-based AI tool, but it's close to the end of the month.

51:17I've burned a lot of my tokens and maybe I don't have many of them left. All of those are reasons why I would turn to something like what you mentioned, John, right? The GPT OSS family of models. I think it's super telling that a company like OpenAI, which has to be very intentional about what it does and where it spends its very valuable human capital resources, chose as part of their open source launch to have one of the models be a smaller model. I think that is really telling because it tells me that in this agentic world that's coming, where it's not just chat back and forth anymore. It's like, by the time you have your AI tool, your 50th poem, you're kind of like over it.

52:02And it's like, all right, data scientist or otherwise, what is the applied practical use of this tool? Like, CodeGen, creating little applets to help you do the things that you do on a day-to-day basis, and it's not something you want to go subscribe to. Local is a phenomenal application of that. Right? Like, so determining that I'm going to do my code gem locally versus in the cloud is a question of how much do I want to spend on it? How fast do I need it to be? How capable do I need it to be? For me, a person like me, that local use case makes a lot of sense. I'm not burning through tokens and watching my credit card, right?

52:39So it is a frustrating answer because it's about choice and that agency that you have to decide where to run what and when. All the way at the top of this conversation, John, we talked about Windows versus Unix versus Linux versus et cetera, et cetera, et cetera. Choice has remained the theme here. Choice of chipset, choice of operating system, choice of local versus cloud. there will be a right tool and a right method at the right time. So I know that that's not the answer everyone wants. And it's always met with this crestfallen look in the room when we talk to our customers, because it's like, I thought you were going to tell me what the answer was.

53:21The answer is that this is hard. We got to do the work.

53:25Jon Krohn:There were some clues in there, though, things like data privacy edge you in the direction of not doing things on the edge. I'll add a spin to it. I think the other thing to consider is, apart from I was going to bring up privacy, and you nailed it, John. That's an important consideration. We're seeing more and more of that concern, right? Like sovereign AI, enterprise AI, they're almost anonymous now. where customers, they're trying to get to initial value as quickly as possible to start their AI journey, right? Because they want to show outcomes, show value. And sometimes the answer is, let's just go to the frontier models in a hyperscaler environment, right?

54:16Because it's the easiest way. I don't have to scale up my team. I don't have to invest a whole lot and increase my time horizon. I can get there quickly. But they're also trading off control. They're trading off, you know, they're at risk of getting locked in to the ecosystem. And they're also trading off security, right? And privacy of data, especially IP, which is starting to become a big deal. And we hear about policy decisions being made across the world, right, that are changing things or putting guardrails on what governments can, cannot do. And then you have, of course, private enterprises that they have their own sets of concerns, right, and the regulatory compliance requirements that they have to adhere to.

55:09So I do think that from that perspective, if you think about cloud versus local, to me, the answer is hybrid, really. The future is hybrid because it is not practical or feasible to move all workloads to the PC. And at the same time, you want optionality because you don't want to be locked into just one type of node. You want to be able to run the right workload on the right compute engine at the right time. So I foresee there being intelligence in an agentic future where there are either basic ML classifiers or much smarter agents that are able to decide where a workload goes and then just be able to use distributed nodes, right?

56:03Whether it's the PC fleet, a GPU cluster on a data center or wherever. So I think that's the future, to be honest. Most people have not thought about computation as a resource like this in a long time. Like as a token to be spent, as a dollar to be spent, like how many units of compute do you have? And now on YouTube, you see every which person stringing together 50 of the same device stacked all the way to the ceiling. Look at this node I made, right? There's a particular Formula One team that Dell works a lot with. Particular team that happens to be winning, I'll add. A particular team Dell works a lot with.

56:41and they are very big on sandbox experimentation. They're very big on like, I don't want to worry about, like, I just want to test something quickly. I have an idea, I have a hypothesis and this is data science we're talking about. This is, you know, thousands of points of telemetry being captured off of these cars, rows upon rows upon rows of data. I just want to see something real quick. I don't want to have to worry about, Do I have to spin up my instance? Is there a queue? Are we like, just give me a device. Let me go look inside. But this file is much too big to open in Excel and scroll to row 750 to check a particular telemetry point.

57:25Right? Like sandboxing is another big reason why this stuff on device is going to matter. And even if the rest of it isn't clear, experimenting with local AI to see if it is part of your workflow or if it is something that's useful to you, that's the only way you're going to find out you got to do it really cool and as an f1 fan

57:43Jon Krohn:certainly is not a bad time to be affiliated with that particular team their dominance this season is pretty insane yes powered by

57:54Jon Krohn:nice i have one last technical uh question for you that i just like to squeeze in quickly there's a uh there's a big windows refresh coming up uh in a couple of months in october well, I guess at the time of this episode release in a month. And so it sounds like a lot of people might need to be thinking about upgrading their hardware now, especially if they're using a PC. And so tell us a bit about this big refresh. And that'll be kind of where we end the episode. I'll ask you both for a book recommendation after so you can start thinking about that in your subconscious. But otherwise, yeah, let's wrap up with this Windows refresh.

58:32All right. So let me go first and then I shall, you know, we'll get your take. Since, you know, I'll keep an audience conscious. So let's talk about Dell Pro Max first. Right. And if I even look at it broadly, just data scientists and developers. something to look forward to is the new line of pcs that are equipped with nvidia's blackwell rtx pro gpus um really game-changing i think looking at some of the capabilities right you can run uh it's talked about code gen a while back you can run a model like dev stroll small which is the 24 billion parameter one at FP32, you can run that completely locally on the RTX 6000 Blackwell GPU, which is doubling the memory for VRAM.

59:34It's 96 GB VRAM, which is allowing you to do these kind of models, you know, work with these kind of models locally. The other cool things about these devices is, you know we think about the power that these are consuming so these are triple fan cool designs so massive thermals uh investments in thermals and acoustics um to give you that in you know experience so if you're the the cue to your listeners is if you're trying to bring some of these newest sota you know state-of-the-art models and run them locally and you try to do it on the previous gen product, chances are you either just can't do it, or if you're able to, depending on the model of choice, you're going to have pretty much like an airplane on your hands.

1:00:23That's what it's going to feel like. So I strongly recommend your listeners to consider the newest line of performance desktops and laptops. One other thing I'll tell you is, one of my colleagues was running some quick tests. And what he found was for like a video workload, right? Just think of like animation. They were working with a large studio and they were trying to create, based on previous artwork and shows content, they were trying to train a model that would keep the style consistency and be able to deliver sort of stage up the content in the same milieu and you know so different they were trying to create mock-ups with fine-tuned flux lauras and in the previous gen gpus it was taking approximately a week and with the uh with the blackwell gpus they're getting it done in about three days, right?

1:01:31So that's pretty measurable in terms of performance gains that you're getting. So just talking about that very specific persona, I would say, take this opportunity to, you know, with this Win 10 end of life, if you're looking for a new device, definitely look for the latest and greatest, because the gains to be had will set you up for, you know, the next four or five years to come. Excellent.

1:01:57Jon Krohn:I can't believe how you can reel through those technical stats, seemingly without notes. Pretty mind-blowing, Sharish. Ish, how are you going to follow that up? Yeah, I'm going to do this and I'm going to whack some people in the head, right? Do not let this deadline pass you by for the love of God. This is not a selfishly motivated thing that is coming out of my mouth right now. This is a look when operating systems age out, there's a reason they're aging out. There's new stuff coming from a security perspective, from a manageability perspective. perspective, if you have been procrastinating, please stop procrastinating, whether you buy a Dell device or not, just please make sure you understand the gravity of an OS refresh and make sure you're not hanging on to something that might be best retired in its 17th season.

1:02:53Right. So that's, uh, I think the only thing I can add to Sharish is like, Hey, like October is here. Like it is no longer this big thing in a future horizon. Like you need to solve this problem now because it is a problem.

1:03:05Jon Krohn:Nicely said. And by the way, for our listeners, so Ish, when he was hitting you on the side of the head, he was flicking his microphone, which he probably presumed would create some kind of like loud sound or whatever, but there was nothing. Oh, man. Just kind of a pause. So yeah, fantastic. Dang noise cancellation. Yeah, exactly. Too good. Too good. The AI features have gotten too good, guys. Yeah, yeah. Ish and Sharish, This has been a fantastic episode. But yeah, before I let you go, I would love to get a couple book recommendations. Ish, let's go with you first. There's a great book called Deep Learning Illustrated, John.

1:03:46John's book, of course, being a great one. I took a course at Sloan called The Analytics Edge. And it took me as a person from, you know, didn't know anything about anything in terms of data science. And it gave me what I thought was a great starter foundational toolkit for a business and strategy oriented person to become familiar with applied data science and how to use it. My book? No, The Analytic Search. No, no, you're talking about The Analytic Search.

1:04:17Jon Krohn:So that's like a course that people can take. It was, yeah. I thought maybe you were saying that my book was like the companion text for that course or something. No, honestly, it should be, though. Having looked at both, I think it should be. So we should have that conversation and see who we can bother over at MIT to get that used. But no, it's about feeling like this is within reach for a lot of people. You don't have to feel like AI and data science in 2025. Like, gosh, there's these 16-year-old wonder kids who are out there. Like, for me to start now, mid-career, trying to learn this stuff, why even bother, right?

1:04:56So on the technical side, go check out the Analytics Edge. It made me a somewhat competent human being, which I was less so before I read it and before I took the course. The second thing I'll say is I have a renewed interest in fiction in this world because what these models are doing for creativity and what these models are doing for content and storytelling more broadly is a really fascinating element to this that I spend a lot of time thinking about, which is if everything is derivative of something older, like, is that really a new phenomenon or has it always been that way? So I'm actually blazing my way right now through the Reacher series, like the Jack Reacher series, and inspired by the Amazon show.

1:05:41I never actually picked up those books and read through them. And as I'm reading them, I'm also thinking about, gosh, what is the future of fiction going to look like? So that's my plug for Lee Childs.

1:05:51Jon Krohn:Very nice. I like the fiction recommendation there. And just so that I'm understanding, So The Analytics Edge is both a course and a book. Correct. And so, okay, yeah, great, perfect. And we'll take a look at that, John, after we hang up just to make sure that I'm not making something up. But yes, it's the short answer. We'll have something in the show notes for you listeners. And Sharish, do you have another book recommendation for us? I do. I'm actually reading Seven Powers by Hamilton Helmer. Fantastic. It's, you know, I think it's a very fresh take on business strategy and, you know, sustainable, long-term profitable companies.

1:06:38So I highly recommend it. I'm still reading it. I'm not done with it, but it's brilliant. The second one is actually one that I picked up based on your recommendation. It's the Hands-On Machine Learning Guide by Aurelian. which I haven't started yet, but it's sitting on my desk. So it's going to be my companion as I revisit some Python after years.

1:07:03Jon Krohn:Yeah, so that will actually be, so this episode that we're filming right now, it will be a week after. We just had Aurelian Juran on the show. He's the best-selling machine learning author of all time with his hands-on machine learning series. And yeah, so thank you for that nice cross-reference there, Sharish. Fantastic. Thank you both so much for being on the show. This is another wildly informative episode from you, Sharish. And Ish, I really enjoyed getting to meet you as well. You do have so much color to add and you can't hide that intellect. I'm going to tell your mom that you do have it.

1:07:42Coming from you, John, that's going to mean a lot. Also confirmed, Analytics Edge is a book, not making it up. I wasn't hallucinating the years 2018 through 2020.

1:07:52Jon Krohn:Nice. All right. Thanks so much, both of you. And I wouldn't be surprised if we're hearing from you again on the show sometime soon. It's been awesome, John. Thank you. Thanks, John. Fantastic.

1:08:22Jon Krohn:use, we can locally train models of up to 200 billion parameters, how NPUs are optimized for AI workloads with superior performance per watt for battery efficiency, while GPUs offer higher absolute performance and scalability today. And lots of other considerations around cloud versus local AI work and what hardware you might want to use for either situation. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for my social media profiles, as well as my guests at superdatascience.com slash 921.

1:08:59Jon Krohn:Thanks, of course, to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, media editor, Mario Pombo, partnerships manager, Natalie Zajski, researcher, Serge Massis, writer, Dr. Zara Karche, and our founder, Kirill Arimenko. Thanks to all of them for producing another super episode for us today, for enabling that super team to create this free podcast for you. We're deeply grateful to our sponsors. You can support the show by checking out our sponsors links in the show notes. And yeah, otherwise you can help us out by sharing the episode with people who would value it, reviewing the episode, wherever you listen to or watch podcasts, subscribe if you're not already a subscriber, but most importantly, I just hope you'll keep on tuning in.

1:09:41Jon Krohn:I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come. Until next time, keep on rocking it out there and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

From the publisher

Using Windows for AI development and the bleeding edge of NPUs: Shirish Gupta and Ish Shah from Dell Technologies speak to Jon Krohn about the latest products from Dell, the future of neural-processing units (NPUs), and how AI developers can make sound hardware investments. 

This episode is brought to you by the Trainium2, the latest AI chip from AWS, by ODSC, the Open Data Science Conference and by Gurobi.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/921⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(04:18) Why Windows still outranks other operating systems

(20:58) The difference between GPUs and NPUs

(32:44) How to access and use Dell’s NPUs and GPUs

(49:08) Using processing units on the cloud versus locally

(57:43) About the Dell Pro Max

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
921: NPUs vs GPUs vs CPUs for Local AI Workloads, with Dell’s Ish Shah and Shirish GuptaSuper Data Science: ML & AI Podcast with Jon Krohn · 1 h 12 min
Listen in VO