1031: Tokenomics: Why Your Agentic AI Bill Is Exploding (and How to Fix It), with Tyler Cox and Ish Shah

29 Sep 2026 · 1 h 15 min · 30 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Agentic AI tokenomics—why agent systems consume far more tokens than chatbots, and how enterprises can cut costs by moving workloads from per-token cloud APIs to on-prem or “desk-side” hardware.

Guests

Ish Shaw and Tyler Cox, distinguished engineers in Dell Technologies’ Office of the CTO for the Client Group. They work on running powerful AI models on machines closest to users.

Key claims

  • Agents use tools in a loop and can act beyond the chat window (e.g., moving through tasks, spawning sub-agents).
  • Token usage is exploding due to agent/sub-agent parallelism and “Jevon’s paradox” (capability improvements drive higher usage volumes).
  • Enterprises are deploying agents widely (Dell cites 90%+ of enterprises using agents in some form).
  • Local compute can deliver major cost savings: Dell/Signal 65 studies report up to 87% lower cost vs cloud; some workloads pay back in ~2 months; Dell lab example cites ~$160k/year cost avoidance and up to 93% savings for ~20-agent deployments.

Notable examples

  • Ish’s weekend project: building a Pokémon fan game; sub-agents and supervision loops drove multi-billion token burn.
  • Dell Pro Max GB300 hardware specs discussed (Blackwell Ultra GPU, liquid cooled).
  • “METR” model evaluation charts and Pareto-curve model selection (speed vs intelligence vs cost).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Agentic AI

0:57 to 3:18

Ish and Tyler discuss what makes an AI system agentic and how it differs from traditional models.

“This episode of Super Data Science is made possible by Dell Technologies, NVIDIA, Anthropic, Origin, Palo Alto Networks, and the Open Data Science Conference.”

The Rise of Agents in Enterprises

3:18 to 6:15

The hosts explore how prevalent agentic systems are in enterprises and their various applications.

“And we actually have this conversation a lot.”

Challenges in IT Departments

6:15 to 8:27

The discussion covers the push-pull between enterprise needs and IT department constraints regarding AI.

“Like similar vibes, right, with some of our big accounts.”

Workflow and Productivity with AI Tools

8:27 to 12:20

Ish shares insights on optimizing workflow using multiple AI tools and managing tasks effectively.

“How am I supposed to reconcile that as an IT leader with the reality that is like AI in the enterprise?”

Choosing the Right AI Tool for Tasks

12:20 to 14:00

The hosts discuss criteria for selecting the best AI tool for specific tasks and their experiences.

“It's like, you know, if you're our boss, Tyler and my boss, hello, Mr.”

Introduction to Model Evaluation

14:00 to 14:48

Learn about the importance of model evaluation and benchmarks in AI.

“Fable 5.1, which is their latest model, Anthropics latest model, tends to be at the top end of a lot of the leaderboards for various kinds of tasks.”

Understanding Model Evaluation Protocols

15:44 to 18:00

Explore how to evaluate AI models effectively for enterprise use.

“And then, you know, that's something that's internal to you and you can be pretty confident the big labs aren't trying to optimize for your particular internal task.”

Exploring Pareto Curves in AI

18:01 to 20:06

Learn how Pareto curves help in selecting AI models based on performance.

“So what we typically will do for Pareto curve is look at how fast is it, right?”

Tokenomics Explained

20:07 to 23:28

Understand tokenomics and its impact on AI model performance and cost.

“And Sharish and I talked to you about this last time, John.”

Cost Efficiency in Using AI Models

23:29 to 28:00

Discover how to achieve cost efficiency while utilizing AI models effectively.

“That is a system we purchased one time that we are operating with open software, open models for free in perpetuity.”
Show all 30 chapters

Building a Personalized Pokemon Game

28:00 to 29:40

Ish shares a creative project involving a fan game based on his life and pets.

“Like Ish, how did you burn through billions of tokens on the weekend on a side project?”

The Increasing Efficiency of AI Agents

29:40 to 35:10

Discussion on how AI agents work together and the implications for token usage.

“So now you've got like an agent in charge of a bunch of other agents, right?”

Cost Efficiency in AI Workloads

35:22 to 39:20

Exploration of cost savings when using on-premise solutions versus cloud services.

“What are the potential kind of cost savings there?”

Potential Savings from AI Infrastructure

39:20 to 42:06

Discussion on significant cost savings associated with different AI infrastructures.

“So we worked with Signal 65 team across the Dell Technologies portfolio, right?”

The Evolution of Desk-side AI

42:06 to 46:21

Discover how advancements in AI software have simplified setup and functionality.

“into this concept we've been playing with around desk-side authentic AI, it used to be really hard.”

Tokenomics and Cost Efficiency in AI

46:21 to 49:10

Learn how AI hardware can achieve cost-effectiveness and quick returns.

“Um, it sounds like with these, with these options of having our own hardware, whether it's desk side or on-prem, it sounds like we can, uh, get a return very quickly.”

Desk-side Agentic AI Solutions Overview

49:24 to 54:09

Understand Dell's desk-side agentic AI solutions and their applications.

“But do you have kind of one big takeaway for me on the tokenomics conversation that has enriched our conversation so far?”

Hardware Tiers for Desk-side AI

54:09 to 56:00

Examine the different hardware tiers available in Dell's offerings.

“That I want to go put a GB10 or a T2 and say, hey, go nuts, right?”

Introduction to Edge Deployments

56:00 to 56:55

Learn why organizations are moving towards edge deployments for sensitive data.

“or tens of thousands of agents working on one problem together and experimenting to try to find what the right answer is, right?”

NVIDIA's NemoClaw Overview

56:55 to 58:02

Discover the key features of NVIDIA's NemoClaw and its importance in AI.

“And let's now move from hardware on to software.”

Use Cases for Agentic AI Solutions

58:02 to 1:01:15

Explore practical applications of agentic AI in various fields such as healthcare and academia.

“A lot of this stuff has been increasingly possible over time.”

Understanding Token Explosion

1:01:15 to 1:02:33

Uncover the reasons behind the token explosion phenomenon in enterprises.

“And it's not just big industrial use cases, John, right?”

Tokenomics Solutions and Savings

1:03:30 to 1:07:31

Find out how desk-side AI solutions can lead to significant cost savings in enterprises.

“And this works, as you said, across different kinds of scales.”

Simplifying AI Integration

1:07:31 to 1:09:14

Understand how easy it is to implement the discussed solutions in various organizations.

“and getting the same kinds of results faster in a lot of cases.”

Book Recommendations and Closing Thoughts

1:09:14 to 1:10:00

Get insights into recommended readings and wrap up the episode.

“I think last time I cheated and I gave you two.”

Exploring AI Personalities and Ethics

1:10:00 to 1:10:29

Discusses the importance of understanding key figures in AI and the emerging ethical considerations.

“And it's got names of folks that at this point are canon for anyone who is interested in this industry.”

Book Recommendations and Closing Reflections

1:10:31 to 1:11:06

Hosts share personal book recommendations and reflect on the informative nature of the episode.

“And I think book number 30 is coming out soon.”

Follow-Up and Social Media Engagement

1:11:06 to 1:12:20

Hosts discuss how listeners can engage with guests and their work on social media.

“For people who want to risk it, risk more danger in the future, how should they follow you, Tyler, Del, about what's going on?”

Key Insights on Token Consumption and AI Costs

1:12:22 to 1:13:11

Discusses the reasons behind the surge in token consumption and cost-saving strategies in AI.

“In it is Shaw and Tyler Cox covered why token consumption is exploding.”

Disclaimer and Methodology Overview

1:13:11 to 1:13:35

Presents a disclaimer about the findings discussed, including methodology and assumptions.

“Outcomes may vary based on workload requirements, deployment environment, system configuration, and implementation approach.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:Over a single weekend, one of today's guests burned through two billion tokens building a video game for his wife. He and his colleague are here to explain why agentic AI bills are exploding and how a box under your desk can cut them by up to 93%. Welcome to another episode of the Super Data Science Podcast. I'm your host, Jon Krohn. Today, I've got two returning guests, but they are on the show together for the first time. Those guys are Ish Shaw and Tyler Cox. Both are distinguished engineers in the office of the CTO for the client group at Dell Technologies, where they work out how to run powerful AI models on the machines closest to you.

0:39Jon Krohn:In this episode, they dig into tokenomics, why agents and their sub-agents devour so many more tokens than chatbots ever did, how to pick the right model for the job, and how moving agentic workloads off paper token cloud APIs and onto your own hardware can pay for itself in as little as two months. Enjoy. This episode of Super Data Science is made possible by Dell Technologies, NVIDIA, Anthropic, Origin, Palo Alto Networks, and the Open Data Science Conference. Ish and Tyler, welcome both of you back to the Super Data Science podcast. You've been on separately with different guests, never together.

1:16Jon Krohn:I think this is a dangerous combination. and I think everyone knows why. Let's start with Ish on why this is so dangerous to have both of you together. Good to be back, John. Thank you for having us back. I think that Tyler and I have spent the last year out of our CTO office at Dell working on a lot of the things that we talked about the last time I was on with you, John, with our friend Sharish and Tyler was also on with Sharish. Our job is to think about the client device. And I remember the first time we talked to you about client devices, you're like, hmm, laptops? I'm like, yes, yes, all end user compute.

1:57And I think the role of those devices has changed a lot over the last, call it 12 months since we last talked to you. So just as a refresher, Tyler and I are both distinguished engineers in the office of the CTO for the client group at Dell Technologies. And we've got a lot to talk about. So I hope you're ready. I hope the audience is ready, John. I think we've got a lot to cover.

2:22Jon Krohn:Tyler can confirm this, or maybe Ish can confirm this about Tyler, but we should let Tyler speak, which is that Tyler is full of detailed facts about things. And then Ish provides great color commentary. That's what they pay me the big bucks for. And by big bucks, I mean not that. Yeah, we do a little bit of play-by-play. I think last time we were talking about linear attention and the rise of hybrid and state space models. And we've only gone from there. Oh, yeah. That's right. That was fun. That was great. We'll have links in the show notes to previous episodes that Ish and Tyler have been on separately.

2:55Jon Krohn:And they are both exceptional. This one, multiplying them together, as I said, dangerous combination. It might be too much for one podcast, but we're going to try to make it happen. We're going to do our best to contain the danger. What was it used last time? The Ishiness of it all? Ishiness. Yeah. Yeah, ishiness. It's a tongue twister. All right, let's get into the technical content here. So this episode is about agentic AI, tokenomics, solutions to get you better results faster, cheaper. But for those of our listeners who aren't totally sure, or maybe it doesn't even hurt to get this definition back to you every once in a while, what makes an AI system agentic versus a traditional model or a standard chatbot?

3:45I think it's the ability to use tools. And we actually have this conversation a lot. I mean, even within a company like Dell, there's like a very broad range of folks and how deep they've gone on this technology. And I would actually say like, as a company, we're pretty far ahead of a lot of others in terms of the baseline sort of AI literacy. That phrase, agent, still confuses the heck out of people. Like if I'm using ChatGPT on my app, is that an agent? If I'm doing something like what Tyler does day in and day out, which looks like the matrix flying across his screen, like, is that an agent?

4:22To me, and I think to a lot of people, an agent is when you take the LLM and you give it abilities that allow it to kind of break out of the tab, so to speak. It can act beyond the tab. It can do things with other data that lives elsewhere. It can do things on your computer and move your mouse around and click on things for you. That to me is the line of like agentic versus, hey, I'm having a conversation with an LLM and it's like writing a poem for me versus like, go check my inbox for this esoteric piece of data that I saw once six years ago and then do this other thing with it on this other website.

4:59That's the difference.

5:00Jon Krohn:Yeah, I like the definition of tool use in a loop to achieve a goal for agents. let's because i assume probably most listeners have a vague idea in their own minds of what an agent is so we can probably move swiftly on from that to how widespread are agents these days like let's say in enterprises how how many enterprises are using agents over 90 percent in our experience and according to signal 65 which is a research firm that dell does a lot of work with we are seeing over 90 % of enterprises in some way, shape, or form have deployed agents into day-to-day use or production use. And they define these things differently, obviously, enterprise to enterprise.

5:43But at this point, unless you have sort of a literal reason that you wouldn't or couldn't, I think the thread that's been pulled through is a lot of folks have been using this stuff in their personal lives. And I think they show up at work and they bang on the door of Mr. services, IT decision maker, and they say, hello, I would like this thing, please, for work purposes. And I think that that's how that number got so high so quickly, John. But Tyler also spends a lot of time with customers. Like similar vibes, right, with some of our big accounts. Yeah, definitely. And what's interesting is the number of different ways we see agents getting used, right?

6:24Software development is obviously a huge use case for agents having the ability to go off and build new tools and software. But we're seeing through our customer engagements a lot of really interesting and exciting different uses of agents, right? You're doing lots of research tasks. You're doing lots of knowledge work, aggregation of data, creation of reports and things like that. But you're also seeing more kind of domain-centric pieces in healthcare and finance and other pieces.

6:56Jon Krohn:Really exciting. there is a huge amount of potential. I feel like we're only scratching the surface of what can be improved in organizations, whether they're enterprises or not. And I think a lot of our listeners will have that experience. When you're using Cloud Cowork, Cloud Code, lots of desktop tools, OpenAI products, Grok products, there's tons of different options out there for you to be using these kinds of tools personally. But as you alluded to there, there's a lot you can run into a lot of roadblocks within organizations from in particular the IT department. There's a joke that I heard a while ago, which is probably rude to IT folks out there.

7:38Jon Krohn:Sorry, IT folks listening to the show. But there's this quip that if IT could remove your keyboard, they would. And it makes sense, right, John? Like these are the people you ask to protect these systems and keep them running at like 99.999 % uptime. like they got to keep the org moving. They got to keep the org going. And it's like, it's a risk first mentality because it has to be. So there's this very organic push and pull happening inside of enterprises right now, where from the top down, take a CEO who read a thing in the wall street journal to the bottom up, take a, you know, independent contributor who just finished their internship.

8:15Those two folks are, holy cow, look at this thing that I did. and they're sort of converging the IT person in the middle. So the IT person from the top down and the bottom up, there's signal in the noise. And they're like, okay, how in a world where I am being told to do more with less every single year, smaller budget, fewer people, but keep my 99.9999 % bulletproof uptime, I don't want anyone to lose a day of productivity. How am I supposed to reconcile that as an IT leader with the reality that is like AI in the enterprise? How do I do those two things? And for us, we spend a lot of time talking to our customers because this is what keeps them up at night.

9:01And the answer is that you have to do a really good job of explaining the types of things we're talking about on this podcast here to your superiors, to the folks in the C-suite of like, this is not your grandpa's IT, right? Like this is the name of the game is changing. And the costs that come with that sometimes are not ones that people have been used to for the IT department spending, right? So IT procurement, pick your poison where it lives within your org.

9:32Jon Krohn:For sure. Speaking of spending, we're going to talk about tokens in just one moment. But before we get there, apologies to Tyler that I'm letting Ish talk way too much here. And I'm going to give him even more floor. Though there might be some aspect of this that you have to add to this, Tyler, that we didn't talk about before hitting the record button. But just before we hit the record button, we were talking about how well. So Ish brought up on screen, he said, oh, I need to bring up my buddies. and then he said, okay, here's my Grok buddy, my Codex buddy, my Claude buddy. And I said, what? What was going on on your side, on your monitor when you were saying those words, Ish?

10:08So, I mean, a workflow is such a sensitive thing. It's such a personal thing. Like when it's time for John to do work or Tyler to do work or Ish to do work, there are things we reach for, right? You reach for your favorite pen, you reach for your favorite mouse, you reach for, if you're my wife, you reach for your favorite snack, right? It's like, depending on where you are and what industry you're in, your workflow is going to look different. But my workflow has changed so much over the last 12 months. Because it's almost like I'm on a construction site and I'm constantly reaching for, like, what is your handiest, dandiest tool?

10:43And there's a very nice, these are called multiplexers, but to normal people, they take lots of screens and they put them into one screen, right? It allows you to run these different agents side by side. Because what my workflow now looks like a lot is I'll kick a task off in one of these panes and my little buddy, so to speak, will go run off in the world and try to do the thing I've asked it to do. Well, I don't just want to sit there and twiddle my thumbs, right? I want to move on to the next thing. So I'll spin up the second session and set it off on the second task and the third and the fourth and so on and so on.

11:20And your workflow is so important because what really gets you is when one of these sessions hangs and you kicked it off in the morning and you come back at noon and you look at it and it's been, you launched it at 9am and it's been stuck since 9.01 asking you to approve this itty bitty prompt, right? Like, hey, are you sure you want me to do blank? That's why it's so important to like get your workflow set up right. It's so you don't burn. Like there's a lot of productivity that happens when humans are away from keyboard now. And my setup is a big part of it. So that's hence the little buddy.

11:56But it turns out when you've got four of them running at once, that starts to like put new kinds of loads on your system. And there are all these like downstream effects of what a modern workflow for a person looks like.

12:08Jon Krohn:We're going to get into those downstream effects in a second. And I did also now think of a great question, Tyler, to ask you related to this conversation that is going to lead us to tokens momentarily and token usage. But really quickly, before we get there, one last one for Ish, which is you alluded to before we started recording that you have rules of thumb for why you choose a particular buddy for a particular task. So why do you, you know, for your usage on your personal, well, I guess your personal work machine, if that makes sense, or your personal machine, whatever, how do you choose when you're going to use Claude or Grok or Codex?

12:47So they're all good. It's like, you know, if you're our boss, Tyler and my boss, hello, Mr. Rob Bruckner. If you are Tyler and my boss, right, and something crosses your desk, you're going to think to yourself for a sec, is this a Tyler job? Is this an ish job? a sure-ish job. Who does this go to? It's the same for me. I have opinions and preferences on if I want speed, for example, of late, Grok seems to be fastest. If I want depth and thoroughness, for example, right now, Codex is one I trust a lot to go tackle these. It's like a rottweiler it just goes after the thing until i tell it to stop sometimes to a fault and it burns all my tokens and i have to i have to sit quietly for a couple hours until it resets um and then claude which you know arguably claude is responsible if chat gpt was responsible for that first inflection point in like bringing people into this world of llms and ai i would argue like claude and claude code over the last 12 to 18 months are it's one of the biggest drivers of adoption into the enterprise.

14:00So for the things Claude is good at, I turn to Claude and Fable 5.1, which is their latest model, Anthropics latest model, tends to be at the top end of a lot of the leaderboards for various kinds of tasks. So artificial analysis is a very good website. If people haven't checked it out, they should. Very easy to understand rankings and benchmarks. And if you have an image task, you might want to consider this one. But people also need to be mindful of bench maxing, which is all of these labs are now training their models to beat the benchmarks against which they are tested. So it's a lived experience question.

14:35And that's how I determine.

14:38Jon Krohn:For all you listeners who want to level up your AI career through hands-on learning, ODSC AI West, October 27th to 29th in San Francisco is the place to be. ODSC AI West is my favorite conference and what sets it apart is it's all about doing. You'll gain practical skills by working directly with the latest AI tools and frameworks in immersive hands-on workshops and tutorials led by experts who are actually building and shipping AI. I myself will even be doing a keynote at ODSC AI West this year on how individuals and organizations can thrive in the agentic era. The full program covers where AI is moving now, including AI engineering, AI-powered software development, physical AI, robotics, and data science.

15:17Jon Krohn:Beyond the training, odsc ai west brings the ai community together with networking events meetups the ai expo and more giving you the chance to learn practice and connect all in one place super data science listeners can use the code super at checkout on odsc.ai for an additional 15 off your pass see you there odsc ai west october 27th to 29th in san francisco for sure and as an enterprise at least this would be kind of probably it might be overkill for an individual, but to get around that artificial analysis or just general bench maxing that all of the vendors are doing, all of the model creators are doing, the best thing to be doing, if you have some repeated task that you're going to want to be automating in your organization, you've got to have a great eval for that.

16:08Jon Krohn:And then, you know, that's something that's internal to you and you can be pretty confident the big labs aren't trying to optimize for your particular internal task. And John, to that point, I'm going to turn this one to Tyler because Tyler has built, along with several very talented technologists within Dell, a model evaluation protocol framework that we use internally, that when these models come out, particularly open source models, open weight models that are capable of running on Dell devices locally, like Tyler used to in the beginning, be up late at night trying to like benchmark all of these, like, hey, how does this work?

16:45Does this work well? Ty, you got to tell the folks about MEP because I think this is Dell's secret sauce, John. This is how we make sure models are going to run great on our stuff out of the box. There's a whole army of people and agents working on it. Yeah. And so we hit the 3 million model mark of public models on Hugging Face within the last month or so. Right. And that's a big milestone. Each of those models are well suited for different things. Right. And so we decided 18 months ago that we needed a more systematic way to do that. Today, that capability for us means that I can tell an agent, hey, there's this new model out there.

17:25Go run it on these 10 platforms and tell me where it fits on the Pareto curve. So that when I go talk to a customer in the financial industry or in a healthcare industry, or should I buy this system or this system, I'm coming in with, well, here's what I see from the data, right? If you're operating point, you need to put 10 users on this thing. Here's what you need, right? You need this model on this platform for this use case. And so what Ish is talking about, that is what we are kind of bringing in across the board. It's not just how smart is the model, it's how well does it fit for your use case.

17:55Jon Krohn:I love it. What's a Pareto curve, Tyler? Pareto curve is looking at kind of what's the best thing to run if I care about a couple different KPIs, right? So what we typically will do for Pareto curve is look at how fast is it, right? versus how intelligent is it. And so if I need a model that has at least this intelligence, then I probably want to pick one that's there and fast, right? If I need something that's at least of particular speed, then I want to pick the smartest model that I can do. And what Ish was mentioning earlier on the artificial analysis site, they have a bunch of burrito curves that are looking at other types of comparisons, right?

18:37It's cost per task for intelligence on different capabilities. That's a really interesting way to look at it. And I think the closer that you can kind of tune in your use case for it, that's a great way to pick when to move, when to upgrade. Right.

18:53Jon Krohn:Fantastic. Really well defined there. All right. We should probably get back on track, although it did look like you might have just inhaled. Did you inhale? I did inhale. Is that President Obama who said that on the campaign trail? No. Oh, yes. Yes. Yeah. Right. You know, you know. And he was talking about breathing in to say something really important on a podcast. Absolutely. That's what he was talking about. Pareto curves and fundamentally what is at the heart of this platform Tyler has built and what is at the heart of every decision that every enterprise is having to make right now are tradeoffs.

19:28Cost versus speed. Speed versus performance. Performance versus cost. Like you can pick your x-axis and your y-axis and plot to your heart's content and that curve that pops up. You're looking for the part where the curve starts to flatten out and it's probably your best bang for your buck, so to speak. I think that it's hard to understate the world of tokenomics and this economics of tokens and this economics of what model you pick, why you pick it, where it runs, what size is the model, how many NVIDIA GPUs are in your PC. Like all of these things are part of this decision matrix of how to get to the end state that your C-suite is demanding.

20:15And Sharish and I talked to you about this last time, John. The hardware is one dimension of that. Over the last 12 months, it's become clear the software is another dimension. Now you've got hardware, you've got software, you've got use case, you've got the literacy of the people using the production stack. All of these things matter. so the reason I inhaled was to say trade-offs are super important and defining what trade-offs you are and are not willing to make as an organization that's going to be the ballgame so you got to stay

20:46Jon Krohn:on top of it love it so many great pieces of information for us to work with practically already in this episode the next one is the long promised tokens and token economics or tokenomics to make a portmanteau. I think portmanteau is the right word. Might have to look that up when you guys are speaking. Oh, I got some headdons. Great. So probably 95 % of our listeners know what tokens are, but we can really quickly define that and then talk about how token consumption differs between the AI of 12 months ago or more that was this kind of generative or conversational only tokenomics relative to the tokenomics of today in 2026, which is so agentic.

21:32Yeah. Ty, I'll leave the what is a token to you and then I'll take the second half of the question.

21:39Jon Krohn:And Tyler, with you speaking about this, I understand that in your lab in Round Rock, Texas at Dell Technologies, you have a big screen showing token usage and somehow you're getting tons of free token usage. That seems to be something prominent on the screen. Yeah, yeah, yeah. So tokens are basically, it's the atomic unit of compute for a AI system, right? You can break up a paragraph into tokens. So you can break up an image to tokens. You can break up video to tokens, audio. All of it is just the mathematical representation of a particular chunk of information, right? And at every unit of compute, every cycle of a model, it's producing the next one, right?

22:27And so what we've got in the lab, we've got a bunch of the systems that we were talking about earlier. And we're going to talk a little bit more about what exactly those are that we're doing in the AI space with our Dell platforms. But we've got a bunch of them hooked up with a variety of the latest and greatest models on there. And we're running workloads on, right? We're doing it for the lab infrastructure for development. We're doing it for tools and analysis and reporting and visualization. We're doing it for customer pilots, right? Of, hey, here's the use case that I need to size for you on our hardware.

23:02And so we've got, it's not a leaderboard, right? We're not token maxing here, right? But we're taking all that work and just visualizing what is the cost deflection from that, right? If I went and ran that on the equivalent cloud frontier model, what is the cost of that. And we're doing hundreds of millions of tokens worth of volume a day inside the lab. And when we say free, what we're really saying is none of that is incremental cost. We're not being billed for any of that. That is a system we purchased one time that we are operating with open software, open models for free in perpetuity. And right now, my run rate, I'll tell you is about$160 ,000 off of a lab of five or 10 people hitting this thing with just normal usage, right?

Read the full transcript

23:50So, and what I'll add to what Tyler just said, heuristically, a token, you can consider like three fourths of an English word, like take an average English word, consider three fourths of it. And tokens include spaces and dashes and commas. And to Tyler's point, like pictures can be converted into tokens. And that conversion is like what allows you to like have a conversation with ChatGPT. Hello, ChatGPT. Good morning. Good morning-ish. Like, hey, I'm going to give you a picture. I need you to take a look at it. Here's the picture. Hit send. What's happening on the back end? Like that picture is getting tokenized.

24:25Like it has to be converted before a model, this black box engine thing that someone has made and trained and tied a bow on and hand it to you, can intake it, process it, figure out how it wants to respond to it, spit those tokens out on the back end. Now, what's really important here is not only is it the atomic unit, like Tyler said, it's how you get billed. For frontier models, a million tokens of output, 50 US dollars per million. Now, to give you an example, over the weekend, I burned about 2 billion tokens working on a side project. Right. And again, there was a brief moment a couple of months ago, John, where like the token maxing news cycle really picked up.

25:11And it was like all these companies had these leaderboards and productivity equals token burn. Right. It's an incredibly crude heuristic to use, but it's what we had. And to a certain extent, it's what we still have. So early in the adoption curve of a company in AI, how many tokens people are using is a good heuristic for our people using your AI tools at all. Right. But the key here is, and this is kind of the drug you get hooked on, it's$50 per million tokens of output. Very, very smart listener base. I don't have to tell them what$2.5 billion would have cost me. So it's important to understand that link because it's the gas.

25:50It's how it gets measured. And the gas is expensive.

25:55Jon Krohn:Who knew? So how do we, well, I guess we'll get to that later in the episode. It seems like, yeah, we have a solution, obviously, involving hardware so that we can be churning through billions of tokens. There's a stat you said, Tyler, that I didn't quite understand if it was dollars or tokens. You're talking about 160 ,000. What were the units? 160 ,000 what per day? 160. Well, our run rate in the lab, right, is about$160 ,000 per year of what we would spend that we are using the devices we got in the lab that cost a heck of a heck of a lot less than that. It's the cost avoidance. I see. I see.

26:36Jon Krohn:And what is like roughly back of the envelope orders of magnitude? How much do you think the equipment costs like 10 grand kind of thing so to be doing that 160 so so the one the one that we're using the most right now and we'll talk about the the lineup here in a minute but it's the dell pro max with gb300 which is basically the biggest baddest thing you can plug into a wall in an office space uh we it's got a blackwell ultra gpu from nvidia it's got 1300 watts of GPU capacity, 252 gigabytes of HBM3E. I can go through the spec sheet on it. This is the naturally aspirated V12 of AI computers is probably the best way to put it.

27:20Except it's liquid cool. Right now on Dell.com it's somewhere over$100 ,000. It's a serious piece of iron. He got me on the natural aspiration, John. You're still saving.

27:36Jon Krohn:So it's like, you know, roughly a hundred thousand dollar piece of hardware, but your team is spending$160 ,000 per year and you could have that hardware for multiple years and it will still be current. So pretty obvious how that is major cost savings. Um, before we get into reeling off tons and tons of stats, which Tyler just did from memory, and I guess it's your job, but it still was impressive. Um, why does token consumption go up so much with Agentic AI? Like Ish, how did you burn through billions of tokens on the weekend on a side project? And can you tell us what it is? I can. You may have to bleep out a word if I commit some sort of IP issue.

28:17So, okay. So it's my one year anniversary this Sunday. And as my wedding gift to my wife last year, what I did was I took, everybody played Pokemon as a kid. Pokemon's making a comeback. It's cool again. Everything old is new again. I basically built a fan game in the art of Pokemon, where the map is my home, my like area that we live in Atlanta, where my wife and I met, where we got engaged, where, and I have these little pixel art maps and I had her caricature done as pixel art. And I replaced the Pokemon with my dogs, right? That's the project. Every year, every major life event that we have, I build a chapter into the game.

29:04And that's my get out of jail free card on the present part of things. And so what I've been working on is these models and their capabilities over the last couple of months have shot through the roof. The artwork has gotten considerably better. The game mechanics and how much I need to supervise my little buddies as they go off and work. I can go have a cup of coffee. And when I come back, the chapter is built. The reason the burn was so high is because what these agents are doing in order to achieve the task, just like humans, they're divvying up the work and they're spawning sub-agents. So now you've got like an agent in charge of a bunch of other agents, right?

29:49And yes, the pie of work is finite, right? Like you have your finite pie of work. But because you've got all these sub-agents in action, like are the sub-agents doing things to like the nth level of token efficiency that a single agent would have done? It's the same thing anthropologically as when you think about like humans in a workplace, right? If one person says, everybody get out of my way, I'm going to own this task single-handedly. I'm going to do it as efficiently as possible, but I'm one person. This is like queuing theory. How much throughput do you have? Multiple sub-agents means that you go faster, means the work gets divvied up, but the pie of work might get a little bit bigger because the sub-agents are at liberty to do certain things.

30:31right? The point of this is best probably articulated by something that has almost nothing to do with what we've talked about, although I'm sure it'll come up. It's this organization called METR, M-E-T-R, Model Evaluation of Research. John, you're nodding,

30:48Jon Krohn:so I'm not sure if they've been on the pod or... I talk about METR probably more than any other single thing on the podcast. And then almost every talk that I've given for a year or two now, Right. Near the beginning, I show meter charts. Ah, so our presentation and your presentation are basically starting the same way. And then yours continue to be smart and mine kind of plateau. Meter, model evaluation threat research, and SDS listeners are going to be familiar with this at this point, has a chart, which when you land on their website, maybe we can put it in the show notes here. like it shows on one dimension time, like 2021 until now.

31:34And then on the other dimension, it shows the ability of a model to operate unsupervised, to achieve a certain goal at a certain fidelity of accuracy compared to a human given the same task. Now, Meter, the reason they have this big, scary name, which says threat research inside of it, their whole point was like, hey, at what point is AI going to cause harm to human beings? And like, we should probably be tracking that. And the heuristic they came up to track that with is like this chart, like how much can it do by itself? And that chart is just like, not only is it up and to the right, it's just like, I mean, it's gone vertical, right?

32:09And at a certain point, they just kind of said, I don't know, it just keeps going up.

32:13Jon Krohn:Since the release of Mythos, they can't really track, like it has gone off of the meter charts because in order to be able to benchmark the performance of models effectively on one of these charts, you have to have had humans doing these tasks and know how long it takes humans to do these tasks as supervised. And that was easy, like 2021, when you're looking at GPT-3 level capability and the tasks are only seconds long or then minutes long with GPT-4 on average, it's very easy to come up with tasks that you can give humans to do. And it's not that expensive to pay them to do it and figure out how long it actually takes them on average to do it.

32:49Jon Krohn:But now that Mythos is doing, or Fable or Astra, GPT-6 from OpenAI, that class of models is now doing dozens of hours of work. Work that would take a human dozens of hours. It could take the AI model 30 minutes or whatever to do something that takes a human 16 hours or 24 hours or 36 hours. We don't know how long those tasks, we don't know how capable these models are because we don't have any human benchmarks. Because how do you even, it's hard to even think of like write a book chapter, write a book. It is quite literally off the charts, quite literally off the charts. And this, they accidentally invented a chart for, for one purpose is now like the best visual we have for like capabilities of models over time.

33:38Right. But as these capabilities go up, like it's Jevin's paradox here, like even if token costs get cheaper, over time, like the base is going to move on you because people are going to realize they could do things like, you know, it took Nintendo how many years to develop a Pokemon game? Like they'd come out every two or three years when we were kids. Now it's like in a weekend, someone can sit down and build a video game to the same level of fidelity. Like the token consumption is growing and it kind of doesn't matter how cheap you make the individual token. If the order of magnitude of usage is just constantly chain reacting on itself to get bigger and bigger and bigger.

34:23So that's kind of where we are.

34:26Jon Krohn:If you're like me, you've spent a lot of this year spinning up AI agents and cloud workloads. Here's something I hadn't thought hard about until recently. Every one of those deployments creates a new identity with permanent access to your critical systems. And the legacy tools most companies rely on were built to manage human employees. So all that machine and AI access goes largely unmanaged. That's the problem IDERA by Palo Alto Networks was built to solve. Human, machine, AI, one identity security platform for all of them. IDERA replaces standing permissions with dynamic access, so you can lock down every identity without slowing your team down.

35:03Jon Krohn:For those of us shipping agents into production, that's a big deal. Secure every identity with IDERA by Palo Alto Networks. visit paloaltonetworks.com slash idira again that's paloaltonetworks.com slash idira yeah and so that is why token use is exploding javon's paradox another thing i'll have in the show notes we'll want to read more into that but it's something we do also talk about on the show for a bit and it sounds like clearly instead of like i do for the most part today instead of buying tokens instead of paying for tokens or running into token thresholds with one of the major cloud providers, we could be using our own hardware instead.

35:52Jon Krohn:That's the other big option. We're going to talk about specific examples, but just kind of generally at a high level first, what is the big, like if organizations move agentic workloads from paying per token to some cloud provider relative to doing it on their own desk side infrastructure? What are the potential kind of cost savings there? So one thing I'll add before I answer the question, John, is that there's a middle that it's really important that we talk about the middle before we even get to desk side, right? That middle is on-prem compute. It's your own big computer as opposed to your own under desk computer.

36:33And obviously Dell is a, it's a very uniquely situated company because we do both, right? Like we have our infrastructure solutions and we have our client solutions. Dell is building the backbone for training and inference for massive companies all over the world, frontier labs and all, and they're doing so with that middle, right? So like this decision of like, Hey, I don't want to pay a cloud service for inference, or I don't want to pay a cloud service for training, or I don't want to pay a cloud service to host my deployed enterprise workload, you then sort of hit a fork in the decision chart, right?

37:09Which is, okay, how big are we talking? How many people are we talking about? How much compute do you need? Oftentimes, the answer is going to be an on-prem server, not an under-desk GB300. But the under-desk GB300 is going to see that class of device and that class of ability closer and closer to the person. I didn't need one of those devices a year ago. I mean, arguably, I don't need one now. But I didn't need one of those a year ago, right? Now that my workflows are actively getting interrupted by token caps, I'm interested, right? So that middle is really, really important to acknowledge because, and Dell has papers on this that we've published which articulate literally the answer to your question, John, which is here's what it would cost in the cloud, Here's what it costs on-prem on your own server.

37:59And here's what it costs on a T6 tower, which is one of our best AI devices, which you can cram full of NVIDIA GPUs and you can host a little server, right? Like you can put some models on this thing that have some serious capabilities. So the cost savings, I'm going to give you my recovering consultant answer on this. It depends, right? I know, I know, I know. My BCG bosses would be so proud of me. It depends. And it highly is contingent on are you Tyler's team of five software engineers who are, you know, they're driving the H2 Hummer. Like it's a gas guzzling pedal to the metal. How much code can I write?

38:43How many of these problems that I've been wrestling with for years can I try to solve quickly? they're going to experience the cost-saving curve a lot faster than a more casual user or more casual work case or workload. This is why you're seeing enterprises adopt AI for software engineering faster than arguably any other function within the company.

39:04Jon Krohn:Nice. And after all that, which was very interesting, and thank you for the tour of the middle ground. I mean, yeah, I know that you have the it depends answer on the kind of cost savings thing, But I do also know that the research groups that you work with at Dell, like Signal 65, Solution Brief, I know that you do have some rough figures that you can give us. Yeah. So we worked with Signal 65 team across the Dell Technologies portfolio, right? We took devices like our Dell Pro Max, the GB10, right? So compact form factor, workstations. Tyler just held one up on the screen for those of you who aren't watching on YouTube.

39:44Jon Krohn:It looks like... this is one of those where ai has gotten so powerful that you can unlock some really uh really nice form factors we also have scale up from there right we did a t2 we did the gb300 we did gpu servers in here too right and so uh from that work um that it depends the answer we've seen as high as up to 87 percent cost savings versus uh the equivalent workloads running in cloud, right? Those are all studies of, am I modeling for knowledge workers? Am I modeling for sales workers? Am I acknowledging for coding workers? How many agents am I running? What is, how many times are they using it per day?

40:25What's the volume of work that they're doing there, right? And what that really translates to, and this is a really, really different way to think about PC buying, right? Is that you're not looking at the CapEx, how much does the system cost? You're looking at how much does this system save me or make me, right? And so on some of these systems, right, for some of the workloads, we've seen they'll pay for themselves in two months versus running that same workload up in cloud. And over the lifetime of that system, you'll get over a million dollars worth of equivalent tokens spent, right? So I went and did the homework while Tyler was covering my rear, John.

41:07Jon Krohn:Which of your agents did the homework, Ish? I can't disclose that. I need to be on commission to disclose that. No. Okay. So we're talking a high complexity workload, AI agent, software assistant, software development assistant, right? A T2 workstation running an RTX Pro 6000 Blackwell achieved 93 % cost savings for deployments supporting approximately 20 agents. So you're talking about like orders of magnitude of potential savings, both in that middle layer, right? Like if you invest in an AI factory and you have these massive use cases, and for the under desk layer, which is if you buy a T2 tower, which is much more accessible entry point into local AI and running your own inference than a T6 or GB300.

41:55Progressively, those get more expensive as you move up the stack. But what used to prevent people from realizing that 93 % cost savings, and this kind of moves into this concept we've been playing with around desk-side authentic AI, it used to be really hard. It used to be complicated. And it used to be like Tyler and Ish over a weekend would spend hours getting set up on it and getting it all tuned and pecked out right so it would work so I could use it from my phone so I could do all kinds. Over the last 12 months, the strides in software, the strides in ease of setup, the strides like the ecosystem has come together to make it such that while it's not the same level of point and click as like opening a website on your phone and just starting to talk to a model, the savings of up to 93 % are certainly worth now the amount of effort it would take to set up a system this way.

42:52So you can hit it when you need it.

42:53Jon Krohn:So it can run the model you need. And I think some people might worry about not having the capabilities they need, but there are probably, there's not that many use cases where you need a Frontier, Fable, or GPT-6 Astra capability, especially when there are open weight models, Kimi series, QN series that you could be using and getting so close to the frontier? GLM is another one by ZAI. A couple of weeks ago, perhaps a month ago now, many, many, many organizations signed on to letters supporting open models, right? The Lama series of models for Meta back when all of this was kind of getting started, you know, it was the articulation of like, hey, we need open models because we need people to have choice and we need things that people can fine tune.

43:52And we need like the model layer of control in Meta's early opinion of all this was like, hey, we need people to have options and we're going to build the best option, right, for people to use. Since then, many, many companies have entered the fray. Inkling is one of my favorite series of models right now. And it's actually Miramarati, who was at OpenAI and now I think it's Thinking Machines. Thinking Machines. Yeah, Thinking Machines is her company. They intentionally did not do a model to compete at the front. And you'll hear this term a lot, and I know John's heard this term, and I know Tyler's heard this term, the jagged frontier of AI, right?

44:35It's not a clean frontier. It's a jagged frontier, which means that for different tasks, and this goes back to the trade-off discussion we were having, different models are going to be right or good enough, right? Like in order to do the artwork for my game, I have found that yes, I need frontier model level image generation capability. Otherwise it doesn't look how I want it. But the code underlying my game engine, I don't need frontier for that. So I will have those agents running on a babier model with a lesser level of thinking and that'll save me some token bird. But right now I'm the human and I'm routing all that.

45:16Pretty soon, you're not going to have to do that either.

45:18Jon Krohn:Yeah. This jagged frontier thing is critical to mention because it also, even those crazy meter charts that we were talking about, that is specific to areas where we have a lot of training data. And the places that we have a lot of training data are things like mathematics problems, computer science, machine learning, where we can simulate tons of data and know that it's accurate because the math works or the code runs or the machine learning model works. And so that is where the frontier is sharpest or furthest ahead. Whereas, yeah, the jaggedness, you know, if you try to, good job, good luck getting an AI model to like clean bed sores off a hospital patient.

46:01Yeah, right. Although physical AI, man, I think 12 months from now, if you have a SPAC, John, we're going to be having a conversation that may not be that far away from that. So never say never.

46:13Jon Krohn:Yeah, I'd love to see it. But anyway, back to, I'm going to try to keep us on track a bit more so we get through everything we wanted to cover in this episode. Um, it sounds like with these, with these options of having our own hardware, whether it's desk side or on-prem, it sounds like we can, uh, get a return very quickly. It looks like I'm kind of, I'm, I'm leading the witness here, or I'm, I'm actually, I'm just going to, I know that you can get break even as quickly as three months after you buy that hardware. Um, I don't know if you want to tell me more about that stat. If you use Tyler's lab as the example, earlier on, we had fewer engineers running even more on it, and then we were exploring concurrency.

47:00Our early math on the first GB300 that we put in Tyler's lab is that it broke even in three months. like we saw it so that stat and that stat they tie out for me at least because if you also know that you're not paying marginal token costs right and empirically you're achieving the objectives that you sought out to achieve and you have evidence that like hey i'm not using the the tip of the spear frontier model that costs 50 per million tokens of output i'm using deep seek or I'm using Quen or I'm using GLM or I'm using one of these Nemotron or Poolside or Inkling, all of these folks who make these models intended to run on smaller hardware than a full-blown data center.

47:48If you do that math, you are very quickly going to come to very short break-even periods, but you're inclined to use it more because it's empirically solving for your need. So you will realize very quickly, I don't need the bleeding edge to do this. I can do this this way. Therefore, your usage will go up so that breakeven time will get pulled in.

48:11Jon Krohn:If an AI agent caused an incident at your company tomorrow, what evidence could you produce? Most teams have the prompt and the final answer. Everything in between, the commands it ran, the files it changed, the credentials it picked up, are gone the moment the terminal closes. Origin closes that gap. A sensor on the endpoint records the agent's work as a trace. Who started the session? What was asked? What the agent reached and what changed on one timeline? And because the sensor sits on the machine, you don't rely on the agent's own account of itself. Origin lives on the endpoint because that's where the work happens.

48:46Jon Krohn:Coding agents in a terminal, local agents, agents calling MCP servers on a laptop, none of that passes through a cloud gateway and Origin sees it anyway. When something unexpected happens, you read the session in order from prompt to outcome. Origin is endpoint AI observability. See what a trace looks like for yourself at originhq.com slash SDS. Cool. Really cool. It is much faster than I would have anticipated. And so if you guys have, like, I think we're going to start moving, I'm going to start moving you to solutions, you know, your specific solutions. I think we want to talk specifically in this episode about desk side agentic AI solutions from Dell to all the problems we've been talking about in this episode.

49:31Jon Krohn:But do you have kind of one big takeaway for me on the tokenomics conversation that has enriched our conversation so far? So I think the really interesting thing that is definitely a challenge to how we think about IT is that we have this lived inherent assumption that the day that I put a device on the user's desk is its best day of life. Right. There are lots of things that happen after that. Right. You have policy updates. You have OS updates. You have the users doing crazy things on the systems. and what we've seen is with AI and we've been talking about the trend of models getting better, that carries through for the platforms you're buying, right?

50:14The systems that we're talking about here, they can do way more things than they could a year ago. They'll be able to do way more in another year, right? So just as the frontier is increasing its capabilities on different models, even for the same size of hardware, because of what the industry is doing right now, you're going to be able to achieve harder and harder problems over time as well.

50:41Jon Krohn:Really cool. All right, let's move on to the solution part of the episode, which is, yeah, Dell, Deskside, Agentec AI. We already talked a bit about on-prem and we already kind of got an introduction to these Deskside solutions that y 'all offer at Dell. So tell me what is included in one of these packages. I think it's more than just being a workstation, right? Yeah. So with our desk-side agentic AI, what we're really saying is, hey, we have these set of platforms, this portion of our high-end AI portfolio that are agent-ready, right? We know, we can tell you, they run powerful enough models.

51:23They'll do them at scale. They'll do them efficiently. You buy one of these platforms, right? Everything from the GB10 to the Dell Pro Precision 9 series of scalable workstation towers, a T2, a T4, a T6, one to five GPUs, or the Dell Pro Max, the GB300 we were talking about earlier, right? You buy one of those systems, you go put the NVIDIA agent toolkit, you put NVIDIA NEMA Claw, go run OpenClaw or Hermes agent or your agent harness of choice on top of that, right? There's lots of very easy ways to get to value. And so with these solutions, we have a partner ecosystem in there. We bring in security tooling for it.

52:02We bring in management tooling for it. We really take it from, I can do this thing as an exploration, right? And to move it to, I can do this in my business, right? And that's where we've seen in a lot of customer conversations, that's where we're trying to help, right? Is how do I take this out of my lab and get it into my workflows broader than that, right? The other piece of the Deskite Agentic AI offering is we have this professional services team that will come in and help you get started. We have an adoption services to help you get started with local AI. Lots of people are using cloud and frontier models right now because it's easy.

52:43and what we're trying to do with Dustside, agentic AI is make it so that it's easy to, it's as easy to do the work with the value realization we've been talking around with the tokenomics piece. Yeah, tokenomics is this big theoretical thing where I can find the point on the curve that is best for me and my business, right? Whether I'm a small business that has a couple of retail locations and I need inferencing happening at those locations, whatever the case might be, all the way up to, you know, I'm McDonald's and I have many retail locations and I need all of those things to work together. Like if a lot of what Tyler said out loud just now sounded complicated, that's what desk side agentic AI services from Dell and NVIDIA, that like, that's, that's what we're trying to solve for.

53:30We are trying to make this as easy for people as what they're used to on the consumer software side, which is it just works. And we're solving that problem by forward deploying folks like Tyler. intent, to come help you out. And like, you know, we're not just going to ship you this box and say, ta-da, here's this Dell PC. Like, boy, do I have a solution for you. Like the box comes with the tiler. And that is, I think, a big part of the services offer, a big part of what turns this into, we're going to sell you a PC. No, no, we're going to help you achieve a business outcome. And we're going to leave you with the keys to a car that runs.

54:09And it's not a DIY, assemble it yourself unless you want that in which in which case we're happy to provide that too nice yeah so

54:19Jon Krohn:the the offering here with i guess this is the del ai factory with nvidia this kind of like end-to-end offering of hardware software so things like the nvidia nemo claw stack that we're talking about we'll get more into software again in a second and software options people have there it includes security and it includes services like having a tyler although i kind of say i don't know does tyler know that much he hasn't impressed me that much i know underachiever um all right so let's talk about we're going to talk about hardware specifically and then software specifically after that so uh what are the three specific different tiers of hardware that are available in this desk side agentic solution yeah so we have we have a kind of an exploration tier right it's where i i may have a power user that i just want to get out of my my inbox asking for more tokens, right?

55:11That I want to go put a GB10 or a T2 and say, hey, go nuts, right? Or I may have a team or a lab or a site where I just need dedicated intelligence at some small scale. We've got multi-GPU towers that you can go in there, deploy up to a 500 billion parameter model there and get to a better tier of intelligence. And then we have larger scale-out solutions, right? with the GB300, with the T6, and in multiples, where you're really looking at up to a trillion parameter models. You might be looking at hundreds of different agent instances. You might be doing things like the self-improving agents or self-optimizing problem sets that we've seen some buzz around the industry where there's hundreds or thousands or tens of thousands of agents working on one problem together and experimenting to try to find what the right answer is, right?

56:09And so I think there's a lot of different problems that map well into those different pieces. And one of the reasons when we talk about where am I going for data center, where am I going for the edge, a lot of the customer use cases that are driving more towards these edge deployments with the desk side systems, it's because they want to bring intelligence into where the data is, right? Because it's IP, because it's sensitive data, because it's data that has legal agreements, governing where exactly that can be moved around to, where it can be processed, what types of tools and systems. It's a really wide variety of reasons of why you would use this, but it's a flexible operating model with a level of capability that we've never had before.

56:55Jon Krohn:Nice. And let's now move from hardware on to software. So we talked earlier about NVIDIA's NemoClaw. What does that include? Yeah, so NVIDIA Nemoclaw has a couple of major pieces. So NVIDIA, with a lot of the AI ecosystem software that they're promoting, they're doing some great open source contributions, right? For me, one of the key pieces of Nemoclaw is OpenShell, which is a guardrail layer that wraps around your agent harness. You can put it into an OpenShell sandbox. You have fine-grained permissions over the policies to really dictate what that agent can and can't do, right? You can read from these websites, from these data sets.

57:41You can use these tools. You can use these tools to access these sites. You can get and put and post and patch or not for all these different pieces. And so with NEMA Claw, they're really making it easy to deploy that consistently across different environments. That manageability part, John, is super important. A lot of this stuff has been increasingly possible over time. We mentioned the difficulty part of it. Yes, there's a technical difficulty component to this problem. There's also a manageability and security component to this problem. manageability, security, and all of the observability that that entails, which like for any of your listeners who are IT folks, like this is going to be old hat to them.

58:26It's the classic problems of who's on my network, who's on my devices, what are they doing? How are they doing it? How do I track spend? How do I track if what they're claiming to be doing is actually what they're doing? All of that now are dimensions of a new order, right? Like the AI question within the enterprise. So what NemoClaw allows an enterprise to do is bring some of those traditional IT guardrails into the AI context where you can have more control at a granular level over what is and is not happening within your IT environment.

59:00Jon Krohn:Love it. And so now we have a good understanding of what these Dell deskside agentic AI solutions have in terms of hardware, in terms of software, let's, if you can, get into some specific use cases that illustrate for our listeners how that hardware, that software can work together to provide, you know, token efficient, highly accurate solutions. Yeah. So I think the number one use case that is common across most of our customer engagements is how do I get my developers to stop spending millions of dollars per month, right? How do I put caps on that? I like the productivity. I like what I'm getting out of it, but the line keeps going up.

59:47And so a lot of them will come in with, well, how do I get good enough models at a capitalized operating model so that I can move more of that volume down? But past that, there's a lot of really interesting places. We have healthcare researchers who are looking at how do I go off and scale out so that I can do these tests and tasks and look at different papers or pull in different research information because there's so much going on right now. How do I use an agent to give me extra hands? Right. We also have with with where we're at in the in the year and in the cycle, we have a lot of academic institutions coming in and saying, hey, how do I how do I train my students, whether they're computer scientists or going into data science and data engineering or not?

1:00:40How do I train them to be able to take advantage of these new capabilities that are going out there? And so we have a bunch of these lab deployments where it's, hey, I need 10 seats or 30 seats of these labs. How do I give them access to these great capabilities in a way that's friendly to my academic learning environment? Let them go play with different things, pull different tools, try different harnesses and models and all these capabilities. And so we've seen a lot of engagement around our desk-sided genetic AI systems for those types of use cases as well. And it's not just big industrial use cases, John, right?

1:01:18Like these are of all sizes, shapes, scopes. Something that's really important to understand in the context of the question you asked earlier, what's with the token explosion? Like, why is this happening? Why is this happening the way it's happening? like I'm going to steal a term a customer in in London recently used with us citizen development it's such a nice way to say vibe code but citizen development in the enterprise right so like the everyone whether they know it or not is a software engineer now everybody whether they know it or not you're writing code and code equals tokens right like it's very simple a to b to c here So as far as use cases go, even people who don't realize that what they're asking for is a desk side software assistant.

1:02:10That's what they're asking for, because the types of things they're describing are achieved through code. So it's a similar harness. It's a similar setup that we would come in and do for you. Even if the use case you're describing, like the qualitative words you're using to describe it might be something else.

1:02:27Jon Krohn:regular listeners will already be aware that i'm obsessed with anthropic's fable 5 model and it has taken over my working life i'm writing a technical book that includes latex files mathematical notation python code examples and fable 5 and cloud code handles requests i make across whole chapters with accompanying jupiter notebooks end-to-end work that a few short months ago would have been dozens of separate requests with way more manual fiddling required with fable 5 it just works, essentially like magic, first time. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you.

1:03:04Jon Krohn:Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai slash superdata. That's claude.ai slash superdata. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash superdata. I love it. And this works, as you said, across different kinds of scales. Tyler was talking a few minutes ago about kind of this explore tier where you've got Dell Pro Max with GP10 or Dell Pro Precision T2 tower.

1:03:43Jon Krohn:And so that handles up to like eight agents running concurrently. You've got this orchestrate tier where you've got 40 agents. That's more like a T4 tower. supporting 120 to half a trillion parameter models. And then obviously you've got the big heavy hitters, your GB300s, your Dell Pro Precision T6 tower, where you're talking about trillion parameter models, 150 agents running. And it sounds, or from research that we did prior to recording this episode, Signal 65, we mentioned them already earlier. They're a third party that does research for Dell. They provided some pretty staggering stats That's where if you talk about common kind of workload types that you would be doing in this agentic era that we now find ourselves in, someone who's like a knowledge worker doing low complexity tasks like email writing, text summarization, that kind of thing.

1:04:34Jon Krohn:You can see by having a desk side solution like Dell offers as opposed to paying per token with a proprietary API, you'd be seeing savings of like 28 % to 71%, that kind of thing. A sales agent. So we're talking like, you know, increasing complexity here where you have, you know, a mix of that kind of email writing that the knowledge worker was doing, but also maybe research tasks happening, you know, going out and finding out information about a prospective client or lead, gathering sales contacts, that kind of thing. So if you're talking tens of millions of tokens per agent per day, you could be in that kind of medium complexity scenario saving 76 % to 91 % versus cloud solutions.

1:05:20Jon Krohn:And then for a lot of our listeners doing software development, using agents to be increasingly, you know, looking more like that meter chart, using agents for handling tasks that take hours or days for a human to do. We're now talking about like tens of millions of tokens per agent per day being consumed, many potentially dozens of agents, sub agents running. And it's that as you move more, I mean, you're seeing savings even with a low complexity like knowledge work, kind of agentic work. But once we're getting to this high complexity code generation, code review, bug troubleshooting, testing, it starts to become a no-brainer where you're seeing 90 % savings relative to using a cloud solution.

1:06:06Yeah, and I would say the direction, and remember the same thing, right? We try to design workloads and benchmarks that are representative of what we think the complexity of a task looks like. And Signal 6.5 is independent in the way that they come up with things. So I think that the direction of the number is the most important thing. And the magnitude of the number is the most important thing. Like whether you're actually getting 25 % savings, 75 % savings, 95%, again, it depends. I think what's really important is that there is an opportunity here. it is a lot simpler than it used to be. And now with Dell, DeskCyte, Agentic, AI with NVIDIA, it is push button to get to a point where you can use this stuff in a production environment.

1:06:52And I think that's really the big takeaway. And as much as I would love to think everybody knows what our desktop lineup acronyms are, I know that that's not the case, but don't let that drive you off. There's a lot of people at Dell who can help you answer the question of, hey, what device needs to be part of this package, right? and we'll make sure you get the right one.

1:07:11Jon Krohn:All right, so all of that sounds great. Between all of us, we've given now a good run-through of the tokenomics problem, the potential solution, the compelling solution in a lot of cases for having a desk-side or on-prem solution to be saving money and getting the same kinds of results faster in a lot of cases. but if we have listeners out there wondering for them as individuals or particularly for themselves for organizations that they might work at, how hard is it to get up and running with the kinds of solutions that we've discussed today? I would say it's not hard anymore and I would say that that may not even have been true six months ago.

1:07:54So I never want to undermine that this is complicated and we're trying to reduce people's most difficult business problems into reusable kind of work blocks. But I would say that we are going to make this process as simple for you as possible. We're going to be partners in strategizing, we're going to be partners in procuring, we're going to be partners in deployment, we're going to be partners in sustain over the life cycle of that solution. So I would say that if this is something you are faced with, and this being exploding token bills within your enterprise, and if you do have an inkling, if you do have an inkling of this being something that might help you,

1:08:36Jon Krohn:you should give us a call because I think that that intuition is probably right. Love it. Great takeaway there from Ish. And we actually, while you were giving that response, we lost Tyler because he had a hard stop and I haven't been moderating this podcast recording session diligently enough. And so, yeah, we ran out of time for him. In truth, in truth, Tyler got tired of listening to me talk. So he left. So don't worry about that. That happens all the time, John. But so that means we're going to have to get a Tyler book recommendation in a future Tyler Cox appearance on the podcast. Ish, what do you have for us?

1:09:15Jon Krohn:I think last time I cheated and I gave you two. And I think I'm going to do the same thing again this time. Oh, man, you're going to run out someday at this rate. I know. I know. I just I yeah, I there's just too much to consider. right now. It's truly, truly difficult. And I can't keep up with the SDS podcast cadence at this point either. So, you know, books, podcasts, I got a lot of them. Two. So the first one is topical. It's Genius Makers, which I'm sure someone on the pod must have mentioned, or you have read, John, at some point, which is kind of the origin story of the current lapse. Like where were all these folks distributed across Silicon Valley?

1:09:54What companies incubated them? When did they choose to leave and start their own thing? How did they get folded back in sometimes? And it's got names of folks that at this point are canon for anyone who is interested in this industry. So Genius Makers was a great walk through the personalities, the humans who are responsible for the AI. And I think that that is going to be so important over the next 12, 24, 36 months as we get into discussions about ethics and safety and things that we've never had to, we always thought about them, we never had to think about them. And now it's here. My fun book is I've continued to progress through the Jack Reacher series since the last time we talked.

1:10:36And I'm almost done. I'm on book number 29. And I think book number 30 is coming out soon. In Too Deep is the last one I read and Exit Strategy is the next one. So Jack Reacher, nice fun airplane read always a good time uh those are my two love it great recommendations

1:10:52Jon Krohn:ish thank you so much for those and uh yeah thanks to you and wayward tyler uh for taking so much time with us today having such an information-packed episode as we knew it would be and uh in the end i think we have circumvented all of the danger and we've survived to the end of the episode remarkably. For people who want to risk it, risk more danger in the future, how should they follow you, Tyler, Del, about what's going on? Yeah, right at the frontier. So we've got all of our corporate social media handles and those are, if you punch in Del, you punch up NVIDIA, they'll come up. I'm on X as at Ishan Shaw.

1:11:35It's the full name, not the Ishiness, but maybe I'll get that handle one day. And we talk about these things and we We talk about fun other stuff like Omar Chi and fun other stuff that are frontier nerd kind of deals. So we welcome that engagement from everybody and looking forward to continuing the conversation. And thank you, John, for having us back as always.

1:11:56Jon Krohn:Of course. I can't wait till the next time. It is always such a joy to have you on the show. Such invaluable minds to be able to get time with. I'm really lucky. I couldn't keep it. I'm sorry. I could. No, it's true. I know. I know. Yeah, it's yeah. Always a fun time. Such a such a joy to have on the air. And I hope it won't be long till the next time. Same here. Thanks, John. What a fun episode. In it is Shaw and Tyler Cox covered why token consumption is exploding. Agents spawn sub agents that grow the total pie of work. And thanks to Javon's paradox, cheaper tokens lead people to take on ever more ambitious projects.

1:12:37Jon Krohn:We talked about how Tyler's small app pushes hundreds of millions of tokens a day through local hardware, avoiding about$160 ,000 a year in cloud spend, with their first GB300 breaking even in three months. And how the jagged frontier means most tasks don't need a frontier model, so open-way models like Quen, Kimi, and GLM running locally are often good enough. Quick disclaimer here for you as well. The results and claims discussed in this podcast were based on findings documented in the associated white paper and validated through AI lab testing and reference solution evaluations. Outcomes may vary based on workload requirements, deployment environment, system configuration, and implementation approach.

1:13:17Jon Krohn:The reported payback period of up to three months is derived from the testing and analysis presented in the white paper. For additional details, methodology, assumptions, and findings, please refer to the full white paper, which is called The Economics of Agentic AI, Cost Advantages of Dell AI Factory with NVIDIA, and you can find that, of course, in the show notes. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Ish and Tyler's social media profiles, as well as my own at superdatascience.com slash 1031.

1:13:51Jon Krohn:Thanks, of course, to everyone on the Super Data Science podcast team, our podcast manager, Sonia Brejevich, media editor, Mario Pombo, partnerships manager, Natalie Zajski, our researcher, Serge Massis, and our founder, Kirill Aromengo. Thanks to all of them for producing another excellent episode for us today for enabling that super team to create this free podcast for you. We are deeply grateful to our sponsors. You, yes, you can support this show by heading to the show notes and clicking on sponsors links, checking out what they're offering. And if you yourself are interested in sponsoring an episode, you can get the details on how at johnkrone.com slash podcast.

1:14:31Jon Krohn:Otherwise, please help us out by sharing this episode with other folks who would learn who would love to learn about local AI hardware, review this show on your favorite podcasting app or on YouTube, subscribe, obviously, if you're not already a subscriber. But most importantly, I hope you'll just keep on tuning in. I'm so grateful to have you listening. And I hope I can continue to make episodes you love for years and years to come till next time. Keep on rocking it out there. And I'm looking forward to enjoying another round of the super data science podcast with you very soon.

From the publisher

In Episode #1031, Ish Shah and Tyler Cox (Distinguished Engineers in the Office of the CTO for Dell Technologies' client group) join Jon Krohn to work out why agentic AI bills are exploding and what can be done about it. Over one weekend Ish burned roughly two billion tokens on a side project, and that is the ordinary shape of agentic work now: agents spawn sub-agents, the pie of work grows, and cheaper tokens only invite more ambitious projects. Tyler runs a small Dell lab that pushes hundreds of millions of tokens a day through local hardware instead.

In this episode, they define what makes a system agentic, explain how to read a Pareto curve when choosing models, work through the jagged frontier and why most tasks do not need a frontier model, and lay out what moving agentic workloads onto your own hardware does to the economics.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1031⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(00:03:42) What makes a system agentic

(00:12:45) Picking the right model for the task

(00:16:46) How to read a Pareto curve

(00:27:02) Why agents burn so many more tokens

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
1031: Tokenomics: Why Your Agentic AI Bill Is Exploding (and How to Fix It), with Tyler Cox and Ish ShahSuper Data Science: ML & AI Podcast with Jon Krohn · 1 h 15 min
Listen in VO