Why People Are Paying 10x More for AI - and What That Means for the Chip Market | Sid Sheth, d-Matrix

6 Aug 2026 · 51 min · 18 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Sid Sheth (d-Matrix) discusses how AI is driving a “premium token economy” and why some companies pay 10x more for low-latency inference. He argues organizations will increasingly use agentic systems (agents that coordinate tasks) and that interactive, real-time applications command higher token prices than batch-style inference. He claims inference is not one-size-fits-all: general GPUs are broad but inefficient for latency-sensitive workloads, while d-Matrix targets “low latency compute” with architectures that pair compute with much faster memory bandwidth than HBM. He cites Cloud Code as an example where faster responses keep non-programmers engaged, and OpenClaw/OpenCore as examples of agent-to-agent, machine-to-machine workflows that are latency-sensitive. d-Matrix is scaling deployment speed via “RackScale” (rack-level deployment) and an acquisition of Giga.io for faster server/rack deployment.

Guests

Sid Sheth, founder/leader at d-Matrix; no other guests mentioned.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Inference Market

0:45 to 6:12

Gain insights into the evolving inference space and its implications for AI.

“I mean, I read that you've had some acquisitions since the last time we spoke.”

The Rise of RackScale Deployment

6:12 to 12:00

Discover the significance of rack-scale deployment in modern data centers.

“Manufacturing, just delivering on product, deploying product faster, supporting product faster, making more of our product.”

Agentic Communication and Its Impact

12:00 to 14:00

Understand the implications of agentic communication in AI tools.

“Not everyone because I would say, depending on your business model, right?”

Exploring Agentic Capabilities in AI

14:00 to 17:28

Learn about the advancements in agentic communication and its implications for AI workflows.

“Clawed is kind of incorporating a lot of agentic capabilities.”

Future of Organizational AI

17:29 to 20:29

Discover the potential for AI to influence organizational tasks and strategies at a high level.

“I wish I could sit here and say, I saw it all.”

Implementing AI at D-Matrix

20:30 to 23:52

Understand how D-Matrix is mandating AI usage across all teams for improved efficiency.

“people talk about organizations being entirely run by agents, right?”

Using AI for Strategic Decisions

23:53 to 27:52

Explore how AI tools are transforming strategic decision-making and integration processes.

“There's one, Aerotechnologies, which builds essentially a C-suite co-pilot.”

AI as an Augmentation Tool

28:01 to 29:13

Learn how AI can enhance presentations and discussions by acting as a supportive tool.

“It didn't change it was augmented like look you can present this maybe talk about this this way This is another way you can say the same thing.”

Low Latency Computing Needs

29:14 to 30:28

Explore the demand for low latency computing in AI and its implications for industry players.

“who are the customers that are buying the chips for that low latency?”

The Shift from Cloud to On-Prem Solutions

30:29 to 31:53

Discover the trend of companies moving back from the cloud to on-premises data centers.

“Now, a lot of people have gone through that.”
Show all 18 chapters

Hybrid Approaches in Data Management

31:54 to 34:08

Understand how hybrid data management strategies are evolving with AI integration.

“Companies that are building these new age applications for AI, they're all running in the cloud.”

Regulatory Concerns in AI Data Usage

34:09 to 36:23

Learn about ongoing concerns regarding data privacy and regulatory issues in AI.

“Although in regulated industries, there is concern about sending data to the cloud.”

Focusing on the U.S. AI Market

36:24 to 37:35

Explore the importance of the U.S. market in driving AI innovation and deployment.

“Yeah, I travel a lot, and, you know, every country.”

Challenges in Chip Supply and Manufacturing

37:36 to 39:48

Discuss the current challenges in chip supply and the role of manufacturers like TSMC.

“Once those questions have been answered, then it'll diffuse into a lot of the sovereign applications, into a lot of the enterprise applications.”

AI Applications and Market Dynamics

39:49 to 42:00

Examine the dynamics of the AI market and the companies involved in developing AI solutions.

“So we have memory, you know, chip shortage.”

The Future of Agentic AI

42:00 to 45:30

Explore the implications and potential growth of agentic AI companies.

“I think if there's a very relevant, large problem that a very relevant, large customer is trying to solve, and we think we are helping solve that type of problem, then the supply chain will make a place for you, right?”

AI's Impact on Employment and Industry

45:30 to 48:24

Discussion on how AI will reshape industries and the workforce.

“that's basically an agentic first company.”

AI in Healthcare: Challenges and Opportunities

48:24 to 50:39

Insights into the role of AI in healthcare and drug discovery.

“Do you see much demand for agentic AI in the healthcare space?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00People talk about organizations being entirely run by agents, right? How far are we from that? Eventually, I could run large portions of most companies. Do you see emerging companies, startups that are agent first? The New York Times had a piece recently about a guy and his brother running a$1.6 billion sales company to basically an agent first company. Do you see that segment of the economy growing? It's really an opportunity for everyone to go into something new. I think at the D-Matrix, we are being very proactive about the use of AI across the whole organization for various different use cases.

0:34We have an AI team in the company that is entirely focused on coding AI across the world. So it is not an option. It is mandated. Tell us what you're talking about at HumanX. I mean, I read that you've had some acquisitions since the last time we spoke. We're moving into a larger system, not just providing chips. And I'm interested in that. You know, my interest is generally in algorithmic research, but I do try and follow hardware. You guys are a player in that space. I think I told you last time I talked periodically to Andrew Feldman at Cerebrus, to Rodrigo Liang at Semba Nova. I haven't spoken to Brock.

1:25If I'm not mistaken, you and those three are the primary players in this new inference-ship space, or maybe not. I mean, if you could give me an overview on what you guys are doing and how you're differentiating. The inferencing space specifically, right, is kind of bifurcating, right? Because I think everyone looked at inference as one, you know, kind of, you know, singular entity, right? But I think there is lots of nuance and when it comes to inference, it's not a one size fits all. Yeah. Something we've been saying for a long time. so you know depending on where you are doing the influencing what your target markets are what applications you're going after you can't really build a single chip that serves all of influencing's needs right and so you got to really kind of break it down and not every company is playing in every pocket of the market.

2:32It can't, right? Our GPUs, you know, just because they are so general, the ecosystem is so broad. Clearly, GPUs can go into many different applications for inference, right? But they're not going to be very efficient because of that, right? So yes, they're very broad in their applicability, their general purpose. But when it comes to certain breakaway applications where applications need a very specific metric fully optimized, then GPUs will not do well in that market. And I think, so I think the way to look at it is, okay, you have the influencing market. It's going to be the largest part of AI compute.

3:13It's expected to be over a trillion dollar market, you know, call it in the next five years. But it's, you know, then GPUs will play in that market clearly, but there's going to be portions or segments of that trillion dollars that is, we will need, you know, lots of optimization and highly optimized silicon. And the one that is truly breaking out right now is the segment that is called low latency compute. And low latency compute, low latency inference is specifically around applications that need high levels of interactivity. So, you know, a great example would be Cloud Code. This happened, you know, a few months ago.

3:53And you have people who are not programmers, who are now using Cloud Code and are great programmers, right? I mean, the joke is, right, the new programming language is English, right? And everyone can program and you can be an ace programmer and know how to prompt the model and how to use it. So, but that has led to people wanting to interact. And the faster the model is when it responds back to the user, the more likely the user is to stay with the application, right? So I think interactivity has become very important and companies are beginning to charge more for those high levels of interactivity.

4:34So you can always say, look, I don't want that level of interactivity, in which case, you know, I'll pay, call it$2 for a million tokens. But if I want that high level of interactivity, I'll pay$20 for a million tokens. And people are willing to pay that$20. They want that interactivity. So I think given that this new tier of inferencing is emerging, this new tier of tokens is emerging, right, where we call it a premium token economy. So we'll look at the entire token economy, which is inference. This is the premium token economy. Not many companies play in that segment because you need a specific compute architecture that is well suited for the fast token generation, right?

5:15So it's about generating tokens very quickly and that typically means you have an architecture that puts compute and memory together in a way where you have an order of magnitude more memory bandwidth available to your compute than a GPU would have. Because GPUs are based on HBM and they have a certain cap on the amount of memory bandwidth. Solutions like B-Matrix, and you mentioned Grok and Cerebris are the other, have architectures that allow compute to get access to very fast memory. And it's an order of magnitude faster than HBM. So that is the breakaway category right now. And Dmatrix falls into that breakaway category.

6:03And because that category has exploded on us literally in the last few months, we're just trying to keep up. And I think we need to scale up into that opportunity rapidly. Scale up, meaning manufacturing? Manufacturing, just delivering on product, deploying product faster, supporting product faster, making more of our product. All of these need to happen, bringing on more people to help support more customers. All of this needs to happen simultaneously. And that is why we did this acquisition of a company called Giga.io, which allows us to scale into faster deployment of servers, faster deployment of racks.

6:45We actually had started our RackScale journey last year. We announced the protocol squad rack in October of last year at OCP. You can explain what RackScale means. Yeah, sure. So RackScale is, you know, you have, if you go into a data center, You have these aisles or rows of hardware, and the unit of compute deployment in that data center is typically a rack. And racks are then put together into what they call a cluster. And the clusters are kind of put together into, you know, eventually multiple clusters will form a data center. they have you know sub you can actually break it down into pods I mean different people use different terms right so there's racks and pods and clusters and eventually you kind of scale up to the data center right so the unit of deployment in a data center is a rack and these days I think just given how quickly AI computing is you know the demand is shooting you know through the roof so I think you need lots of compute and the unit Everyone is looking of compute deployment in terms of racks, right?

8:03So I think you need to be... In their own racks. In their own racks or even in data centers, right? I mean, like you go talk to a customer, any hyper-scale customer or a NeoCloud or a Frontier Lab, and they say, okay, how many racks of compute can you deploy? How quickly? That will be the first question they ask you. They're like, okay, D-Matrix, how many racks of compute can you deploy for me and how quickly? Because then they can do the math in their head. you know one rack is about 100 kilowatts you know depending on you know everybody has different ratings sure somebody has 50 kilowatt rack somebody has 100 kilowatt rack somebody has 150 kilowatt racks so depending on the customer they can make the math they can run the math very quickly because they're like yeah you know what okay so d matrix you're telling me you can do like 100 racks in three months or five months or whatever the number is 100 racks times you know 100 kilowatts this is the amount of power and this is the amount of compute capacity i need which is typically measured in terms of power.

8:57So many megawatts of compute or so many kilowatts of compute or whatever. Typically, it's always megawatts. Now we're talking gigawatts, right? So they're kind of working backwards from there. They say, okay, I need X gigawatts or X megawatts of compute. That translates into X number of racks. And so the unit of deployment has become a rack. And so I think rack scale is just, you know, is kind of very prevalent unit in terms of how quickly you can deploy compute. And so that's really why we... Yeah, let me ask, because a hyperscaler, they have their own racks. And you were selling previously packages or...

9:39Cards, yeah. Cards, cards packaged into, as I recall, two cards in a, what do you call it, a unit? Yeah, we had trays and cards, right? Trays and cards, okay. And then the hyperscaler, if they were going to have a D-Matrix cluster, they'd buy a bunch of those cards and they would slot them into their racks. What's the advantage of you having your own racks? Because the hyperscaler, I mean, how does that work? Yeah. So I think two things, right? One is that model has not changed from our perspective from the point of view of the hyperscaler. Keep in mind, even when we work with the hyperscalers or the big neoclouds who also tend to buy trays and cards and will deploy their own racks, you still need to understand rack scale.

10:36As a company, we need to understand because you're talking to people at the other end who understand rack scale. So you can't be like, look, we don't understand any rack scale, but we leave it to you. No, they expect you to understand. We, at the end of the day, have to help them deploy our solutions into the racks. If anything goes wrong, we need to have rack scale people at our end who can help them. So I think there's an expertise and a skill set gap that we are addressing with this in our company. So that's one piece. But then there is another type of customer that we're going after, which is more the sovereign customer, longer term will be the enterprise customers who really want racks.

11:08They want whole racks to be deployed. Even there, we are partnering with Supermicro today to deploy those racks. So it's not like we are going to build our own rack. We're going to have our partners build the racks. but somebody needs to spec out the rack. Somebody needs to design the rack and we would have to do that for our hardware. And then Supermicro can go build it to spec but you still need to have the rack scale engineers. Now that whole process has been accelerated is all I'm saying. The deployment process has just got accelerated because of what we talked about earlier in the conversation which is the low latency computing opportunity has exploded on us and we have a solution that is well suited.

11:44So that whole process of taking our solution from where it is today to deploying it in a data center whether it's a sovereign data center or a hyperscale data, it doesn't matter. We need to have the rack-scale people at our end, the system-scale people at our end who can accelerate the process of deploying the product, and that is what the giga. Because when I read about it, it sounded to me, or I misunderstood, that it's a hardware that you're moving from simply producing cards to producing racks, and then, I don't know, having your own clouds with your own racks or providing the racks directly.

12:23But you're not building racks. It's the expertise that you require. Is everybody doing this? Not everyone. Not everyone because I would say, depending on your business model, right? I think in our case, our business model was, and I'm assuming your question is, is every other chip company looking to acquire rack scale expertise? The answer to that is yes. I would say pretty much every company is looking to acquire, but not every company is looking to build a partner the way we build our partner, right? I think, you know, a lot of the other companies that you mentioned earlier, they're building their own racks.

13:03I mean, in many cases, they sell the whole, they sell a rack. They sell, their unit of sale is a rack. In our case, the unit of sale is still a card or a tray. We partner with, you know, folks like Supermicro who would then sell the rack, right? And we would revenue share with them or there would be another business model. Or in the case of hyperscaler, we would not even worry about selling a rack because they build the rack, right? Yeah, it's fascinating, this breakout premium token market. I haven't heard it described that way. Are there other examples beyond coding that you can give that's driving that market?

13:44I mean, OpenClaw? I mean, what happened with OpenClaw end of last year? The agentic, we haven't even seen the true fallout of that, but I mean, you've seen a lot of agentic tools now being unleashed, right? Clawed is kind of incorporating a lot of agentic capabilities. Pretty much every enterprise class software tool is trying to incorporate agentic capabilities. But I think to me, two things. One is Clawed code and the significant improvement in capabilities there. And then OpenClaw. So this is agentic communication, machine-to-machine communication, where machines are kind of, you know, you can spawn multiple agents across multiple virtual machines, and a task is broken down into multiple agents.

14:33And these agents are talking to each other, essentially orchestrating all the different work or the workflows that you have running on your computer. Now, that is very latency-sensitive communication because machines are just waiting for other machines to finish tasks, right? And you don't want that to happen for too long, so latency begins to matter. Yeah. Right? And in something like OpenCore, It's also writing code on the flyers. Yeah, it could spawn an agent to write code. So it can spawn cloud code to go write some code for you on the side. Wow, that's amazing. What about edge compute? It seems to me that this would apply the latency question.

15:16Are you moving at all into the edge where latency is so important? We could. Absolutely, we could. We are not currently. The platform, the way we have built it is a platform that is built with, you know, like a Lego block approach, right? We have these chiplets that are the Lego block and then we can scale the solution up. We can, you know, put many chiplets together into a very large solution. So for the cloud market and the data center market, we, you know, essentially take our chiplet approach and we package, you know, close to, you know, eight of them on a single card. And we have 64 cards in a rack.

16:03So, you know, we call rack scale. We have 64 cards, each with eight chiplets. Right. So that gives you an idea of how many chiplets we sprinkle through the rack. Right. But you don't have to. I mean, if you went into an edge market application where you're just doing, say, for example, if it's a robot or it's a physical AI application or autonomous driving where you need AI capabilities, you don't need that many chips. You might need a few chips. And we can actually do that because of the Lego block approach. We just scale it down, sprinkle fewer Lego blocks into a physical AI application or an autonomous application.

16:42and the software would not change. The software was essentially built to scale with more chiplets or less chiplets, right? So the platform has been built today to go into the edge market. However, the edge market is very fragmented. So everyone has a slightly different software stack and each vertical, I mean. So if you go to the physical AI vertical, the software stack looks different. Autonomous vertical, software stack looks different. Or industrial, the software stack looks different. So we just don't have the scale today as a company to go address all these markets while we are trying to address the data center market.

17:14So our goal is to stay focused on the data center market. Then at the right time when the company has scale, we will branch out into other markets. Did you foresee this emerging market with coding and agent-to-agent communication? And if so, what do you foresee coming? I wish I could sit here and say, I saw it all. I envisioned it all. No, I think some of it we did, some of it was luck, obviously. And the luck is really about how quickly a lot of this stuff happened and how useful a lot of this stuff became so quickly. So I think that has certainly caught even us by surprise, right? I mean, how quickly this is maturing.

17:59And this is because AI is learning fast, right? So AI is learning fast and it is... AI is actually learning to help itself, right? I mean, so there's this iterative recursive loop whereas AI gets smart, it learns to build tools that it can use in more efficient ways and more capable ways, right? So I think you have this kind of very quick, you know, positive feedback loop that is, you know, the virtuous cycle that is underway. And that's why you're seeing these impact, you know, the impact is just so rapid, right? But, you know, we had a broad idea, right? We said, okay, the bet was on inference first, right, when we started the company.

18:38And then we said, you know, inference is going to be the big application. But then, you know, it's not just going to be about inference. It's going to be about, you know, interactive inference. And then it's not just about interactive inference, but it's about spawning machines that can take all of that interaction and do something with it, right? So I think we kind of had a broader vision of how this could evolve. and then but you know it's even surprised us and pleasantly surprised us at how it's all come together so quickly and I think what we would be most excited by next is I think we have seen agentic capabilities unleashed I think we haven't yet truly seen organizational capabilities unleashed meaning you have human spawning agents and augmenting humans but can you create teams of agents and so the abstraction of the task will keep going higher and higher right now the tasks that are getting abstracted are like okay i have uh you know i want to make i want to extract this data from this database prepare a powerpoint presentation based on the data for a board meeting that i have in a few in a few hours call it um and that can you know the agents can go spawn up and do a task like that now imagine we start abstracting to a higher and higher level where you say, it's not just about mundane tasks that I want to get done on a day-to-day basis.

20:02Imagine if we could describe a task at an extremely high level like, look, I have a customer that I want to win and you go figure out who the decision makers are, what product what my product needs to be changed to and so this thing can cover the whole strategy, like what an entire sales team and a marketing team would essentially do collectively these agents could go spawn off and create a whole strategy plan. So I think the tasks will get more and more abstracted as you go higher. And very soon, I think, you know, people talk about organizations being entirely run by agents, right? And so I think where are we and how far are we from that?

20:37We don't know yet, but I think we are going to levels of abstraction here that, you know, eventually AI could, you know, run large portions of most companies, right? Quite reliably. Let me ask first how Dmatrix is employing agents. So we have an AI team in the company that is entirely focused on rolling AI across the whole company. We just launched Cloud Code to the whole organization. We have everyone who has access. We are mandating that everyone uses AI, finds ways of using AI, and shares any new tricks or tips that they find with the rest of the organization. So it is not an option. it is mandated.

21:26This is across the hardware team, the software team, the operations team, the CTO teams, the marketing teams, the Corp.com teams, everyone. Everyone is going to be using AI. And not everybody really knows what they will do with it, but they have to start using it. And the moment you start using it, the use cases will emerge. Right. So yeah, I think at D-Matrix, we are being very proactive about the use of AI across the whole organization for various different use cases. And claw code is one thing. Open Claw is another. I'm, you know, not a coder, not a technologist. I'm a journalist, but I started using Open Claw through, you know, my claw is the interface I use.

22:12And it's incredible. And I sort of imagine if I'm doing it, you know, I'm 70 years old, I'm a retired journalist. it must be happening all over the place. Do you have an agentic framework that, I mean beyond Cloud Code that you have the organization using to spawn agents? Yeah. So OpenClaw, I mean we don't use OpenClaw because it's an open source frame. So there's obviously security concerns around, there could be, I mean we don't know what the agents are doing behind the scenes yet, right? there is a certain level of transparency that's missing. But we use Cloud Cowork. So now Anthropic is incorporating all the same agentic capabilities that OpenClaw has into Cloud Cowork.

23:09So we have Cloud Cowork Enterprise, Cloud Code Enterprise. So yeah, absolutely. We have this agentic layer that comes along with the Cloud tools that are available to the whole organization. So yes, they can unleash agents across multiple different applications. And the good thing is there is a zero data retention policy with all the enterprise tools. So anything that the employees enter into Cloud is not retained by Cloud, right? It's only used for the accession. And yeah, so I think, yes, we have already got an agentic overlay framework on top of all the Cloud tools that we're using. I've interviewed companies.

23:54There's one, Aerotechnologies, which builds essentially a C-suite co-pilot. And the idea is that a CEO will have this co-pilot in his office and will be talking to it throughout the day or using it in sessions to brainstorm strategy, business strategy. or to change business processes. Do you use anything like that when you're looking at the market, looking at your organization, trying to decide, should I buy a rack scale company? Oh, yeah, absolutely. I mean, again, a lot of this has happened in the last three to four months. keep in mind. I mean, but even before that, I always used ChatGPT at the time.

24:57Really? For that kind of thing? Absolutely. And it's surprisingly, now, of course, we use Claude. We use Claude Code. So Claude Code is, you know, phenomenal. So for both, you know, I'm doing, by the way, we did one acquisition. We are looking at more, right? And so now our M &A strategy, we actually use Claude as a sounding board for our M &A strategy, right? And it's amazing what it can do. It comes up with, you know, beautiful reports on what, how the companies could be integrated. So this used to be entire teams, remember, you know, banking teams and advisory teams that used to come in and help companies integrate and come up with a whole strategy on how to put teams together.

25:35Guess what? Claude does it in 15 minutes, right? So, and it presents a beautiful report on how the companies can be integrated and put together. Yeah. Right. Now, you know, it's not perfect because sometimes the data is, you know, not it's stale but I mean once you have a data room from the other company I mean you can throw cloud at it and it'll help you create a beautiful plan on how to put what an integration plan looks like right now till this was not available I mean till about six months ago I didn't you didn't have access to this oh no I used to use chat gpt just to run strategy questions yeah again we're thinking about product like this I mean what do you think right I would say sometimes you know chat gpt tended to be very polite right i mean it never disagreed with anything that i asked i mean so i think what you're seeing more of now is the models are just more capable and more well reasoned and well thought out um and um they do tend to disagree i mean they will tell you stuff that you're missing um uh right of course you know the personality that each of these chatbots has is is uh you know i think they've been curated to be very uh likable and amenable.

26:42But I think I've seen more debate in the conversation with these chatbots than in the past, right? And that's probably because I think as they acquired this step function improvement in reasoning capability, I think they were able to kind of reason through situations. You could say, you could argue that they became more confident. The models have now understood that they have better capability. It's almost like a human being, right? Once you realize that you have acquired a new capability, you become more confident. And same thing with these models. They tend to become more confident with disagreeing in a polite way, but highlighting, you know, things that you might have not seen, right?

27:24So I see more of that in the conversation. There's a level of trust. I mean, with the sycopancy that's addressing, were you ever concerned that the feedback or the ideation that it's helping you with is just reinforcing your own biases? A little bit of that. Initially, I did. So I wasn't sure how much to use it. I mean, now I'm feeling a lot more comfortable because it feels more like a human conversation. Yeah. I think six to eight months ago, it felt more like an echo chamber conversation. somebody out there just echoing my sentiments and maybe adding a little more color to what I was already saying and presenting it in a slightly different way and so if you are looking where those tools were more useful for me personally was if I wanted to make a point or if I had an idea that I already wanted to convey a presentation I wanted to do like it kind of augmented the way I did it.

28:26It didn't change it was augmented like look you can present this maybe talk about this this way This is another way you can say the same thing. So it really was an augmentation tool. Now it has become more of a, it's a copilot. I mean, I don't know if copilot is the right word, but it's almost like, it's like, you know, another version of me that is willing to debate with me. Right? It's not just an echo chamber. Yeah. And then on the, on the, that's also a problem with employee. Right. CEOs tend to, to surround themselves with people who are eager to agree with the CEO. Correct, exactly. Yeah, that's fascinating.

29:05Is Anthropic a customer? Could be, someday. Yeah, yeah, absolutely. But this kind of inference that you're talking about is, who are the customers that are buying the chips for that low latency? Oh, it's all the Fountual Labs. So your Anthropic would be one of them. OpenAI would be one of them. XAI would be one of them. and Meta Super Intelligence would be one of them. Google DeepMind would, I mean, all of them need this. Everybody needs it. There's no other way, right? I mean, they're going to have to find a way to use a different type of low-latency computing to augment the throughput-based computing that they've been, you know, kind of doing with HBM-based solutions.

29:52So the GPUs or other accelerators that all use HBM just don't have the same levels of interactivity that the solutions we build are having rates. I was asking about, you mentioned sovereign AI as opposed to models from the foundation labs that you're hitting with an API. Can you talk about that market? Is that because there was this huge transition? You got to feel for the IT guys. There was this huge transition, get out of your data. center into the cloud. Now, a lot of people have gone through that. And now it seems like things are flowing back into private data centers. Can you talk about that, how you see that evolving?

30:47You know, I know there was a big outflow we saw over the last, call it 15 years, where cloud became more and more popular. and but I think it was it was you know a lot of the outflow that happened to the cloud was smaller companies you know companies looking to get started I mean I'll talk about Dmatrix when we got started we didn't want to deal with an on-prem on-prem cloud right I mean we wanted to deal with you know we could just offload our entire IT infrastructure to somebody like an AWS or a Microsoft Azure and we were often running in like days you know right as opposed to setting up our own data center, setting up our own hardware, even if it's in a colo, right?

Read the full transcript

31:29I mean, doing all, I mean, that was a traditional model even for small companies once upon a time where you set up your own, you know, small data center or, right, and you kind of ran your own hardware, right? And we didn't have to do any of that. Didn't have to worry about security. Didn't have to worry about scaling that hardware, deploying the tools. None of that, right? I mean, so you see, I mean, that trend has been very strong and will stay strong even to this day. All the AI native startups, right? Companies that are building these new age applications for AI, they're all running in the cloud.

32:01They're all still running in the cloud, right? Now, what you're talking about is maybe companies that were, you know, kind of further along in their journey, much larger, mid-size to large-size software companies, Fortune 500 companies that carry a lot of data. See, they never fully migrated to the cloud, right? There were applications that they migrated to the cloud, but there were certain applications that they didn't migrate into the cloud. or there's certain things that were data sensitive, their data never migrated into the cloud. Now the only thing that is different now is you still have the same structure.

32:31It's a hybrid approach for some of the larger organizations where some of their applications run in the cloud, but some of their applications that tap into their native or domain-specific data, they don't want to have those applications running in the cloud, so they're still running on. So they're maintaining both. They're maintaining an on-prem data center, they're maintaining, you know, they have a cloud data center or partner also, right? Now, AI doesn't change any of that, right? AI just is an overlay on top of that. So now you use AI to access the applications running in the cloud, depending on the type of applications you're running, those applications will get infused with AI.

33:06And you have an AI overlay on top of those applications. The same thing goes for what they're doing inside their enterprises, right? The only thing is that the AI overlay that is happening in the cloud is happening much faster, right? Because the hyperscalers, Google Cloud, Microsoft Azure, Amazon, they are very AI savvy. So they are able to bring AI capabilities into all their cloud offerings a lot faster than an enterprise that has its own data and its own tools internally can do. So bringing AI capabilities natively to a Fortune 500 company, for instance, will take much longer than their applications running in the cloud at say Microsoft Azure will become AI friendly a lot faster.

33:50So I think that's the only thing that you're seeing. It's not like they're looking to pull back stuff that was running in the cloud back into their organizations because they left the organization in the first place because they were not concerned about the data that they were accessing. Yeah. Although in regulated industries, there is concern about sending data to the cloud. Do you deal with a lot of... That concern has always been there. It's not like a new concern with AI. It's concern has been there like 10 years ago. So I don't think anything changed. I think it's all about where the data is resident, right?

34:31And that concern never changed, right? Because this data has been around for decades in many companies and they were never quite comfortable sending that data into the cloud, right? I think the one thing that you might be referring to is most of these companies have access to an OpenAI API or an Anthropic API, right? And now there is new data. I mean, and what people are doing is they're kind of conducting searches. And that tool has to get in and access databases and has access to that data, right? And, you know, is that something you want to do, right? I mean, I think so. We have the same issue.

35:08We have all our proprietary data sitting in a data silo somewhere. And now Anthropic can access all of that data, right, to do things for us. Right. And that's where they have the zero data retention. Right. None of that data ever goes to Anthropic. Right. And at least the enterprise class too. So that's a guarantee that they're giving you that all the data that they touch within your organization is not going to them. Right. So it's only, and I think maybe that was your question originally. is like, hey, there is an API that you guys are accessing through Anthropic or OpenAI, and that API gets access to some of your internal data because you are taking some of your internal databases and asking Anthropic or OpenAI to do things for you.

35:52Yes, so that is a concern, I think, has been addressed by these companies, right? And in some ways, it's probably not very different from you running, for example, Microsoft tools, Microsoft Open Office 365 or M365 has a lot of your information, a lot of emails, but it's sitting, Microsoft tools have access to it, right? So in many ways, it's not very different from that. This market, as you said, it's expanding maybe faster than even you anticipated. Are you primarily focused on the U.S. market? I really am, yeah. Yeah, I travel a lot, and, you know, every country. I ran into Rodrigo Liang in Saudi Arabia a year or so ago.

36:47Are you looking at these other centers that are trying to build out compute and AI infrastructure either for their own country or for, you know, there's a lot of talk about who's going to be the leader of the global south of the AI movement? Right, right, right. We are spending just the right amount of time, not too much. I think we are focused a lot on the US market because that is where the opportunity will emerge and then eventually it will diffuse into other countries in the global south, right? So I think we want to start at the source of the problem, not go to the destination directly. So I think we're starting at the source where a lot of the innovation is happening.

37:33People are trying to figure out how to deploy these solutions, what makes the most sense, which applications really need it, which applications don't, where does it work well, where does it not work well. Once those questions have been answered, then it'll diffuse into a lot of the sovereign applications, into a lot of the enterprise applications. We'll ride that diffusion process. But it has to start from the source, you know. Yeah. Do you pay attention to what's happening in China? Not so much. We, of course, pay attention to all the work that's happening there. and you know I just wonder are they moving in the same direction on you know this premium token market and all of that they have to be they have to be I mean I just don't see why I mean that's a very intuitive thing right like I want to create I want people to access applications I mean everybody wants to access an application in a highly interactive and a real time way why would that be any different for a segment of the population in China, right?

38:33I mean, if you need it in the US, you need it in China, right? Yeah. Are you allowed to sell to China? Under export controls, we would be allowed to sell to China, yeah. I recall most of your chips, if not all, are manufactured in Taiwan at TSMC.

38:51Silicone is etched in Taiwan. That's right. You know, TSMC is building a fab here. I don't know if it's online yet. Yeah, it is not online. It is online. They're building multiple fabs in Arizona, and at least one of them is online since 2024. Have you seen any of your manufacturing or fabrication moving back from Asia? Not yet. Not yet. Still in Taiwan, yeah. Yeah, I would imagine, and I asked you this last time, and I get the same answer from everybody. Oh, we have relationships. But is there any sort of a bottleneck? I mean, you know, can TSMC serve the world as the demand for inference explodes?

39:47Well, you know, the chips are going to be in short supply for the next, you know, call it three to four years at least, if not more. So we have memory, you know, chip shortage. We have compute shortage, right? There's a reason why Elon wants to build his own TerraFab, right? Because he feels TSMC cannot meet all the demand for compute. Yeah. and memory. So we'll see how that goes. And Intel is helping him out with that. So we'll have to see how it all works out. But I think for the demand that we have, D-Matrix, if I would look at the D-Matrix demand picture for the next, call it five years, I think TSMC has got plenty of supply to build our demand.

40:36We are a small company. We're looking to grow into this opportunity. This opportunity is also very nascent. It's just exploded on the scene literally in the last few months. It's got a long, long way to go. So I think we are growing into an opportunity that's pretty new. And the choices we have made on how we build our product, like we don't use the bleeding edge process technology at TSMC. We use kind of N-1, N-2 process nodes. We don't use any HBM technology. We don't use any Covost technology. So all these decisions that we made early on help us. So yeah, if you look at a supply chain that uses those technologies, like the ones I just outlined, I think TSMC should be able to make, you know.

41:19Will it be tough? Yes, for everyone. Tough for us, tough for everyone else. But I think we have a better shot at making supply because of the choices we made. Yeah. Is that a constraint on your growth at all? Not yet. Not yet. But it would be a good problem to have, right? I mean, you know, if I get to the point where my growth is constrained by how much TSMC can supply to me, I think I'm in a good place, right? I think then I have to just kind of work through the problem, which we will work through the problem with them, right? And to me, I think it comes down to the relevance of the problem, right?

41:49If your problem that you're solving is very relevant and there's lots of big customers who care about solving their problem, yes, you have relationships and that's all great, but at the end of the day, you know, just follow the money, right? I think if there's a very relevant, large problem that a very relevant, large customer is trying to solve, and we think we are helping solve that type of problem, then the supply chain will make a place for you, right? Because the world needs this, right? So I think they'll find a way to make place for you. Yeah. You talked about the agentic explosion and about the auto, whatever you call it, AI coding explosion.

42:35Which of those is the larger market for you? And what kinds of companies are you selling to? You were saying you're selling to the foundation labs, not necessarily anthropic, but whose coding co-pilots are you providing chips for and whose agentic systems are you providing chips for? So agentic, I think, is not a separate... Agentic is a capability that I think pretty much every company that is building a coding tool is embracing or is introducing into their product offering. Yeah. Right? So I don't think you should look upon it as, okay, there is a segment of, a section of companies that's only working on agentic.

43:27because Agentic is like, you know, kind of a capability that is now being, OpenClaw demonstrated what's possible, but now you take that capability and introduce it into your products, right? So everyone is doing it. OpenAI is doing it. XAI is doing this. You know, obviously Anthropic, because they sell into the enterprise, they're doing it faster than anyone else. But it's a capability, right? And that capability adds more influencing needs on top of the coding generation tools that are already need inference, right? So you have, it's kind of putting it together. We are working with everyone, right?

44:04I mean, we're working with all the Frontier Labs that build coding tools, agentic coding tools. We're working with Frontier Labs that are building video agents, voice agents. We're working with AI Native Startups building voice agents. we just had to be careful about not going too broad because I think the company is still not at the stage where we can support all these customers simultaneously so we had to be kind of gradually grow into this opportunity with at the right time with the right company but we're doing it with everyone because it's relevant again going back what we are building is relevant for everyone everybody wants it it's not like there's somebody who says I don't care about high levels of interactivity I don't care that my users are able to interact with my application in real time.

44:49Nobody's telling us that, right? Everyone wants fast tokens. Yeah. Or is there a segment of the market where you think, you talked about an agentic enterprise where agents are doing much of the work. Do you see emerging companies, startups that are agent first and do you think that segment of the economy is going to grow? I mean, everyone's sort of watching. You know, there was a New York Times outpiece recently about a guy and his brother running a$1.6 billion in sales company that's basically an agentic first company. Do you see that segment of the economy growing? I mean, that's the future.

45:46That's the future. Is that going to challenge legacy companies in different... I think it depends on what kind of product offering or products or what kind of service you're building. I think there is a place. I mean, you have those companies today. I mean, even pre-AI, right? I mean, there was companies, you know, you had people running a consulting practice, right? I mean, you are a one man. Are you a one man? There you go. So you are right. You're already there. Now you're going to start using AI, right? I mean, so you could be one of those people like, hey, look, I've been far ahead of everyone.

46:25I've been a one-person company for a long time, right? And now I just use AI as an augmentation tool. So it really depends on the product offering and service, right? I mean, I think there is a lot of companies that had a product offering or a service that can be taken over by AI completely. And you don't really need humans for the kind of work that they were doing. So there will be kind of the low-hanging fruit, I would say, that will immediately get consumed. Either they embrace it or then you have what you just said is a new breed of companies just come and say, wait a minute, why are those companies, you know, why do they have humans in those companies?

47:06You know, we can just do it with AI and then they will just become a lot more efficient. Or those companies embrace it themselves and, you know, reskill the humans into something else, right? Or they find that as an opportunity for those companies to grow into something different. Because humans are good at something which AI is not good at. So why don't you reskill the humans to do stuff and expand the opportunities you can go after as opposed to sticking with what you have. It's really an opportunity for everyone to grow into something new. But absolutely, there'll be many more companies that will be tooled only with AI to address those products and those services, right?

47:42But then there are certain products and services you really just cannot build with AI only. I mean, you need manufacturing lines and you need good product design and there's emotion that goes into how the product is sold and how the messaging is done and how you appeal to consumers. I mean, all this is something that you need humans. Because at the end of the day, you're making a sale to a human. At the other end is a human who's buying the product and you want to appeal to their sensibility and their emotions. I think not every section of the economy is going to get consumed by AI-only companies, but there will be a lot of them, right?

48:22A sizable portion. Yeah, healthcare is an industry that's going under through tremendous change. Do you see much demand for agentic AI in the healthcare space? Are you serving that space or are you focused on a few verticals? Yeah, we are not very active in the healthcare space. I can't meaningfully comment on what the latest and greatest is there. We are serving the tools. The tools get used by various verticals. Healthcare is certainly one of them. but I think to me the most exciting thing in healthcare is really a drug discovery you know you know you know disease you know curing you know really complex diseases you know producing new forms of treatments that could accelerate the development of drugs or or find cures right I think so I think to me to me that that segment is is is a lot more exciting and you know that is one of the goals of AI should be one of the goals of AI is yes you know more equitable wealth distribution sure but really can we can we make you know health benefits you know broadly available can be and easily accessible to parts of the population that don't even have access to it right and so can we get humanity to a point where, okay, everybody has enough.

50:03They can live long lives, long, healthy lives, right? And they can live peacefully. I mean, that's maybe the third quest, right? We can... Now moving maybe into a bit of a utopian vision here, but that would be the vision, right? It's like, there's enough wealth for everyone, enough health for everyone, enough peace for everyone. and then AI has really helped humanity, right? Yeah, that's right. Well, we're hoping. It also requires political leaders. That's right. That's right. Exactly. Okay.

From the publisher

The AI chip market looks monolithic from the outside - NVIDIA dominates, and everyone else is fighting for scraps. But d-Matrix's CEO Sid Sheth argues that the market is quietly splitting into two distinct tiers, and the one that's exploding right now is the one NVIDIA's architecture isn't built for. In this episode, Sid joins Craig Smith to explain the "premium token economy": a new class of AI inference where interactivity is the product, users pay ten times more per million tokens for instant responses, and the memory bandwidth limits of GPU-based systems create a structural ceiling that purpose-built architectures don't have.

The conversation is unusually candid about what AI actually looks like at the executive level: Sid describes using Claude as a sounding board for M&A strategy, producing full integration plans in 15 minutes that used to require entire banking advisory teams, and watching AI shift from a tool that echoed his ideas back at him to one that genuinely disagrees, flags what he missed, and pushes back with enough confidence to be useful. He also makes the case that we're at the beginning of a shift from individual agents to what he calls "organizational AI" - teams of agents running entire company functions at a high level of abstraction - and that the infrastructure bet d-Matrix is making positions them directly in the path of that wave.

Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.

More from Eye On A.I.

All 266 episodes
Why People Are Paying 10x More for AI - and What That Means for the Chip MarketEye On A.I. · 51 min
Listen in VO