In short
Podcast Episode Notes: Eye On A.I. - Episode #251 "Sid Sheth: How d-Matrix is Disrupting AI Inference in 2025"
Overview In this episode, Craig S. Smith interviews Sid Sheth, CEO and Co-Founder of d-Matrix, a startup that is revolutionizing AI inference hardware and competing with established players like NVIDIA. The conversation delves into the significance of AI inference, the architecture of d-Matrix's products, and the future landscape of AI hardware.
Key Highlights
Introduction
- Guest: Sid Sheth, CEO and Co-Founder of d-Matrix.
- Focus: The disruption of AI inference through innovative chip architecture.
Main Discussion Topics
The Shift from Training to Inference
- Key Insight: The future of AI is in inference rather than training.
- Market Opportunity: d-Matrix targets the inference market due to its potential for efficiency and performance improvements.
d-Matrix’s Corsair PCIe Accelerator
- Performance Claims: Claims that Corsair outperforms NVIDIA's H200 in efficiency and silicon usage.
- Silicon Area: Corsair utilizes 3200mm² on a 600W card compared to NVIDIA's 800mm².
- Architecture: The Corsair uses a chiplet-based design that integrates multiple chiplets for optimal performance.
Importance of Memory and Bandwidth
- Memory Bandwidth: Corsair's architecture emphasizes high bandwidth memory to tackle memory bottlenecks.
- In-memory Compute: Utilizes SRAM to store model weights, aiming for more energy-efficient operations.
Integration with Hyperscalers
- Business Model: d-Matrix focuses on seamlessly integrating their accelerators into existing hyperscaler environments (e.g., Google, Amazon) without requiring customers to overhaul their infrastructure.
- Cloud Augmentation: d-Matrix products are designed to augment existing cloud server capabilities for AI inference.
Market Dynamics
- Heterogeneous Infrastructure: The future of AI hardware will require a diverse set of solutions tailored to various inference needs.
- Global Market Outlook: d-Matrix's strategic market focus includes the US primarily but is also exploring opportunities in APAC and the Middle East.
Competitive Landscape
- NVIDIA's Dominance: Sid discusses NVIDIA's strong position in training and the complexity of breaking into that market.
- Future of Competition: He predicts a shakeout in the semiconductor industry, suggesting multiple players will emerge but not at the current saturation level.
Future Innovations and Opportunities
- Technological Advancements: Sid emphasizes the ongoing transition in AI architecture, hinting at significant growth opportunities in inference computing.
- Call to Action for Talent: d-Matrix is on the lookout for talent to help solve complex technical challenges in AI inference hardware.
Conclusion
- Sid Sheth's insights provide a comprehensive picture of the evolving AI landscape, particularly focusing on inference and the role of dedicated hardware solutions. The potential for d-Matrix to carve out a substantial market share in a rapidly growing sector is substantial, especially as the need for efficient AI applications increases.
Additional Resources
- Podcast Links:
- Craig Smith on X: [Craig's Profile](https://x.com/craigss)
- Eye on A.I. on X: [Eye on A.I. Profile](https://x.com/EyeOn_AI)
- Sponsor: DFINITY Foundation—focused on decentralized cloud computing through the Internet Computer (ICP).
- Further Listening: Episode with Dominic Williams about the Internet Computer.
Key Takeaways
- The shift towards inference computing is not just a trend but a fundamental change in how AI applications will be developed and run.
- d-Matrix's technological innovations position them well against established competitors by focusing on efficiency and integration.
- The inference market presents vast opportunities, and the landscape will likely consolidate over time, leading to fewer, more robust players.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We work with TSMC today. That is our primary fab partner. And a lot of the prototyping work we do is with TSMC. And, you know, they have a lot of programs where you can try out your ideas on a smaller scale, where you can build smaller chips and test out your technology ideas. And we have done a lot of that. We actually, for every generation of product, we've typically done a few test chips before we productized it. And they've been very supportive. And it is very possible to do that. So that's kind of been our way of getting our products to market and testing out our technology before we actually productize it.
0:34This episode of Eye on AI is sponsored by the DFINITY Foundation. DFINITY Foundation is a Swiss not-for-profit that is home to some of the world's leading cryptographers, computer scientists, and experts in distributed computing. Their mission is to shift cloud computing toward a fully decentralized state by supporting the Internet Computer, also known as ICP. If you don't understand anything about the internet computer, you can see my episode with Dominic Williams. ICP's vision is that most of the world's software will be replaced by network resident software. That's an evolution of traditional smart contracts.
1:23To achieve this vision, ICP is designed to make smart contracts as powerful as traditional software, all while remaining tamper-proof, unstoppable, transparent, and verifiable. Since DFINITY's launch in 2016, they stand as Switzerland's most extensive blockchain research and development initiative and have been awarded more than 500 research grants worldwide. DFINITY remains steadfast in its mission to drive the advancement of the decentralized internet. Network resident software can now be used to run AI models and RAG infrastructure, preventing them from becoming quote-unquote hot wallets from where data can be stolen and increasing resilience.
2:15Furthermore, the technology has been designed to allow AI models to spin up and modify running web applications and internet services solo by addressing several key challenges. If you're interested in reading more about the internet computer, visit internetcomputer.org. I also encourage you to listen to my episode with Dominic Williams in which he explains the internet computer. So thanks for having me on the podcast. You know, I've been a systems and semiconductor guy for many, many, many years. Started my career at Intel in the mid 90s. Actually, I spent the first 10 years of my career being a chip designer.
3:10and i you know spent a lot of time doing different different portions of the design work at intel in the first five years that i was there worked on circuits you know bleeding edge circuits for the pentium processors worked on you know system issues related to bringing up processors After that, I went to a startup called Eloros, which was focused on building interconnect chips for Ethernet applications. So that was the next five years of my career. So really the first 10 years of my career were spent in heavy R &D and really working on both processor chips and chips built for interconnects and networking.
3:58After that, I want to say, you know, I transitioned. I was still at the startup and I was looking to transition into something that was more customer facing and moved into a business role. Spent another five years working on different aspects of what it takes to take a product to market. Really focused on go to market and I had an advantage because I was essentially taking the chips that I had helped build and working with customers to figure out how to deploy them and making them useful so to speak for those customers after which i transitioned to a bigger role at a company called in5 and in5 was in the high-speed networking and high-speed interconnect space they were looking to break into the data center market they had a really good you know footprint in the telco market but uh were looking to break into the data center market so i went there and incubated a business for them which was only focused on data centers and that was a huge success for the team spent about nine years building that business that eventually grew to two and a half billion in revenue from zero and was acquired by marvel in 2021 in a 10 billion acquisition.
5:24I stepped out in 2019 to start Dmatrix because I was getting very excited by the inference computing opportunity. And I know it sounds very obvious today, but it wasn't very obvious back then. A lot of the narrative was about training and training bigger models and doing inference with image classification and image models were really the you know, all the rage back then. And we quickly realized that it didn't make much sense for us to go do something with training because that was a market that was well capitalized, you know, a lot of investment. Didn't make much sense for us to go after the influence market in the edge that was focused on images and image classification.
6:14So we kind of were looking for a new entry point and that is how we started focusing a lot on language and transformers. And I'll talk more about that as we talk about the D-Matrix journey. But that's been my journey. I grew up in India in a family of entrepreneurs. Both my grandfathers were entrepreneurs back in the 1940s, left home to start their businesses. My father was an entrepreneur too, left his home, came to the U.S., worked here, went back and started a business. so entrepreneurship is probably in my DNA somewhere it was a matter of time before I stepped out and did something of my own and here I am taking my shot at this, we'll see where it goes yeah, well let's talk about Dmatrix's products and how you establish market share in a market that is so heavily dominated by NVIDIA and the other big players.
7:24Yeah. So I think I've lived through the process of building a semiconductor business before. As I was stating earlier in my introduction, you know i was when i was at in five we essentially incubated a business to enter the data center market you know building interconnects for data centers and what we saw was and i'll come to your question in a bit but this this this perspective is important because that answers the question that you just asked which is in the semiconductor business especially when you're building chips you know the very capital intensive obviously and take a long time you know multiple generations of product to to get to success and the only way to really be successful is to really intercept some kind of a discontinuity in the market there has to be some kind of transition that is that is underway and you need to intercept that transition at the right time.
8:29So at InFi, the transition that we intercepted, there's really two major transitions that happened at the same time, was that data centers were going optical. You know, most data centers used to be based on copper interconnect, but they were transitioning to optical interconnect because obviously the speeds were getting to a point, bandwidth was becoming very, very critical, and the higher the bandwidth, you know, the harder it gets to run signals over copper. so it was transitioning to optical and the other thing that was also happening at the same time was the signaling standard which used to be nrz which is the one zero standard was transitioning over to pam based signaling so four level signaling so we were going from a two level signaling standard to a four level signaling standard and so that required new electronics new optics and so you had this tech transition and a market transition happening at the same time um and that is what we intercepted very successfully and you know the incumbent back then were companies like broadcom and marvel and they were giants right but in five was successfully able to break into the market and dislodge some of the incumbents and build a very successful business doing that right so same thing applies for where we are today uh we see you know major transition happening in the market again obviously uh there is the whole transition from general purpose computing to accelerated computing which is a mega mega change that is happening across data centers but within that accelerated computing we are also seeing the market transitioning away from training to inference and entrance becoming a very relevant workload and that is also a tech transition because you need a fundamentally different architecture as migrate away from training which is predominantly a performance oriented workload to inference which is an efficiency oriented workload so you can't really build a common platform that does both training and inference um you can it's not like you can't people have done it but it's not going to be uh it's it's like building a swiss army knife right you're trying to do too much with the same thing building a dedicated inferencing solution um given that inference is going to be a much larger portion of the opportunity for AI compute, which means we think a dedicated solution for inference is going to win, right?
10:56And that's what we are focused on at Dmatrix. And that's why we think we can find enough share in the market to build a very valuable company. Yeah, I have to ask, I live not far from IBM Research. and I was over there the other day talking to their head of chip design, and he was saying they have a fab up, I think, in Poughkeepsie where they do pilots or proof of concepts or whatever. so once you've designed a chip which which is a feat in itself uh where do you go to to build a prototype and and see if it's working so we do all our work you know we first of all we are building a product that is pretty right you know is built on standard cmos process technology so we are We are not doing anything funky here with the process technology that requires us to go to boutique fabs and build partnerships with fabs that work on very exotic processes and process technology and whatnot.
12:18So standard CMOS logic process. So we can work with TSMC. We can work with Samsung. We work with TSMC today. That is our primary fab partner. and a lot of the prototyping work we do is with TSMC and you know they have a lot of programs where you can try out your ideas on a smaller scale and you know what they call the the MPW program or a shuttle program where you can build smaller chips and test out your technology ideas and we have done a lot of that we actually for every generation of product we've we've typically done a few test chips before we productized it um and they've been very supportive uh and it is very possible to do that so that's that's kind of been our way of getting our products to market and testing out our technology before uh before we actually productize it yeah uh and and as far as the inference market goes as i was saying before we started recording uh i've spoken to cerebris and SambaNova, with training, NVIDIA has a lock because developers are used to using CUDA and they don't want to have to learn a whole other programming language, I guess, or instruction set and all of that.
14:00But with inference, developers are primarily using hardware in the cloud and, you know, through API. So they don't have to worry about building on top of the chip. Is that the opportunity as well that you see? And how does Dematrix, I mean, I had a conversation yesterday with Google Cloud asking about how these new accelerators get a footprint in the cloud. And I know Sambanova and Cerebrus are both building their own clouds until there's enough of an uptake that the hyperscalers will offer clusters of their chips. So how do you do that? Yeah. Yeah. So, first of all, the team at D-Matrix has successfully sold into the hyperscalers before.
15:18We've been through, call it 10, 15 hyperscaler ramps with products, not at D-Matrix, but through our prior experience. So we know what it takes to ramp a product at a hyperscaler. We've been doing it for 20 years. When Google first had a cloud offering, and not necessarily a cloud offering, but they were setting up infrastructure to power their own applications. They were setting up pretty enormous data center infrastructure back in 2007, 2008. We were building chips that would help them power that infrastructure. This was back you know many many moons ago my my my startup from you know the 2000s when we were building chips to connect networking uh you know equipment we were selling products into google to to help them power their infrastructure so and you know all the way from back then to the infight journey where we had our chips sitting in pretty much every hyperscaler on the planet all all four big hyperscalers here in the US, three in China.
16:30So we have a lot of experience when it comes to deploying product at hyperscalers. We know the anatomy of what a typical hyperscaler ramp looks like. We know what the hyperscalers are typically looking for and how they engage, right? So at Dmatrix, we are kind of modeling our business along similar lines. We are building product that would embed into the equipment and infrastructure that the hyperscalers are deploying. So we are not going to them saying, look, you need to rip and replace what you have to accommodate us. We are looking at their equipment and we try to understand what they deploy and we try to fit inside what they already have without them having to rip out stuff to accommodate us.
17:17So really what we are trying to do is a cloud augmentation play, right? So we can go and we build accelerator cards and maybe I can show you the card while we are on the topic. So yeah, this is our 600 watt Corsair accelerator card. It's a PCIe card. And it's got two chips here that you can see. Each chip has four chiplets. so there's a collection of eight chiplets sitting on a single card that are you know connected to each other in an any to any configuration and then if you turn this card around there's a lot of ddr memory that is sitting at the back this is a pretty dense card it's like a 24 layer card a lot of engineering that has gone into making this because we have a lot of you know interesting chiplet technology that goes on to onto this card that needs to be routed very efficiently um so that's that's you know the the guts of the card and then you know that gets uh packaged into into this accelerator unit you see two cards there are put together um and then they are connected with a bridge uh on top which allows every every car every chip on every car to talk to a chip on the other card directly, right?
18:38So from a software perspective, this is treated like a single unit from a software perspective. So just as NVIDIA has NVLink, we have the ability to connect all these chips sitting across these two accelerators into a single config from a software perspective, right? So that allows us to unlock a lot of efficiencies when we are running inference um so coming back to your question i mean that's that's what we ship that's the card we ship that's a unit we ship and that plugs into any ai server that is built by uh you know folks like dell or hp or super micro or you know lenovo or call it even if it's a custom server sitting at one of the hyperscalers as long as it can house a pcie card we can plug into that socket and augment that server for AI inference.
19:37So that is our go-to-market model where we go to the hyperscalers and we tell them, hey, you've got a fleet of servers that currently don't have our hardware in them, but what if you were to put our hardware into the fleet of servers? We could augment that entire fleet with a solution that will completely change the economics of doing generative AI inference. Now, wouldn't that be attractive to you? and they pretty much across the board love that right because it doesn't require them to pull out and rip out stuff but it allows us to integrate our solution into the ours and and then they can go augment that fleet with more capability and uh talk a little bit about the uh the technologies uh involved in this chip you're you're putting the weights directly on uh on sram is that right which, as I recall, is at least what Cerebris is doing.
20:35I think SambaNova stacks memory on top of the chips. I don't remember. But the static random access memory has the model weights directly in them. can you talk about that how that speeds things up how it's more energy efficient and then talk about the chiplet architecture yeah sure so a few few questions in that one question So let me unpack some of that for the audience, right? So first of all, the chiplets themselves, right? So you see this is, you know, a chip here. You know, there's obviously a lid that if we kind of pull that lid off, you'd see four chiplets inside a chip here and four here.
21:37So there's a total of eight chiplets sitting on this 600-watt PCIe card. if i remember correctly the only other 600 watt pci card on the market is the hopper h200 blackwells are still not available in a pcie form factor still i'm sure they will make those but there's currently not available so the only pci accelerator from nvidia that's the latest and greatest is the h200 now the h200 has about 800 millimeter square of silicon on a 600 watt card right uh we have across these two chips 3200 millimeter square of silicon so that's four times the amount of silicon that is sitting on this card at the same power envelope in the same power envelope 600 watts compared to the latest and greatest nvidia offering and and the reason we are able to do that is because of the chiplet technology and the chiplet approach and obviously a lot of design techniques that we have built into the chiplets that allow us to run the chiplets in a highly energy efficient manner so we have you know almost 4x the amount of silicon on the same 600 watt power envelope right and so what's in that what's in those chiplets right so obviously we've got a lot of compute but again you know the bottleneck is not compute right if you look at generative ai inference the bottleneck is really memory bandwidth and then eventually memory capacity and it's about finding the right balance between compute memory bandwidth and memory capacity and that's what we have strived for on this solution right so we have um obviously 10 petaflops of compute in on that card that I just showed you on 10 petaflops of, you know, four bit compute, sitting on on that one card, which is, you know, plenty for now.
23:35And but the key thing is to optimize around memory bandwidth. Now, you talked about putting, you know, the weights in SRAM, we have a combination of our in memory compute and SRAM on those chiplets. So there's a combination of, you know, call it across that one card we have almost 2000 in-memory computing cores and we have a lot of sram also that is built into those chiplets so we have almost two gigabytes of what we call high performance memory so this is memory that is sitting on the chiplets directly but running at a very very high bandwidth you know call it 150 terabytes per second of bandwidth you compare that to hbm memory that is sitting at you know roughly about you know 4.8 terabytes or you know five terabytes the blackwells are closer to eight or ten but the hopper that i you know was comparing our product to because that's the only other pcie product is sitting at about you know five terabytes right so there's almost like a 20x 20 to 20 30x improvement on memory bandwidth now that is where we kind of punch through the memory wall right uh the you know it's not about just having a lot of compute but it is having a lot of compute and a lot of memory bandwidth to go along with it so you can feed that compute very efficiently from memory and the memory is you know is is local now uh there's about two gigabytes of that which is you know typically not enough for a lot of the ai models so we we scale this up further so we'll have you know eight pci cards sitting in a single server and then we have eight servers sitting in a single rack so we have a collection of 64 cards sitting in a single rack.
25:20So 64 times 2 gigabytes is about 128 gigabytes of that high performance memory that is sitting in a single rack. With that 128 gigabytes, you know, we can easily serve, comfortably serve 70 billion parameter models or 100 billion parameter models out of a single rack, which is for the most part plenty for anyone who, you know, most enterprises, is most, you know, very high end enthusiasts, many sovereign opportunities that we are going after, including some of the hyperscaler opportunities. A hundred billion parameter model sizes are pretty attractive. They can cover a lot of ground with that type of model.
26:00Now there are frontier models that are much larger and we can certainly run the frontier models across multiple racks or even out of the capacity memory that I showed you. There is, you know, DDR memory sitting at the back. You know, we can run trillions of, you know, trillions of, multiple trillion parameter models running out of that capacity memory. They would run a lot slower, but you can't run them out of that capacity memory. So we have the ability to do this, you know, you know, we have this flexible memory architecture where we can run 100 billion parameter models in a single rack, or we can do more across multiple racks.
26:33But it's running at these extremely high speeds because of the memory bandwidth problem that we addressed. And all this is possible while embedding our hardware in a customer's environment without having them having to purchase the whole solution from us, right? We can go augment their infrastructure with our solutions and give them the benefits of running what we call ultra low latency batched inference, right? So extremely fast while allowing the user to batch multiple or the customer to batch multiple users. So really that's what we bring to the table. It's a unique combination of this chiplet, in-memory compute, technology, highly energy efficient design, how we connect the chiplets to each other, the numerics.
27:30You know, we have 4-bit numerics and 8-bit numerics that are block floating point, which is now becoming a standard for inference. All of this goodness built into this one solution, one PCI card form factor that can go augment our customers' fleets with, you know, and supercharge those fleets for generative AI inference, low latency generative AI inference. And are most of your implementations on-premise or is this primarily a cloud solution? It seems with inference, that market growing, and as large enterprises start using generative AI, that a lot of their putting it in on-premise because they don't want their data to leave.
28:29Yeah. Yeah. Yeah. Yeah. I mean, it's bought. I think we are really quite agnostic, right? As I was stating earlier, our go-to market is built on taking our solutions and augmenting our customer solutions, right? So we don't quite care whether the solutions are in the cloud or in the enterprise. But the journey for most AI deployments is starting in the cloud, right? So I think our journey is starting in the cloud today, whether it's sovereign clouds or what we call GPU clouds or Neo clouds, which is the smaller dedicated AI clouds and then there is the hyperscaler clouds so these are really the three focus areas for us today and then over time this will migrate into on-prem enterprise clouds and we'll migrate as the workloads migrate there but right now you know the high velocity deployments are happening in in the clouds that are public clouds and then eventually they will go to the private clouds over time yeah the where do you see this going with as I said I mean Nvidia has such a dominant position is the market opening up I mean as I was saying too if you're in the cloud and this is for inference, what matters is speed and cost.
30:05Do developers care what chip is giving them inference? I mean, do they make that decision? Or if it's in the cloud, is that a decision that's being made by the cloud provider? Yes, yes. So typically, I think you need a person who really understands the application that they're trying to accelerate. And what is happening is there's plenty of data that suggests that applications which are highly interactive with the user tend to be more successful, meaning the users are more likely to stay with an application that really interacts with them. Right. Anything that is more offline in the interaction, you know, doesn't tend to keep users around for too long.
30:51right so anything that draws eyeballs is stuff that is highly interactive so latency matters right clearly so you know you need to look at the application i mean there is obviously you know other constraints like cost and uh energy efficiency and whatnot size of the model uh how big is the model and uh you know you got to kind of look at the whole picture to figure out you know what kind of underlying accelerator hardware could really allow for a better user experience and a you know better uh cost profile right um so it's a combination of these things so yes the hyperscalers have the people have the right expertise and the right uh you know uh folks with the right set of skills to make those calls uh they can they can match the right hardware to the right application right and um and then they can go deploy the and they also have the ability to go deploy new accelerators in their fleet to enable that user experience.
31:48So if a new accelerator solution can demonstrate that we can improve the user experience which is latency or we can improve cost or we can improve things like energy efficiency by a factor of at least 2 to 4x, that is a big win typically for a customer. right so if you have the latest and greatest nvidia hardware and you can show a two to four x improvement over that in that generation that's a big win that's a very big win right so so again for inference i like to say it is it has been hard to beat nvidia on training and you know they've done a terrific job of keep you know you know anybody who was at gtc earlier this week and saw the keynote i mean they've done a terrific job of advancing their roadmap and you know they have a roadmap that goes out you know four years and five years and you know it's It's almost impossible to kind of play catch up there, you know, with all the technology that they bring.
32:40Well, when it comes to inference, I think there is no one size fits all. You can't really take a big rack of blackwells and just throw it at every inferencing problem. It's not going to work. It's not like every single inference request looks the same. Inference requests can be latency, you know, may need to be latency optimized, may need to be cost optimized, may need to be throughput optimized, right? So there is no one size fits all. You can't say like here is one piece of hardware that's going to serve all of inference. It's difficult to do that. So inferencing fleets are going to be heterogeneous.
33:15There's going to be GPUs. There's going to be CPUs. There's going to be accelerators. And that is what we are focused on. We are not saying that we are going to be the best at everything, but we are really good at low latency batched inferencing. So anytime you need multiple users where each user needs a really interactive experience, that's where you need to think of Dmatrix. Yeah. You know, I spoke a couple of years ago, and I was just looking to see if they're still independent. but there was a company, Rescale, that has kind of an orchestration technology that plugs into all the clouds, and then you give them your problem, your inference problem, and they automatically route you to whichever cloud has the correct chip or the correct resources.
34:19resources is that uh something that uh that did you work with people like that yes yes we work so by the way we work uh we worked with rescale uh you know we know them well and um you know they've uh but you're absolutely right i think that kind of skill set by the way already exists at the hyperscalers uh so you know they're they're not uh they either partner or they have the skills inside the company to build that orchestration layer in their case you know they they are doing query profiling they have a query coming from a user they're trying to batch that query based on the profile of that query right whether it's a query that is latency optimized query or is a throughput optimized queries or a cost optimized query so you know depending on the region it comes from right you may have to you know route the query to the right right you know right data center and so on and so forth right so i think there is just a lot of things you need to kind of take into consideration before you decide where you want to send that query and so yes that orchestration layer i mean nvidia announced dynamo at gtc this year that is that is again an orchestration layer for inference queries right so yes that is going to be very important because again going back to the fact that inference is not a one-size-fits-all not every query is the same so you need this orchestration layer that is going to be able to efficiently route queries to the right hardware.
35:45Right. And we are seeing a lot more of that. And that is going to be where influencing hardware is more heterogeneous. It is not just all GPUs. Right. And that is, that is something that gradually I think will become clear to a lot of people, but it's pretty clear to us. Does it matter which model you're using and, and And when you put the weights on the chip or on the PCI card or however you say it, that's a decision that you're making or that's a decision that the customer is making? It's typically done in conjunction with the customer, right? So the customer is trying to solve for a certain problem.
36:35And they work with us to figure out what's the best way to get there, right? So there's again different vectors, right? There's things like, you know, the model size, there is a model quantization, there is a software stack, which is okay. You know, my model is highly customized, right? You know, my own custom kernels to, for that particular model, how do I bring those kernels to your hardware, right? Oh, I want to focus a lot on latency. Oh, I want to focus a lot on cost, right? I mean, so there are all these different things that going to consult cannot be done in isolation uh it is a highly highly uh uh you know it's a typically a very deep engagement with with a customer to kind of get these models deployed now sometimes you know you have model like llama three you know call it 70 billion or 8 billion and you know you want to deploy those models uh it's a lot easier uh and that's why you're seeing you know companies getting into that business which is i'm just i'm just hosting llama models right uh you know deep seek you know these are all open source models and available and readily available and uh they all look a lot like each other i mean they're all you know variants of llama to a certain extent uh variants of transformer models to a certain extent so once you've done work on a few of them you can actually incrementally take that work to others right so you know it becomes a lot easier to host those um so you're seeing more of that that becoming very popular right uh you know we'll be doing that too of course we are not going to get into the hosting business that's not something we do but we'll be enabling a lot of folks to host models of that of that uh you know family uh that variant uh with our hardware yeah as more of these models come out how easy is it to to load a model onto the chip i mean if you've got uh you know chips with llama three on them and then suddenly deep seek comes out with something that blows llama three away which i guess it already has how do you then switch uh can you how do you do that yeah yeah so it depends again you know what that new model looks like like deep seek r1 distill which is a distilled version of deep seek the reasoning model right it uh it is essentially it looks a lot like llama right it is uh so if you've done work on llama 370 billion uh the incremental work needed to do you know bring deep seek r1 distill onto your hardware is it's not it's not like it it's going to take you you know you know six months right it's going to take you probably a few weeks actually doing that solution because the kernels look you know very similar the model architecture you know yes there are changes but it's it's you know for the most part you can leverage a lot of work you've done, right?
39:22So it is a few weeks worth of work and you can onboard a new model, right? So that's really what it comes down to is what does the model architecture look like, you know, things like how many layers in the model, you know, are there any new mathematical operators that the model introduced so that you have to write new kernels, right? What is, how big is the model, what is the scale out strategy, right? What is the scale up strategy? How we partition that model, any kind of, you know, parallelism techniques that you need to apply a lot of this work across the models we just talked about, you know, can carry forward, right?
40:00That means your incremental work is a lot less. So yes, it really depends on on from model to model. I don't think there's a generic rule that you can apply across every single model. And where do you see this going as models proliferate?
40:25Are chips going to proliferate as well? And is the inference market growing so fast that there's plenty of room for for everybody or is there going to be a shakeout at some point uh where i mean the hardware business is expensive and yeah you know i think the market is big right so the opportunity is massive i don't think there's been a creation of an opportunity like this one in our lifetimes right so the the inference computing opportunity, I mean, if you believe what Jensen said at the GTC, I mean, it's, you know, almost 10 times larger, there's 100x more compute needed, the opportunity itself is 10 times larger, right?
41:12So if the training opportunity was 100 billion, then the inferencing opportunity could be 1 trillion, right? I mean, let's, you know, let's scale that down a little bit. I mean, you know, even if it is a 200 to 300 billion dollar opportunity, right? That's massive. You know, that's massive. I mean, you're looking at the creation of a the entire semiconductor market is you know 500 billion so you're looking at the creation of another 300 billion dollar opportunity just in the next 10 years which which is going to support a lot of support a lot of different company and we talked about inference not being a one-size-fits-all so it's not like there's going to be just one solution that runs away with the price right um so yes there will be multiple players uh the question is it's not going to be as many as we have today right so there'll be there'll be uh some kind of consolidation that'll happen over the next you know call it five years maybe um but there's going to be multiple winners because first of all the market just the inherent nature of inference is such that you can't really build a single chip that covers all your needs um and that is going to allow other companies to break out and um and that would be kind of the healthy outcome anyway I think, you know, in the semiconductor industry, we've not had the creation of a yet another large semiconductor company in the last 20 plus years, right?
42:30Because the industry went through a lot of consolidation and, you know, this return on semiconductor investments was not very attractive. So not many companies got funded. So we've not had a very long gap where we have not had any company really break out. So I think this platform shift, this AI platform shift is going to drive that. And this is going to lead to the creation of the next NVIDIA, the next Broadcom, the next Marvel, call it. And I would expect that to happen, right? And it would be good for the ecosystem. Yeah. Are you, you know, TSMC is building this fab in Arizona. I think they're going to start with four nanometer and eventually get to two.
43:14are you going to be manufacturing in Arizona when they're up and running? Probably. Again, the way TSMC is also going to have similar fabs in Taiwan. They're going to have a 4 nanometer in Taiwan. The pricing structure could be different just given the cost of labor between Taiwan and the US. We'll take a look at it. Right now, we run everything out of Taiwan. which is we are at six nanometers at TSMC, and that's kind of exclusively run out of Taiwan. Our product roadmap goes to four and then eventually to two, and we'll have the option of running out of Arizona. We don't have to take that up.
Read the full transcript
43:56We can always still keep running out of Taiwan, but we'll have to just see what makes most sense from a business perspective at that point in time and then make a decision. Yeah, and your market is, I would guess, for the time being primarily in the United States, but I ran into Andrew Feldman and Rodrigo Liang in Riyadh, Saudi Arabia last fall, and they were both there, obviously. We were at a conference, but they were talking to Aramco and some of the big inference companies or companies looking for inference. How much of your market is outside the United States and where would it be concentrated?
44:50Yeah. So yes, we are working with customers outside the US. We have customers in the Middle East. We have customers in the APAC region. we have customers in Japan right so we are working with customers outside of the US we you know we do believe that our primary market will be in the US because that is really where you know most of the high volume deployments are happening and we talked about our go-to-market model it's a little different from the companies you mentioned right we are focused more on embedded uh embedding our hardware into a customer's environment right we and so it's a volume driven business for us and um for that you know by that token uh we are looking for you know very large fleets of uh server ai server hardware that can be augmented with our solution and that typically today at least given that china is off limits with all the trade you know the trade tension and trade wars uh for now we can't ship our product into into china so that means the u.s market is going to be a huge focus the markets i talked about apac in the middle east still don't have the ability to house our product they're looking for total solutions so our our dependence on those markets is going to be limited for some time why or maybe i'm just unaware of it but i mean not including china why are there not chip companies sprouting up in in india for example or are there that are competing with you guys there they are they've begun to spring up uh i mean there's a there's a couple of reasons One is, I think just the perception is that the software ecosystem in India is a lot stronger than the hardware ecosystem.
46:53And that is true, right? There's just a lot more software talent in India than there is hardware talent. But there is good hardware talent, right? But the startup ecosystem around hardware in India is not that robust yet, right? So there is not enough capital to feed these companies. companies, the markets are not in India. The markets are predominantly in the US and China or Japan or elsewhere. So it's hard for a company to establish itself and say, I'm going to build product for international markets and be successful. Because there is always a company that is already doing that in the US for the US markets or there is a company in China doing it for the China market, right?
47:44So then building a company in India to serve those markets when the startup ecosystem for hardware is already not that robust is making, it makes it very, very challenging. Yeah. You know, China has done a lot since the chip in Huawei has done a lot. And it was kind of the leading edge of the chip embargo. and China has a lot of talent it has a lot of hardware talent and they're working very hard at developing their own ultra what is it the UH I'll cut this out but the lithography machines needed to etch you know, sub-7 nanometer chips. Are they going to be a player in this market eventually, or are they already?
48:47You're talking about the EUV machines? Ultra high... Extreme ultraviolet. Extreme ultraviolet, yeah. And your question is, are they going to be a player building those machines? in in providing inference chips to the global market oh inference chips to the global market okay got it um to be seen to be seen i think um i mean there's just a lot of uh you know embargoes right now on china with you know getting access to the latest and greatest technology so it's hard for them to really keep pace with the us on on a certain class of technology and inferencing chips uh um so yeah it's difficult to say whether they will uh we you know they will be able to keep pace with uh innovation in the u.s um i think to me you know the u.s is currently on the forefront of of uh of you know just the latest and greatest techniques when it comes to chip design right chip manufacturing chip not necessarily manufacturing but chip design specifically and architecture and And we just have a very robust ecosystem of engineers and customers and partners who put it all together.
50:05And by the way, even before these embargoes came and the trade tension escalated, I think the U.S. always had a lead when it came to building cutting-edge chips. So it is independent, by the way, of what's happened in the last, call it eight years or 10 years. right um so uh and it's only it's only the gap is only widening now because because of the trade tension so so i would i would argue that the us you know will will be the one to produce uh bleeding edge cutting edge inferencing chips you know at least for the next five to ten years right for the global market and it'll be difficult for other markets to really really keep up yeah yeah okay well we're uh almost up to an hour is there anything i haven't asked that that you think listeners should know yeah i mean um you know dmatrix is a great place to come work so if anyone in your audience uh is looking to work on really bleeding edge uh you know inferencing chips uh where we kind of are solving some really hard uh you know technical problems so both software and hardware uh right you know you know compilers uh infrastructure software uh you know runtime systems software scale out scale up triplet technology triplet interconnects compute technology packaging a board design i mean uh come talk to us uh we'd love to we'd love to talk to you and we are growing the team and uh you know we are the only only company that has a purpose built generative ai inferencing solution built from the grounds of for inference it's not like we started the company doing something else and then pivoted into this opportunity we built this company for for generative ai inference right just given the timing of when we started it worked out that way so you know we we are super excited about the opportunity that we are growing into and this is uh you know your chance to come work on something super exciting this episode of eye on ai is sponsored by the divinity foundation divinity foundation is a swiss not-for-profit that is home to some of the world's leading cryptographers computer scientists, and experts in distributed computing.
52:37Their mission is to shift cloud computing toward a fully decentralized state by supporting the Internet Computer, also known as ICP. If you don't understand anything about the Internet Computer, you can see my episode with Dominic Williams. ICP's vision is that most of the world's software will be replaced by network resident software that's an evolution of traditional smart contracts. To achieve this vision, ICP is designed to make smart contracts as powerful as traditional software, all while remaining tamper-proof, unstoppable, transparent, and verifiable. Since DFINITY's launch in 2016, they stand as Switzerland's most extensive blockchain research and development initiative and have been awarded more than 500 research grants worldwide.
53:38DFINITY remains steadfast in its mission to drive the advancement of the decentralized internet. Network resident software can now be used to run AI models and RAG infrastructure, preventing them from becoming quote-unquote hot wallets from where data can be stolen and increasing resilience. Furthermore, the technology has been designed to allow AI models to spin up and modify running web applications and internet services solo by addressing several key challenges. If you're interested in reading more about the internet computer, visit internetcomputer.org. I also encourage you to listen to my episode with Dominic Williams in which he explains the internet computer.
From the publisher
This episode is sponsored by the DFINITY Foundation.
DFINITY Foundation's mission is to develop and contribute technology that enables the Internet Computer (ICP) blockchain and its ecosystem, aiming to shift cloud computing into a fully decentralized state.
Find out more at https://internetcomputer.org/
In this episode of Eye on AI, we sit down with Sid Sheth, CEO and Co-Founder of d-Matrix, to explore how his company is revolutionizing AI inference hardware and taking on industry giants like NVIDIA.
Sid shares his journey from building multi-billion-dollar businesses in semiconductors to founding d-Matrix—a startup focused on generative AI inference, chiplet-based architecture, and ultra-low latency AI acceleration.
We break down:
-
Why the future of AI lies in inference, not training
-
How d-Matrix’s Corsair PCIe accelerator outperforms NVIDIA's H200
-
The role of in-memory compute and high bandwidth memory in next-gen AI chips
-
How d-Matrix integrates seamlessly into hyperscaler and enterprise cloud environments
-
Why AI infrastructure is becoming heterogeneous and what that means for developers
-
The global outlook on inference chips—from the US to APAC and beyond
-
How Sid plans to build the next NVIDIA-level company from the ground up.
Whether you're building in AI infrastructure, investing in semiconductors, or just curious about the future of generative AI at scale, this episode is packed with value.
Stay Updated:
Craig Smith on X:https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Intro
(02:46) Introducing Sid Sheth
(05:27) Why He Started d-Matrix
(07:28) Lessons from Building a $2.5B Chip Business
(11:52) How d-Matrix Prototypes New Chips
(15:06) Working with Hyperscalers Like Google & Amazon
(17:27) What’s Inside the Corsair AI Accelerator
(21:12) How d-Matrix Beats NVIDIA on Chip Efficiency
(24:10) The Memory Bandwidth Advantage Explained
(26:27) Running Massive AI Models at High Speed
(30:20) Why Inference Isn’t One-Size-Fits-All
(32:40) The Future of AI Hardware
(36:28) Supporting Llama 3 and Other Open Models
(40:16) Is the Inference Market Big Enough?
(43:21) Why the US Is Still the Key Market
(46:39) Can India Compete in the AI Chip Race?
(49:09) Will China Catch Up on AI Hardware?




