In short
Eye On A.I. Podcast Episode Notes
Episode Title: #228 Rodrigo Liang: How SambaNova Systems Is Disrupting AI Inference Host: Craig S. Smith Guest: Rodrigo Liang, Co-founder and CEO of SambaNova Systems Sponsor: RapidSOS
---
Episode Overview
In this episode, Craig Smith interviews Rodrigo Liang, who discusses SambaNova Systems' innovative approach to AI inference technology. They explore the transition from AI training to inference, SambaNova's achievements in speed and efficiency, and the competitive landscape of AI hardware, particularly in relation to NVIDIA's dominance.
---
Key Concepts
- Introduction to SambaNova Systems
- Founded by Rodrigo Liang and two Stanford professors.
- Focuses on revolutionizing AI by creating scalable and power-efficient solutions.
- Originated from a background in high-performance chip design.
- Significance of AI Inference
- A shift from AI training to inference is underway.
- Inference is becoming crucial for real-time applications, emphasizing speed and efficiency.
- SambaNova has developed record-breaking inference models (Lama 405B and 70B) that excel in accuracy and performance.
- Technical Achievements
- SambaNova’s hardware can achieve:
- 132 tokens/second for the 405B model.
- 570 tokens/second for the 70B model.
- Operates on a single rack consuming less than 10 kilowatts of power.
- Focus on power efficiency unlocks opportunities for private and secure AI systems.
- Competitive Landscape
- NVIDIA's dominance in AI training is challenged in the inference space.
- Other competitors like Cerebras and Grok struggle against NVIDIA's established cloud presence and developer lock-in via CUDA.
- SambaNova offers an API inference service that enables developers to work with open-source models, facilitating easier access.
- Challenges and Strategies
- SambaNova addresses challenges by focusing on:
- Efficient scaling without requiring significant infrastructure changes.
- Modular deployment in existing data centers.
- Multi-tenancy capabilities, allowing concurrent use of multiple models on fewer racks.
Notable Discussions
- Power Efficiency: Liang emphasizes the need for power-efficient hardware to handle the increasing demands of AI inference.
- Market Positioning: SambaNova aims to redefine AI for enterprise adoption, providing alternatives to the entrenched solutions offered by NVIDIA and other major players.
- Emerging Trends: The conversation highlights trends towards open-source models, customization, and the need for real-time processing in AI applications.
- Future of AI Hardware: Anticipates a landscape where fewer but more capable players emerge, driven by the need for better performance and power management.
---
Conclusion
Rodrigo Liang's insights into SambaNova's strategy and achievements provide a compelling view of the evolving AI landscape, particularly in the realm of inference technology. The focus on speed, power efficiency, and innovative deployment strategies positions SambaNova as a significant player in redefining how enterprises leverage AI.
---
Stay Updated
- Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
---
This summary highlights the main themes and insights from the podcast episode, providing a comprehensive overview for readers interested in the developments in AI technology and SambaNova's role in it.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00If you look at the strong grip that NVIDIA had on training, it's not like that. It's coming to an end with inference. Right. Not to say they won't be a player, but when the power and the performance and the cost are so far off from what others are able to provide. Now you're starting to see a large number of developers coming into other offerings and get their APIs from there. You can come into Google Cloud, you can come into AWS, you can come in from wherever you want because it's just an API. It's just an API. And so we're able to host those APIs wherever the customer wants. In fact, here we're in Saudi Arabia today and we power Saudi Aramco.
0:38Hi, this episode is sponsored by RapidSOS. Terabytes and petabytes data is exploding in our daily lives. How many connected devices do you own? You ever wonder how all this data could be used in emergencies? By 2030, we'll have over 32 billion IoT devices worldwide, double what we have now. Despite this data abundance, there's a critical safety gap. Vital information isn't reaching emergency responders in time. This gap results in reactive rather than proactive emergency response, putting lives and property at unnecessary risk. Many of us have had emergencies in our homes where we depend on EMS to respond quickly.
1:32Rapid SOS is closing this safety gap with their AI-powered intelligent safety platform. They connect life-saving data from over 540 million devices to more than 21 ,000 public safety agencies across six countries. Despite$200 billion in annual safety system investments, enterprises often lack timely, actionable information during emergencies. Rapid SOS is changing that. It's not just about avoiding problems. Safety investments are linked to increased customer satisfaction, employee retention, and long-term firm value. Close the safety gap and transform your emergency response with RapidSOS. Visit RapidSOS.com slash IonAI.
2:32That's IonAI, E-Y-E-O-N-A-I, all run together. Visit RapidSOS.com slash IonAI today to learn how AI-powered safety can protect your people and boost your bottom line. That's rapidSOS, R-A-P-I-D-S-O-S, dot com slash IonAI. Visit them today. Can you just introduce yourself and what your background is, your educational background, and how you started Salmonova? Yeah, I'm Rodrigo León, co-founder and CEO of Salmonova. I did my undergraduate and graduate work at Stanford in the high-performance chip design business for 30 years now, starting with PA Risk at HP, and then ran some microsystem Spark processors for 15 years, and started this in 2017 with two Stanford professors who are still there, really driving a grounds-up way of thinking about artificial intelligence of semiconductors and driving for power efficiency, for performance, and ultimately for scale.
3:43So I wanted to talk to you about the new inference system on a chip that beats Grok and Cerebris' new inference service. We don't have much time, so why don't you tell us what what that is and why inference is becoming the new field of competition. Yeah, really proud of it. Today, we're now Summon Over Cloud, and it's an inference service, API service, for developers to be able to engage and use the best open source models as a consumption model, so token services. And so we've seen other players do this now, and you're providing tokens for some of these models. And so in that game, it's all about accuracy and performance.
4:30And so today we announced that we're able to take the best model from Meta 405B, the LAMA 3.1 405B. We're actually influencing full precision at 132 tokens per second. In comparison, NVIDIA is taking, for the 405B, they're having to get down to 8-bit precision, which you lose accuracy in that. when you quantize down and still running at something close to about 30 to 40 tokens per second and so so you look at the comparison not in terms of just accuracy but performance it's just number one it's actually only one in the world because cerebris and brock cannot do it today they don't offer it today um oh and and then we also set a new world record in the 70b model full precision full precision um 16-bit uh running again um and uh on llama 70b we're doing 570 tokens per second and again number one in the world and so really excited about that but probably what's most uh exciting to me greg is we're doing that in a single rack a single rack that's running less than 10 kilowatts wow yeah right i've said all along that AI is going to production.
5:52The cost of inferencing will be 10 times more than what we spend on training. I'll say 10 times more, right? Production AI is going to be all about inferencing and fine-tuning your private data into those models. And so you need to find a way to inference those efficiently. Efficient means speed and performance. It means accuracy, but it also means power. And we think that today, most of the other services, whether it's the startups or the big companies, they try to not talk about power. Right, right. Right? Yeah. Let me ask you, Cerebris has had trouble getting traction in the market. Grok has had some trouble getting traction in the market because NVIDIA has a lock on the cloud.
6:42Yeah. and they have a lock on developers with CUDA. This API inference service sidesteps that because people can build. This is a fast inference, much faster. They can build on top of it. But how do you get the clouds to adopt SambaNova and switch from NVIDIA. Yeah, that's a beautiful thing today. If you look at the strong grip that NVIDIA had on training, it's not like that. It's coming to an end with inference. Not to say there won't be a player, but when the power and the performance and the cost are so far off from what others are able to provide. Now you're starting to see a large number of developers coming into other offerings and run them, you know, and get their APIs from there.
7:45Now, you can come in through Google Cloud. You can come in through AWS. You can come in from wherever you want because it's just an API. Yeah. Right? It looks like OpenAI is just an API. And so we're able to host those APIs wherever the customer wants. In fact, here we're in Saudi Arabia today and we power Saudi Aramco. Oh, is that right? We power their internal MetaBrain, right? which is we take 90 years of data, trained it into a private and secure model, and we inference it completely privately. Deployed in their own data centers here in Saudi Arabia, completely secure, and they're inferencing their data in these models for their employees.
8:26And so this is the use model that we think that is going to be extremely fast for people to adopt because you no longer need to worry about CUDA and what you train them with. You can take any open source model, dump it in. actually. What about production and capacity? Because that's, you know, I wrote an article on Cerebrus. One of the pushbacks I got is, yeah, but does Cerebrus have the production to meet demand? Yeah. And are, you know, how much capacity do they have to serve APIs through the cloud? Yeah, that's a really valid question. And one of the things I tell people, but take a Lama 70B, right?
9:07The most important question people should ask is how many sockets does it take to do that? Or how many wafers does it take to do that? How many racks does it take to do that? And why is that important? You're running one model. If that one model takes hundreds of chips or many, many very expensive wafers to run, or in NVIDIA's case, still saying a lot of power to run, that doesn't scale, right? And so for some of them, we're running these models in a single rat at less than 10 kilowatts. And so what that allows us to do in a very modular way is just deploy more racks very efficiently. And going into existing data centers, you don't need to go build a liquid, cool, brand new data center to power these systems because it's too expensive.
9:52Yeah, which is what NVIDIA is doing, right? They're starting to build their own data centers for inference as well. Exactly. You have to because the chips are so power hungry and they need all this new liquid cooling. And that cooling has to be done a special way. Some of them would decide we don't want to do any of that. What we want to do is delivering performance, delivering availability, all in the existing infrastructure that there is. Data centers that have 10 kilowatts are able to actually push this rack in and achieve 132 tokens per second on 4 or 5B. Single rack. Where's your production?
10:27Where are you making the chips? The chips have been out of TSMC in Taiwan. Yeah. Are you having trouble getting, I mean, everyone is looking for it. Yeah. I mean, this is one of the really great things about the RDU, the reconfigurable data flow chip that's ours. We're able, because it's data flow, we're able to collapse the total number of chips by an order of magnitude. Yeah. So what you would have thought would take, say, a hundred chips of nvidia to do we can do with 10 ish chips and so so i think that's kind of allowed us to significantly reduce the demand of chips which then has all these other supply chain benefits yeah but uh but as a well funded startup we raised 1.1 billion in venture capital in the first three years of the company yeah how long have you been around we've been uh six almost seven years yeah our first three years were able to raise a lot of capital from some of the top investors up you know this is google and intel and you know black rock and soft bank and tomasek you know some of the you know uh great investors the saudis i'm guessing well you know so so yeah yeah we're the ones that we've announced are are the ones that are public but uh uh but these are the the investors that have allowed us to actually be able to get these chips way in advance yeah get the chips early so that we have them we can supply them just supply them quickly but as you're seeing this drive for demand growing at this rate, there's always gonna start having a lead time.
11:55We're starting, over the last couple of months, we've started kind of really actively managing allocation because it's starting to get to the point where we have to do that despite the fact that we pre-bought so much material. Yeah, yeah. So you have a good pipeline of chips coming through. You're negotiating presumably with the cloud. I mean, one thing, you know, I'm a journalist. I wasn't aware of Sambanova. Why is that? I know, I know. And are you focusing on a particular kind of inference, or is this wide open? Yeah, our history has been that we go into enterprises and we go on-prem. Today, we're the most deployed AI chip startup in the U.S.
12:39government. People don't know that. Yeah. Right? We go into all the major labs. We go into a bunch of different places. we're in banks, we're in three continents today. We'll go into enterprises where people have a lot of private and secure data. They don't want to disclose it to somebody else. 83 % of enterprise data sits on-prem today. Right? They have all that data. They don't know what it says. So San Manovo, for the last several years, we were able to bring our racks in because we're so efficient. Instead of having to create an entire data center to read your models, I can just roll in a couple of racks.
13:13You've got a model. I can read that data and now you have your own private GPT. Right, right. And so I think that's kind of for us, is just something that we're able to do very efficiently and get into engagements with enterprises very quickly. Now, as we've moved since the world's going to production, we see this tremendous energy around inference. And frankly, our technology is really good for inference, as we've now shown with 405B being world record, 70B being world record, all 16-bit. And so we're able to drive these results in a way that others can't and do it in a fraction of the power that everybody's doing.
13:51So that's what drove us to do this cloud now, because now we can take what we're doing at the customer's premise. We're just putting them into cloud environments. And these aren't our data centers. These are our partners' data centers. We run a service through them. Right. And so putting it through our partners' data centers so that then we can offer these services to a broad range of people that perhaps they don't want to do it on-prem by themselves. Yeah. And this move toward, well, first of all, the higher speed gives you a much larger context window. Is that right? Well, there are three things that you get from Salmonova.
14:24One is speed, right? The speed that you're getting is, you can use it for a couple of different things. One, the world wants to go real time. Right. A lot of AI wants to be real time. And if you don't have the speed, it's too slow. We're in an impatient world. Right. Right. And so after a couple of seconds, we think the machine hot. Right. And so so real time, you need the speed. Most of the slower models just won't work. So that's one thing that we do with someone over that we're able to give you that speed. Second, because we're driving the 405, 400 billion all the way to a trillion models. We're giving you the accuracy.
14:58The bigger the models, the more accurate. Unlike what others says, the bigger the chip is better. No, the bigger the model is better. The bigger the chip just brings more power. If the big chip can't run the big model, then you're not getting the benefit. And so we want the bigger models. 45B from Meta is the best open source. Yeah, and how is that architected? Are you putting the weights directly on the chip? Is it in memory just close? Yeah, so Sambanova has a very sophisticated memory hierarchy with a lot of SRAMs on chip. But we have HBM and we have DDR. So in a single eight socket system, we have 12 terabytes of DDR directly attached to some of those boxes, some of those chips, which allows us to do all sorts of different things.
15:46Right. And so we're able to, in a very small footprint, host a 400 billion parameter model and run it really efficiency. But the other thing that we can do now is we can also concurrently host hundreds of your AP model. Yeah. in a virtual manner sure right so now i just brought virtualization into world of ai which nvidia and all the other startups don't do right right and you say why is that important because one day you will have your own llama checkpoint he will have his checkpoint everybody wants their own and today in the in the existing legacy world your model one rack of hardware his model another rack of hardware we can multi-tenant hundreds of these checks in one system and dynamically switch virtually in the way that VMware made an entire business for two, three decades, right?
16:37And so that's what that memory has allowed us to do. I mean, there's a lot more than just the memory. The software is very intelligent to do that, but we give you that efficiency because we give you more dependency, we give you speed, and we give you the actual raw power of the rack. How fast is the API? Because the inference can be fast, but you still need the bandwidth out to whoever's using the API. And some of what we do here is we, you know, I mean, the model, our model, the AP model, 0.09 second time to first token, that's number one in the world. Yeah. That's number one in the world. But like you said, it depends on where you are.
17:17So why are we able to actually get our service in this way is we're dropping these racks across the world because I don't need to build gigawatt data centers. Right. We can go into existing data centers and turn that into a micropod for service. We can put eight racks, 80 racks, 800 racks. Because it's very modular and it's all a single rack and a single rack runs the full model. And so it gives us the flexibility of bringing AI locally. And we've already done it because we do it for people on-prem. And just on the on-prem point, right now you're doing the open source, not as open source, presumably other open source models, Command R or Mistral.
18:02We have a broad range of models that we support, all open source from Hugging Face. Our top model is called Samba One. Samba One is a composition of experts. And what that is is basically a virtual platform that allows you to bring many checkpoints of open source models and then the model route to it based on your prompt. Right. Right. And that's the architecture of how we achieve multi-tenancy. There are 91 different models today offered on Samba One. OpenAI is still the king. They're proprietary. I mean, there are other proprietary models that are powerful. What's it going to take to get OpenAI to use Sambanova instead of NVIDIA GPUs?
18:47Well, we've got to run it better than what they have with NVIDIA, for sure. And as these models are getting bigger, Sambanova's technology advantage on the big models start getting bigger and bigger and bigger. And so this is what we play, and we play really, really well. We can run these large models better than anybody else. and even the modest models were already showing number one, but the bigger the model, the more our architecture shines. Why is OpenAI so married to NVIDIA, do you think? I don't know. I don't know. I think it's time. I think a lot of the decisions there were probably tied to legacy decisions that they've made, investments you've already made on GPUs because GPUs just were there before us.
19:27Sure. But we're at this point in time now where the scale, the next phase, the scale of hardware investment you have to make is an order magnitude larger. Where is it going with inference? Because there was so much focus on training, foundation models, but now the focus has shifted to inference. You mentioned someday you're each going to have your own LAMA checkpoint. Where do you see that going? What kinds of products do you think this is going to be? Well, three things. I think the world's going to be open source. So you're going open source and that's going to continue that way. I think it's going to be more and more.
20:06The models are really good. Good now too. I think it's going to be bigger and bigger over time. Why? Because people want accuracy. They want multimodality, right? They want a bunch of different things that pushes the model to be bigger. And so that's, you have to handle that. And three people want customizations. People want to train their private data into it. They want to understand me. I want to create my own checkpoint. I want to create my own value. And so those three things are driving us to do the things that we do to run that future outcome much, much better, which is it's got to be able to handle fast, big, and it's got to be able to handle lots and lots and lots of concurrent checkpoints.
20:46And do you think that there are going to be a few big players? You know, it's been NVIDIA for a long time. Cerebris is making noise. You guys have not popped up. You've been around, but you have this new service. Are we going to end up with 10 different options? I don't know that you're going to have 10. It's a very capital-intensive venture. So I think you have to be able to do it and do it consistently. We're in our fourth generation hardware already, right? in the market, and so you have to do consistently. But the world wants choice. There wasn't one's choice, and the best product is going to attract people because power is a real problem in this world.
21:27And if you don't have power, you don't have capacity. Hi, this episode is sponsored by RapidSOS. Terabytes and petabytes data is exploding in our daily lives. How many connected devices do you own? Do you ever wonder how all this data could be used in emergencies? By 2030, we'll have over 32 billion IoT devices worldwide, double what we have now. Despite this data abundance, there's a critical safety gap. Vital information isn't reaching emergency responders in time. This gap results in reactive rather than proactive emergency response, putting lives and property at unnecessary risk. Many of us have had emergencies in our homes where we depend on EMS to respond quickly.
22:23Rapid SOS is closing this safety gap with their AI-powered intelligent safety platform. They connect life-saving data from over 540 million devices to more than 21 ,000 public safety agencies across six countries. Despite$200 billion in annual safety system investments, enterprises often lack timely, actionable information during emergencies. Rapid SOS is changing that. It's not just about avoiding problems. Safety investments are linked to increased customer satisfaction, employee retention, and long-term firm value. Close the safety gap and transform your emergency response with RapidSOS. Visit rapidsos.com slash ionai.
23:23That's I-ON-A-I-E-Y-E-O-N-A-I all run together. Visit rapidsos.com slash IONAI today to learn how AI-powered safety can protect your people and boost your bottom line. That's rapidsos, R-A-P-I-D-S-O-S dot com slash IONAI. Visit them today.
From the publisher
This episode is sponsored by RapidSOS. Close the safety gap and transform your emergency response with RapidSOS.
Visit https://rapidsos.com/eyeonai/ today to learn how AI-powered safety can protect your people and boost your bottom line.
In this episode of the Eye on AI podcast, we explore the world of AI inference technology with Rodrigo Liang, co-founder and CEO of SambaNova Systems.
Rodrigo shares his journey from high-performance chip design to building SambaNova, a company revolutionizing how enterprises leverage AI through scalable, power-efficient solutions. We dive into SambaNova’s groundbreaking achievements, including their record-breaking inference models, the Lama 405B and 70B, which deliver unparalleled speed and accuracy—all on a single rack consuming less than 10 kilowatts of power.
Throughout the conversation, Rodrigo highlights the seismic shift from AI training to inference, explaining why production AI is now about speed, efficiency, and real-time applications. He details SambaNova’s approach to open-source models, modular deployment, and multi-tenancy, enabling enterprises to scale AI without costly infrastructure overhauls.
We also discuss the competitive landscape of AI hardware, the challenges of NVIDIA’s dominance, and how SambaNova is paving the way for a new era of AI innovation. Rodrigo explains the critical importance of power efficiency and how SambaNova’s technology is unlocking opportunities for enterprises to deploy private, secure AI systems on-premises and in the cloud.
Discover how SambaNova is redefining AI for enterprise adoption, enabling real-time AI, and setting new standards in efficiency and scalability.
Don’t forget to like, subscribe, and hit the notification bell to stay updated on the latest breakthroughs in AI, technology, and enterprise innovation!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI




