In short
SambaNova CEO Rodrigo Liang discusses scaling AI inference efficiently, a $1B+ fundraise at an $11B valuation, and why “premium inference” needs low power, low latency, and fast deployment (often in single air-cooled racks) for global and edge use.
Guest backgrounds
Rodrigo Liang is SambaNova’s CEO; he has 32 years in high-performance computing hardware (“high performance trips”) and has led the company since its 2017 founding, focusing on inference efficiency.
Key claims
- Interest in semiconductors is at an all-time high due to AI’s scale (millions of daily users via Anthropic/OpenAI/Gemini).
- Inference is now the bottleneck; training is less the problem than deploying at scale without exploding power, space, and latency.
- SambaNova’s approach avoids liquid cooling and reduces minimum deployment to one rack using standard 19-inch hardware, air cooling, Kubernetes/Red Hat/Linux, and Ethernet.
- “Premium inference” is defined by accuracy (bigger models) and speed (running full precision without quantization).
Notable examples
- SN40: outperforming a ~130–140 kW NVIDIA GPU rack with a ~10 kW air-cooled SN40 rack; running a trillion-parameter model in one rack vs dozens elsewhere.
- Edge/modular deployments: 10–20 rack clusters in shipping-container data centers; Armada is a partner.
- Routing: inference providers already route requests via standard APIs; SambaNova racks can be inserted to improve tokens-per-cost and margins.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to SambaNova's Success
0:00 to 1:00
Learn about SambaNova's recent billion dollar fundraise and its implications.
“We just did the first close of a billion dollar fund raise at an 11 billion valuation.”
SambaNova's Momentum and Investor Interest
1:19 to 2:08
Discover the details of SambaNova's funding round and investor enthusiasm.
“I think this will come out about a week after the news drops, but we'll still make a clip on that.”
The Importance of Semiconductors in AI
2:08 to 3:01
Explore why chips are crucial for AI advancements and SambaNova's role.
“I'm sure this was a very hyped-up round in some way or another.”
Evolution of SambaNova's Products for Inference
3:01 to 3:52
Learn about SambaNova's product evolution and its focus on inference technologies.
“Inference has really taken the stage and it's been the next evolution of computing and where everything's going with AI.”
Scaling Challenges in AI Deployment
3:52 to 4:50
Understand the scaling challenges related to AI inference and solutions.
“It was people trying to test the model that they trained.”
The SN40 Rack and Performance Advantages
4:50 to 6:41
Discover the performance advantages of the SN40 rack over traditional options.
“So we're on, we've taped out six chips in the last seven years, and we'll tape out seven for the next year.”
Changing Dynamics in Rack Composition for AI
6:41 to 9:10
Explore how rack composition has changed from training to inference in AI.
“If you look at kind of for training, one of the challenges that you have is you have to aggregate all of these racks because you need thousands of GPUs together and they have to work in sync.”
Deployment Efficiency in Data Centers
9:10 to 10:00
Learn how SambaNova improves deployment efficiency in data centers.
“Here in Paris, downtown Paris, you can find an existing data center and deploy five racks, 10 racks there for ultra-low latency where the users are, right?”
The Future of Data Centers and Latency Considerations
10:00 to 11:33
Understand future data center trends and the importance of low latency.
“maybe as long as 18 months to build a gigawatt data center, to bring liquid cooling in, to bring all this new power.”
Defining Premium Inference in AI
11:33 to 14:00
Learn what constitutes premium inference and its significance in AI applications.
“They're mid-size, and it's going to be even more important as you go into this agentic world because, you know, in the world of agents, you're not dealing with a single model and a single prompt, right?”
Show all 34 chapters
The Importance of Premium Inference
14:00 to 17:29
Learn about the dimensions of premium inference and model accuracy.
“And thinking banking, healthcare, like there are many use cases where your latency is really, really important.”
The Importance of Premium Inference
17:33 to 18:28
Learn about the dimensions of premium inference and model accuracy.
“And I refuse to spend my time on work that shouldn't exist.”
Market Bifurcation in AI Speed
18:28 to 20:43
Explore how speed in AI services impacts market dynamics and customer preferences.
“I know that token costs and everything is a really big topic right now, But for consumers, we do expect speed.”
Global Access to AI Technology
20:43 to 21:44
Discuss the challenges and opportunities in providing global access to AI.
“There's also a couple other macro themes that are happening that are just going to explode data and usage across the world.”
Deploying Edge Data Centers
21:44 to 24:18
Understand the importance of edge computing and modular data centers.
“Or you might be in a place where you don't have that and you need something significantly smaller.”
Collaborative Competition in AI Hardware
24:18 to 27:49
Learn about the dynamics of collaboration and competition among AI hardware providers.
“Because, again, access to them will be quite varied, right?”
Routing Performance in Data Centers
27:49 to 28:06
Discover how routing algorithms enhance performance in data centers.
“And how are they measuring the routing performance?”
Routing Models in HPC
28:06 to 30:02
Learn how models are routed in high-performance computing environments.
“when a request comes in, you are already routing those models to certain racks.”
Measuring Revenue in AI Infrastructure
30:02 to 31:03
Discover how service providers measure revenue in AI infrastructure.
“And this is the way we see most providers measuring.”
Standard Definitions in AI
31:03 to 32:16
Explore the importance of standard definitions in AI and cloud computing.
“You're like, well, it's very clear with hardware.”
Land Grab in AI Services
32:16 to 35:01
Analyze the competitive landscape in AI service provisioning and the race for scale.
“What are the biggest bottlenecks that you're seeing and what are the biggest challenges there?”
Differentiation in AI Service Providers
35:01 to 37:06
Understand how AI service providers differentiate themselves in a crowded market.
“And so now that everybody's kind of building out data centers, so you have to use that capital really, really efficiently and secure users, secure companies as quickly as you can.”
Collaborative Cloud Services
37:06 to 40:09
Learn about the role of partnerships in delivering cloud services efficiently.
“Because today, if you look at NeoCloud, What's Neocloud A and Neocloud B?”
Sovereignty in AI and Data
41:03 to 42:06
Discuss the importance of data sovereignty and national models in AI.
“the other week, really talking about sovereignty and owning your data, your products, everything in the pipeline, Berkeley.”
National Models for AI Development
42:06 to 44:28
Discover how countries are developing their own AI models to protect data privacy.
“And so that's what people are thinking about.”
The Shift Back to On-Prem Infrastructure
44:28 to 46:46
Learn about the trend of companies returning to on-prem solutions for data security and privacy.
“And, you know, we always have these like, you know, fun cyclical patterns and behaviors of, okay, like let's unbundle everything.”
Transforming Business with AI
46:46 to 49:25
Explore how businesses are using AI to differentiate their services and create new revenue streams.
“I'm going to do some paperwork, some operational things.”
AI and the Future of Job Services
49:25 to 53:24
Understand how AI can change job roles and enhance service offerings in the market.
“because you don't know what we don't know.”
Scaling in the New AI Landscape
53:24 to 56:00
Examine the challenges and opportunities for businesses to scale effectively in the AI era.
“Starting with, I hope I don't have to do taxes.”
The Journey of SambaNova's Growth
56:00 to 56:21
Discover how Rodrigo Liang scaled SambaNova to an $11 billion valuation.
“It's hard to catch up in a race where everybody's going really fast.”
The Journey of SambaNova's Growth
56:25 to 56:52
Discover how Rodrigo Liang scaled SambaNova to an $11 billion valuation.
Resilience and Team Dynamics in Business
56:52 to 57:36
Learn about the importance of resilience and teamwork in building a successful business.
“You can't control what the politics do, right?”
Introduction of the Latest Salmono Chip
57:36 to 58:05
Rodrigo introduces the new SM50 chip and its potential impact on technology.
“We surrounded ourselves with great people that are incredibly hardworking, but more than that, incredibly resilient, right?”
Transforming Inference at Scale
58:05 to 58:53
Explore how the SM50 chip will revolutionize deployment of inference at scale.
“It's something we announced earlier this year, but we're now actually producing some really incredible results on this.”
Transcript
Automatic transcript. May contain errors.0:00Rodrigo Liang:We just did the first close of a billion dollar fund raise at an 11 billion valuation. I've been in this industry for 32 years, building high performance trips for a long time. I've never seen the interest in semiconductors higher. Now what you're seeing at scale with Anthropic and with OpenAI and with Gemini, you've got millions and millions of people using it every day. We released SN40 a couple years ago. It became incredibly popular because instead of 130, 140 kilowatt rack of NVIDIA GPU, We were outperforming it with a 10 kilowatt SM40 rack. We could take a trillion parameter model and run it in a single rack, where we take dozens of racks of other people's equipment to run the same model.
0:38Rodrigo Liang:We're$2.5 billion raised in the history of the company, and there aren't really that many companies that have raised into the multiple billions. It's all about scaling. It's all about who can get to scale faster.
0:59Rodrigo, welcome to Sorcery.
1:01Rodrigo Liang:Thanks for having me. Well, we're here for context. We're in Paris right now for the RAISE Summit. And right now we're sitting right in front of where the conference is. I don't really actually know what this park is called, but it's next to the Louvre. Yeah. I don't know if you know. Yeah. You've been here a bunch, right? I've been here, but I'm not sure if I know exactly the name of the park. We're right in front of the Louvre. Well, you have some big news. I think this will come out about a week after the news drops, but we'll still make a clip on that. So what is the big news? Well, we're super excited.
1:33Rodrigo Liang:We just did the first close of a billion dollar fundraise at an 11 billion valuation. This is a great show of momentum for the company and great show of support. The round was led by General Atlantic with a number of incredible investors that came in. Seligman Ventures, T. Rowe Price, Capital Group. which these are all significant American investors that are coming in, that shows that the company's got momentum, we're driving towards the scale, and a significant amount of capital infusion to help us do that. All the energy right now is going into semiconductors. I'm sure this was a very hyped-up round in some way or another.
2:12Maybe it's been faster than others. What was the process like for you?
2:16Rodrigo Liang:I've been in this industry for 32 years, building high-performance trips for a long time. I've never seen the interest in semiconductors higher. And I think it's a realization that chips at the center of this transformation, if you look at what AI is doing in the world and the build-offs of the data centers, you can't do it without chips that run and run efficiently. And so with Salmanova, we're coming in and providing technology that is able to take it to scale, take it to a level of inference scaling that's just really not that practical to achieve just with traditional GPUs. And so I think the world sees that and the excitement is coming in from some of the top investors in the world.
3:00So where we are at today is inference. Inference has really taken the stage and it's been the next evolution of computing and where everything's going with AI. So for people that don't know SambaNova, can you walk through the products and how you've evolved them for inference?
3:17Rodrigo Liang:Yeah, I mean, look, with AI, you've got a sophisticated audience, so they know. With AI, there was always a training in the inference. There's no point of training a model if you aren't going to inference it, if you're not going to use it. And so the example I use with people is you don't go and invent the search algorithm if you're never going to do search. right and so saying we're not going to train a model if you're not going to use it and now we're in the phase of using these models we've always used them we've always inferenced them but it was still research to train models better and better and so we when we started the company in 2017 we're very focused on how do we actually lower the cost of training right and back at the time we're training models for image recognition can we tell the difference between dogs and cats and you know can we recognize voices can we make voices you know we're doing all that research But in the end, inference wasn't really a problem yet because the number of people using it were very small.
4:15Rodrigo Liang:It was people trying to test the model that they trained. Now what you're seeing at scale with Anthropik and with OpenAI and with Gemini, and you're at scale, you've got millions and millions of people using it every day. And so now you have the problem that Sumitava was originally focused on, which is around efficiency. How do you actually deploy at scale so the whole planet can use it? without burning up the planet, without kind of running out of data center space, without blowing up your infrastructure cost. Because at scale, the number of chips deployed for inferencing will be orders of magnitude greater than whatever you're doing for training.
4:52And so walk through some of your chips.
4:53Rodrigo Liang:Yeah, so we started a company. So we're on, we've taped out six chips in the last seven years, and we'll tape out seven for the next year. We're really excited about Generation 5 that's shipping later this year. When we started with SN10 and SN20, these are the early chips, they were really focused on training. They're focused on can we train those models faster with fewer chips? As you know, some of the largest AI labs are deploying thousands of chips just to train one model. And they were trained for months and months and months. And so we're focused on training. As the world moved along, it became very clear that the bigger challenge, the bigger problem was inferencing.
5:33Rodrigo Liang:by that once you were at scale, how do you actually deploy those same models you trained on all of these different data centers for people to use? And now you can say, I need a gigawatt data center somewhere in West Texas. That's one way to do it. But how do you serve all of the countries, all of the planets, all the different users that are worldwide? And so you need to think about power. You need to think about data center. You need to think about latency. And these are all things that we started taking on. And so by the time we released SN40 a couple years ago, it became incredibly popular.
6:11Rodrigo Liang:Because instead of a 130, 140 kilowatt rack of NVIDIA GPU, we were outperforming it with a 10 kilowatt SN40 rack. But in a 10 kilowatt SN40 rack, now suddenly, and it was air-cooled. You didn't need liquid cooling upgrades. suddenly every data center that's around the world that you're using traditional CPU, traditional storage for, you could just roll in the summit of a rack, air-cooled, and you got state-of-the-art inferencing faster than on an NVIDIA GPU. And so that became a really, really popular way to deploy, especially if you don't want to put hundreds of racks of NVIDIA GPUs, we would collapse the footprint because we could take a trillion parameter model and run it in a single rack, where it would take dozens of racks of other people's equipment to run the same model.
6:57Rodrigo Liang:And so that's kind of what Sumnova is really focused and known for, is really driving premium inference, the largest models that are very low cost, very low power, and delivering ultra high performance, which becomes really, really valuable if you're actually starting to deploy these incredible models at scale. How has rack composition changed? If you look at kind of for training, one of the challenges that you have is you have to aggregate all of these racks because you need thousands of GPUs together and they have to work in sync. As you're training each loop that you think, you're working in sync.
7:36Rodrigo Liang:And the problem there was that if any one fails, the whole cluster fails. right and so you have to put these checkpoints you have to kind of stop you know uh every so often to make sure that you um uh you store the progress you made up to that point in case the next cycle fails right and so there's all of this work and then you need really and you talk to tech tech folks all the time you really you need really really high performance networking to connect all those things through and so your networking equipment becomes really really expensive you need a lot of memory you need a lot of software uh coordination and things like that As you go into inference, the beauty of inference is it's scale out.
8:16Rodrigo Liang:And so basically you're adding racks as your users grow. And so with Samanova, that minimum quantum is down to one rack. Where if you have other service providers, you just run, say, a DeepSeq model, which is now one and a half trillion parameters. Just to run that, the minimum for some of the other providers might be 10 to 20 racks. And so you're starting to think about, okay, well, if I want to be an inference provider and I want to run these large models, my minimum cluster just to start serving is 20 racks and the cost stroke is high, the power needs are high. Some of it, we can actually reduce the minimum down to a single rack.
8:58Rodrigo Liang:And so now you just grow as your user base grows. Significantly more efficient to deploy, significantly more flexible if you're going into environments where you don't have gigawatt, you can go into your data centers. Here in Paris, downtown Paris, you can find an existing data center and deploy five racks, 10 racks there for ultra-low latency where the users are, right? And so this is kind of what's changed in the racks is that we're very focused on bringing it to broad-based deployment, standard everything, standard 19-inch rack, standard air cooling, no complicated liquid cooling retrofit in the data center.
9:37Rodrigo Liang:We're using standard Kubernetes, standard Red Hat Linux, standard Ethernet at the top for networking. We don't have to use all this kind of really expensive networking equipment to gang these things together. And so that allows people to go in to existing data centers, roll this thing in, pull out the old gear, and you're up and running with new services, which otherwise might take you nine months to a year, maybe as long as 18 months to build a gigawatt data center, to bring liquid cooling in, to bring all this new power. And sometimes you have to figure out how to secure new power and new energy and nuclear power plants.
10:11Rodrigo Liang:And all those things that people are talking about, it just takes a lot of money, a lot of time, but you don't have to do it if you use some of the equipment. That's pretty good. Yeah. And it's operating at the velocity of inference, which is people want to stand these things up quickly, right? It doesn't take months to inference because these models have been already trained. Whether that's an OSS model from OpenAI or Mistral more in France without all Mistral or you've got the DeepSeek and Minimax models, amazing models that are out there or the frontier models or the closed source models that people are offering.
10:49Rodrigo Liang:It doesn't matter. You deploy this technology, you can bring those models in immediately and you're up and running with the latest and greatest AI models in the world, and then you can upgrade them as you go along. So do you buy the headlines that new data centers should be 50 or$100 billion? Well, I think you're going to have some data centers like that because I think there's still going to be large-scale deployments, large-scale access that people want. And I think you're going to find that the world is going to be heterogeneous, that there's going to be those large data centers that near people go and secure a lot of capacity for some of the things that they want to do.
11:26Rodrigo Liang:And I think you're going to see this new wave of companies that are doing distributed data centers, right? So these data centers are mid-size, right? They're mid-size, and it's going to be even more important as you go into this agentic world because, you know, in the world of agents, you're not dealing with a single model and a single prompt, right? If I go to ChatGPT, you know, we're here in France, and what should I do if I have an extra day in France, right? You can talk to ChatGPT, it'll generate an itinerary for you. That's between me and the model, and it produces the result. In the world of agents, these agents are orchestrating within themselves without us.
12:04Rodrigo Liang:My problem starts in the beginning. Ten agents are all intercommunicating, and each of them taking some amount of time. So if you actually have a lead time or response time that's, say, two seconds, which for a one user to one model is not very long, right? For our eyes, it takes longer than two seconds to read the output, right? And so that's okay. In the world of agents, where you have, say, 20 agents orchestrating with each other, each of them takes two seconds to respond. Now I've got 40 seconds. The end user did an initial prompt here, like, please move some money from my bank account, my Bank of America bank account, over to PayPal, and then give me a report of all my expenses over the last six months.
12:46Rodrigo Liang:Okay, that's my prompt. Then the orchestration of security and balances and all of that has to happen. And the output. If each of those 20 agents took two seconds, that's 40 seconds. You've already given up on that prompt. So the response time by the user's expectation is, say, one to two seconds. Divide that by 20, it's less than a second per, it's 0.1 seconds per. And so your response time is going to be really, really important. And that's why latency matters. And so you're now seeing us coming in and saying, look, we're going to deploy the hardware where the users are in large metropolitan cities, right?
13:22Rodrigo Liang:Because that's where business is being run. And so latency is really important. So we're going to deploy that. Well, unfortunately, in those large metropolitan cities, you don't have those gigawatt data centers. There's no space for Manhattan in Paris, right? Where are we going to find space to drop a, what did you say? How much money did you say? 50 to 100 billion. I mean, where are you going to find even the space to build that? You're going to find some space in some place to build that, but in terms of ultra low latency in the big cities, you're going to have to find smaller quantums, smaller spaces that allow you to deploy what you need for the users there that require that really, really low latency.
14:05Rodrigo Liang:And thinking banking, healthcare, like there are many use cases where your latency is really, really important. you're just not going to want to wait. On the topic of inference, what is premium inference? The way we define premium is ultimately the highest value use cases. And there's two dimensions. It's basically on one dimension is the size of the model because it's about accuracy. Okay. And so years ago, we used to talk about hallucinations, right? You talk to Jackie Pt and come back and say, well, what was this? That was a fun time. I think, wait, like we should probably look back. Like hallucinations was a fun time.
14:40Rodrigo Liang:It does less of it now. It does less of it, right? These models are getting pretty good and does less of it. And yet accuracy is still incredibly important, right? And so why are these models going bigger and bigger? It's not that people want to spend the hundreds of millions of dollars to train, right? I mean, they're fighting for that bit of accuracy because you look at a model like Claude, Anthropic, right? You look at that model. why did it become the most popular code generation model? When software developers go and type, it generates really good code. And you take a different model, sometimes not as much.
15:19Rodrigo Liang:And this is where also the open source, the Minimax model that's out there as an open source model, became very popular, incredibly accurate when generating code. And so these models still are being valued significantly for the output they generate. Because if you can trust it to produce good output, you don't have to invest as much human energy to go double check it. If you have to go double check it, then certainly you're starting to invest more time. So on one dimension, premium is how accurate the model is. And today, that's proportional to the size of the model. right and so you have models that are very you know llama 8b for example is an 8 billion parameter model very small by 70b pretty small used to be 70b used to be the big it's tiny today right when you had chat gpt and gpt5 at 5 trillion right the new models are heading towards 10 trillion even the open source models are already one to two trillion parameter models and so now you're starting to see these models getting very big because people are looking for accuracy right And they want a model that handles a broad range of things, but handles it correctly.
16:29Rodrigo Liang:And so we're very focused on making sure that we handle the largest models well. And then the second dimension of premium is learn it fast. For the reasons I just described about agents, that you don't want to take a long time. We live in an impatient world. You and I, I mean, a few seconds we're starting to tap the phone, something happened. And so if the service behind it is agentic, you need to actually run really, really fast. There's not time for you to actually wait. And so that combination of running really big models, as you know, they don't run fast. Or you look at services like RockCerebus that run fast, you can only run the small models.
17:08Rodrigo Liang:So how do you find the ones that run really fast on the big models? And that's where some of the strength is, that we take the biggest models and run them in the original precision. We don't quantize. Quantizing is you chop half the weights off. So we don't chop the model down. We just run original precision, full precision, run faster than anybody else. This episode is brought to you by Brex, my favorite. You become what you spend on. And I refuse to spend my time on work that shouldn't exist. Expense reports, receipt chasing, and manual closes. The companies building what's next from Vercel, OpenAI, Anthropic, Granola, and Deepgram all made the same call.
17:51They all run on Brex. Brex is the intelligent finance platform that combines cards, expenses, and banking into a single stack with agentic finance built in. AI agents that handle expenses automatically, enforce policy before spend happens, and close your books in minutes. That's why Sorcery runs on Brex so I can spend time on building and not busy work. It's time to get Brex AF. Learn more at brex.com slash sorcery. That's B-R-E-X dot com slash S-O-U-R-C-E-R-Y. Bye. At what point do you think speed is going to bifurcate the market much more into different pricing units? I know that token costs and everything is a really big topic right now, But for consumers, we do expect speed.
18:42But should we be getting enterprise quality speed?
18:46Rodrigo Liang:Well, I think you're going to find that people pay for it. And it's always been true. If you look at the internet, people paid. Initially, they would pay a premium for the upper end of the packages for faster internet. And then some people would still remain on the basic, right? And same with the cell phone. But I'll say this. Look, when 5G showed up, nobody's signing up for 2G. Right? If your phone, if your cell phone is not transferring very fast, you're getting very frustrated, right? So I think on the curve is over time, as fast inference is broadly available, who's going to want slow? right and so it's going to not only be the premium today where kind of the the people who need it co-generation people real-time banking real-time healthcare there's a number of industries where speed matters because it's their livelihood they're going to pay up for it that's kind of the premium service that we're offering but over time what you're going to find is that all of us are going to want that like i i don't see a use case where the average population either consumer or enterprise, is going to say, actually, I prefer this low.
19:56Rodrigo Liang:That wasn't the case on the internet. That's not the case on cellular, you know, on mobile data, right? It's never been the case. Let me pay more for this low, right? Or even let me pay a little bit less for this low, right? Most people over time are going to say, no, I want the fastest, right? And so as the cost of delivering fast goes down, you're going to see most people switch over to the fast. And this is what I feel like, you know, the premium inference, which is large models, which equals the most accurate. Most accurate models and fast, ultimately steady state is what everybody's going to want, right?
20:29Rodrigo Liang:Is whether today we can offer a little enough cost that everybody can afford it, right? And so until then, you're going to see people very quickly moving over if they have a need for that fast inference and have a need for the accuracy, which is, I think, a large part of the enterprise. There's also a couple other macro themes that are happening that are just going to explode data and usage across the world. Yeah. One of them is Starlink. Like, I don't know if you've been flying around on planes that have Starlink, but even having access to that in remote areas, it just increases the amount of work you can do and work for, you know, like edge cases.
21:04Another part of that is edge computing. I think, do you guys deploy in like edge remote areas as well?
21:12Rodrigo Liang:Well, what happens here is, again, this is tying into your question, is a$50 billion data center ubiquitous? It's going to be hard to say that when you have economies that can't afford a$50 billion data center, right? And you're going to see across the world. If you believe that AI is going to be a technology as pervasive as Internet, which I do, right? This is something that everybody on the planet should have access to, right? And so mobile service, internet, AI, everybody on this planet should have access to it. And so if that's the case, then it's not going to be equally deployed because in some countries, you can't afford to put$100 billion gigawatt data center somewhere and then let tens of hundreds of millions of people come in and use it.
21:57Rodrigo Liang:Or you might be in a place where you don't have that and you need something significantly smaller. And today, because of some of the technology being as little as 10 kilowatts per rack, we can put them inside shipping containers. So we build out these data centers, clusters of 10, 20 racks, inside the shipping containers that you see on these ships. And then you deploy them in these edge data center use cases in a much, much more cost-efficient, power-efficient footprint than having to build out this liquid-cool gigawatt data center with its own nuclear power plant. And then you can put a startling connection to it.
22:35Rodrigo Liang:You can bring the internet in. You can actually have solar farms next to it. And then you can actually power that. So there are lots of different ways that people are creating these data center clusters that allow their communities, whether that's in regions or in industries or in sectors, to be able to get access to the best models. Yeah, one particular company I was thinking of is Armada. I don't know if you know Armada, but Armada, they deploy modular data centers on the edge. And so, I mean, it's not just in remote communities, but it's for critical industries. Critical industries that don't have that real-time data that they've had before.
23:14So it's like oil and gas, it's mining, it's oil rigs out in the water, all that kind of stuff. And even for military and defense and that kind of thing. Yeah. No, look, Armada is a partner of ours.
Read the full transcript
23:29Rodrigo Liang:They've been a partner for a few years now. Dan's a good friend. Oh, yeah. And yeah, yeah. And they've got our racks. And the same concept that if you can actually take, instead of having to put a 100 kilowatt NVIDIA rack in there, you can put a 10 kilowatt NVIDIA rack and generate more tokens, right? That becomes very valuable because you can actually fit a lot more output in a smaller amount of space, a small amount of power. And you then deploy it into regions where you don't have the traditional data center available. You are in remote areas where you want AI to actually manage your operation out there.
24:04Rodrigo Liang:And, you know, oil raids and things like that is an example, right? And so it's an incredible opportunity to actually get this type of technology deployed in different vehicles in different ways. Because, again, access to them will be quite varied, right? It won't be just in that very large$100 billion data center for them. talking back into the data center and the racks you're now working pretty much with other chip companies in a way that like you weren't before so how has that evolved is it is it weird do you think that's going to last the world of cooperation you've got you know partners you know that uh competing in certain areas and collaborate in other areas look at the the core tech industry is incredibly small.
24:53Rodrigo Liang:But who is building tech? But if you really look at it, how many people are truly building and deploying chips? How many people are truly building and deploying systems and building and deploying racks and building and deploying data centers? There's actually not that many, right, if you really look at it, in the construct of the entire economy, right? And so we do collaborate. But here's what I'll say. In the end, we're all in service of customers, right? And so customers are coming and telling us, look, we have these NVIDIA racks, we have Xeon racks, we have AMD racks, we have other chips, and how can you actually get my total business operating more efficiently?
25:31Rodrigo Liang:And now that the world's going to inference, I'll give you this example. If you look at any service provider, like inference cloud provider, they've got NVIDIA racks sitting there. And it's doing all sorts of things. Let's say 1 ,000 racks, you're doing training some models from some customers, You're influencing models for other customers. You might be doing HPC, right? You may be doing some biology, some physics. Who knows, right? Games, you're doing gaming. You're doing all sorts of different things. And you are then, you bought those racks and you're trying to monetize. So someone comes in and so we are focused on things, right?
26:10And we know that when you're influencing these models,
26:14Rodrigo Liang:you run it significantly faster at a fraction of the cost, right? And so if I go into that data center, the inference has arrived, and you look at the percentage, 70%, 80 % of those racks are running inference. And so why would you run inference on those racks when you can run it at a fraction of the cost at a higher performance on some Innova? And so route that traffic to some Innova, frees up all these racks for you to resell to all these other things that people want anyway. right and so that the economy starts getting much better because now without buying more hardware they generate more revenue they actually run the most popular models in a much more efficient rack at a much lower opex and a much lower capex and they're generating faster tokens which then you can charge more right and so your premium service is is able to charge more for faster and so you're making more money and in general than just lifting your margins on the same hardware infrastructure And so that's usually kind of what the customers are asking us to do is how do you get our business actually generating better margins?
27:18Rodrigo Liang:Because for them to sustain themselves, as you know, today, inference services, they're not making enough margin. They're generating lots of revenue, but you're not generating enough margin. And in order for them to sustain, they've got to be more profitable. And this is all the investors. You mentioned a couple of investors you're talking to. They're all thinking about, well, how do I sustain this? Well, you sustain it by having every service provider make more money, right? If they're making more money, they can continue to invest. And what we do is we generate more margins by giving better infant service at a much lower cost.
27:49How are customers measuring that difference? And how are they measuring the routing performance?
27:55Rodrigo Liang:Well, the routing's already happening today. So I'll start there. So if you look at a large data center, let's just say you're running open source models, closed source models. You're running some HPC. when a request comes in, you are already routing those models to certain racks. And so these racks are running Enthropic or these racks are running Minimax and DeepSeek. And so that routing is already coming in. You say, hey, I want to run my service over to you, do a prompt, a router over to a Minimax model. It will already go over there. It's automatic. Is there like software in there? Like how does that work?
28:34Rodrigo Liang:Well, there's software that's on top. And so if you look at kind of what these API services are, and this is one of the beauties of what the open AIs and Enthropics really set up, is they always set up these open standard interfaces. And so there's these API calls that you will run, and these are standard API calls, and some of them we actually match that as well. And so that when you prompt, actually it will look for a particular API to a particular model running on a particular IP address. And so once you actually have that, then it's a standard interface. So when you're deploying your racks at a NeoCloud or Hyperscale Cloud, you just have the same API interfaces, and then you can leverage kind of what's out in the open community for routing.
29:17Rodrigo Liang:And so you were already doing that before. So if I go and I'm a customer of GPUs, for example, the GPUs have A100s, H100s, B200s, B200s, Even the different versions of it, those are being routed too, right? Because, you know, if I pay for the newest chip, I don't want to be routed to a no chip, right? And so same thing with now you can just stand up other chips. You can stand up EMD chips and Samanova chips and other chips next to it. And they're all just already part of that ecosystem for routing, right? And so that's kind of the routing. But most service providers, and I think before we started, we were talking about KPIs and economics and how do you measure.
30:02How do you measure?
30:03Rodrigo Liang:And this is the way we see most providers measuring. They purchase per rack. They operate per rack. And so they want to generate revenue per rack. And the revenue is generated per token. If I put a rack of hardware, I'm just seeing how many tokens I might generate in a particular model, and that model has a price per token. Multiply that by 30 days per month, 24 hours per day, number of tokens per second, and you can figure out how much money that rack is generating, and you look at how much it's costing you to operate. So that's as simple as that, right? And so that's what we focus on. We're very focused on making sure that when you deploy a Rack of Summit, you generate great margins relative to the model.
30:46Rodrigo Liang:And you can change the model because different models have different pricing per token. But you're still generating a significant number of tokens per month so that you're making profit on that Rack that you actually spent and are operating on a per month basis. Yeah, to your point, I did wonder this because with some of the conversations I've been having, a common theme is there are no standard definitions for anything. You're like, well, it's very clear with hardware. This is how you do it. And I was like, oh, well, that is the clearest definition. It does get a little murky when you go into like, how do you determine what an AI agent is?
31:24What does that even mean? What does a jailbreak mean? We just talked about this with Don Field. And it's just so interesting. And even way back when, it was like maybe a couple of months ago, I interviewed Teresa Carlson. She's CEO of the General Catalyst Institute in D.C. She used to be CEO of AWS for public, the public sector. And so she worked with government and she was the first one to get AWS actually like a CIA contract. It was crazy. Crazy story. She said the biggest difference for getting cloud into the public sphere to start selling to government was to have a standard definition. And then once you had that standard definition, they were like off the races.
32:05Cloud really started to get adopted. So I guess between all the lines of all these different topics that we've been covering, from the chips to the data centers to power, edge, and also tokens and the costs there, What are the biggest bottlenecks that you're seeing and what are the biggest challenges there?
32:25Rodrigo Liang:Well, I think, look, it's a land grab right now, right? From inference providers, whether that's at the hyperscale level, the frontier lab level, neocloud level, right, or sovereign clouds. It's a land grab. You know, you can come to France, right? There's going to be the hyperscalers coming in and trying to compete there. You've got all the model makers trying to compete here. There are even regional cloud players trying to compete in this space. Regional meaning not just France, but Europe, right? And so you're neoclouds. And so you're going to have all the competition. And so it's all about scaling.
32:59Rodrigo Liang:It's all about who can get to scale faster, right? Because as we've seen over history, the large players globally or regionally end up having this enduring, lasting impact in the market, right? And so people are investing a lot to go out and grab the users, grab the customers, because usually once you're in, once you're kind of using, say, Microsoft or Google Gemini, you're pretty much kind of in that ecosystem for a while. And so you see that investment happening aggressively. And so for Summono, what we're really focused on is providing people with that edge and advantage. What you want is you want to be able to come in and secure as many users quickly without having to deploy the traditional amount of capital that you otherwise would.
33:46Rodrigo Liang:People forget, as much as NVIDIA costs, it's commodity. Yeah. Right? Because what you offer is the same as what your neighbor offers. and your differentiation is, I can save you a little bit of money because maybe I got a discount from NVIDIA, right? Or maybe I got a subsidy from the government on my data center space, right? But in the end, your cost advantage is very limited, right? And so what Summonova does is we come in and we offer differentiated service with premium token, premium inference, fast inference, the largest models, and allow you to remix them with existing infrastructure and now give you the flexibility to create different services.
34:29Rodrigo Liang:Now I can provide a sovereign service. I can provide my own national model on some of our infrastructure running faster. And so for that, now you can blend in premium pricing. And so you can go and capture more users because you offer something different. You can generate more revenue because you're premium. And you're actually deploying greater capacity at a lower cost. And so that land grab that's happening, you've got to find a way to actually do it efficiently because the capital cost today is getting more and more expensive. And so now that everybody's kind of building out data centers, so you have to use that capital really, really efficiently and secure users, secure companies as quickly as you can.
35:12I mean, I feel like you've raised capital pretty efficiently. What do you make of all these other companies that are raising ridiculous amounts? Like, how do you justify that?
35:21Rodrigo Liang:I don't know. Well, look, I mean, well, one, so we're$2.5 billion raised in the history of the company. There aren't really that many companies that have raised into the multiple billions, right, if you really count them, right? There's a lot of money that's kind of been thrown around. But in the end, if you're able to raise that much, it means that there's something about your technology that makes investors comfortable with the fact that you are going to be a long-term player, right? You don't, investors are incredibly astute in this space, right? They don't put their money in if they don't think you're going to last.
35:53Rodrigo Liang:And so that's one. I think two, if you look at why the biggest, you know, the biggest players and kind of the most credible names are able to raise that is because there is a race to, you know, land grab. There is a race to be the dominant player. And in the end, I don't think, I think as much as these data centers and service providers are heterogeneous, right, they're using different chips, NVIDIA and other chips, it's not going to be 100 different chips, right? It might be two or three, maybe three or four, right? That's as heterogeneous as AI infrastructure is going to get. It's not going to be thousands of different versions because there's diminishing returns, right?
36:36Rodrigo Liang:You take someone over that offers premium. Why do I need another premium? and I don't want incrementally more variety in the data center just for that reason because you have inefficiency there. But bringing in providers that offer significantly different services from what they have initially is going to allow you to blend it. And so I think you're going to see that being able to actually take technologies that are different, deploy them, and offer different services will start becoming the way that these service providers are differentiating in the marketplace. Because today, if you look at NeoCloud, What's Neocloud A and Neocloud B?
37:10Rodrigo Liang:What's the difference? And we struggle, right? And we struggle as well. There's a black hole there. There's a black hole here. They run, you know, I mean, similar price. What is really the difference, right? And so I think you're going to see these inference providers coming in and say, oh, I offer ultra low latency agents. I offer very low latency in metropolitan areas. I offer it for banking for most secure and data private inferencing. So now you're kind of adding this layer of technology capability that allows these service providers to hone in on something that's differentiated, allow them to charge these premium services, and allow them to average down the cost instead of being completely dependent on just one provider.
37:53How important do you think it is to have cloud as a product?
37:56Rodrigo Liang:It's really important. And I think being able to give people very low entry costs, developers being able to access for free, being able to kind of get in is really, really important. Some of it, we do that. We do that through our partners. So some of it, for the most part, many of the ship companies have chosen to go build their own cloud and compete with the AWSs of the world. We have chosen not to do that. What we decided that we want to do is focus our energy on creating technology that we can ship. So we ship racks. But what we've done is we've created a broad range of partners that are building the NeoCloud services.
38:33Rodrigo Liang:And so last month we announced this great partnership with Vista Equity and Cambium on this new NeoCloud Vector Core Compute, VC2. And what they're doing there is, you know, so they're deploying these ultra low latency data centers because for them, you know, for them, using some kind of technology opens up these markets that was not possible before. And so we now are able to then offer cloud services through that partnership and actually bringing some of the biggest model providers into that ecosystem because they are able to actually support all the data center, the energy and the facilities that service providers should do.
39:18Rodrigo Liang:And now we can actually then just ship the infrastructure that they need. And so that type of collaboration creates velocity, creates capital efficiency, right? Because that's their business to actually go and build the capital, raise their capital to build data centers. And then we can actually continue to focus on shipping racks, shipping racks, shipping chips, shipping software. If you're building what's next in AI, you need to know MongoDB, the database platform developers love and built for the agents you're running. MongoDB stores searches and reasons over your data in real time. with vector search and embeddings from Voyage AI all in the same system.
39:56No separate pipelines, no stitching together 10 different tools. It's why 75 % of the Fortune 100 and leading AI-native startups run on MongoDB. Build and scale from your first user to billions of vectors. Go to mongodb.com slash AI to learn more. That's mongodb.com slash AI to learn more. Bye. assembly ai is a voice ai infrastructure layer millions of developers build on they build the industry's best speech-to-text voice agent and speech understanding models that serve as critical infrastructure for companies like granola kgen ashby and clickup their speech-to-text models lead the industry in accuracy and quality and their speech understanding models help you go beyond transcription by uncovering insights identifying speakers and highlighting key information from voice data.
40:49You can get started today at assemblyai.com slash sorcery and get $50 of free credits to start building voice AI products. That's assemblyai.com slash S-O-U-R-C-E-R-Y. On the topic of sovereignty, this is becoming a hot topic again because of Alex Carp on CNBC the other week, really talking about sovereignty and owning your data, your products, everything in the pipeline, Berkeley.
41:20Rodrigo Liang:Yeah. And so in terms of your stance, we're in Europe, right? Yeah. So how important is it for Europe to own their own chips? It's been a big part of our business, incredible part of our business. I think owning their own chips, I think every company wants to buy their own infrastructure in order to be able to actually run privately. What's more important is running a service and running models where their private data went into it. Right. That's what they're trying to protect. You know, and so if you look at whether that's sovereignty at a national level or sovereignty at a corporate level. Right.
41:54Rodrigo Liang:I don't want my data trained into a model and had that model shipped worldwide. But I can't imagine if your bank account information starts showing up in ChatGPT in some other place in the world without your permission. And so that's what people are thinking about. How do we protect our information in a way that it doesn't accidentally become part of just the models? And so in order to do that, what a lot of countries are doing is they're saying, look, we don't want to base off of a global model or an American model. We want to base it off of our own national model. And so countries have started doing this work and you see this in Japan and Korea announced the same thing.
42:33Rodrigo Liang:You see in other parts of the world where they're investing a significant, significant amount of money to actually train from scratch, train their own national model for use cases in the government for their own citizens, et cetera, et cetera. And so it's not derived from an American model is homegrown. And then some of the other, we're helping them in that. We're helping them in part of the training process and the inference process of those models. And so you're going to see this globally, that there's going to be more and more emphasis around the fact that we should have, as part of the mix, as part of this heterogeneous AI computing, where you're going to use Anthropik and OpenAI because they're great models.
43:15You're going to use some open source models that are out there, but you're also going to have models that are, at the very least, privately fine-tuned by that country, right, through a partner.
43:28Rodrigo Liang:Or, you know, companies, right, banks and, you know, certain, you know, tech companies creating their own models because they have IP that they don't want disclosed and they want their own user base using those models. And so I think you're going to see that become more and more prevalent. And for that, you need your own infrastructure. Yeah, the point being competition-wise, you don't want to be giving your data up, one, for a security measure on PII or anything like that. But two, you don't want your competitors. This happened with Anthropic and Figma. That was just an insane account to have someone internally and the company itself pretty much copy the product.
44:11Yeah. And so that then disaggregates the types of people that feel comfortable training on that. We recently had Harvey AI on their legal model. And so what they see their biggest competition as anthropic. Everybody is on prem. There's like now a shift back to on prem. It's like the hot thing. And, you know, we always have these like, you know, fun cyclical patterns and behaviors of, okay, like let's unbundle everything. Let's bundle it back up again. Now we're like unbundling again. How do you think that effect will change prices and the service for customers?
44:47Rodrigo Liang:Look, repatriation of infrastructure into on-prem is definitely happening, right? You saw this big shift. Everybody's got a cloud called cloud, you know, everything's got a cloud. and 20 years later, you still have companies just starting the migration to the cloud, right? I was talking to a CIO recently of a big bank, and they were never going to go to a cloud. And it's like, see, I knew all along. I knew all along that, you know, the answer is on-prem. I was like, yeah, it's cyclical, right? And so a 20-year cycle, wait long enough, it'll come back, you know? And so you're seeing that today for exactly the reasons you're talking about, right?
45:17Rodrigo Liang:That people want data privacy, data security. But here's the other thing. if you think about the global 2000, right, most of the largest companies in the world, in the end, what do they have? What is their kind of differentiate? Why do I get to charge a certain amount of money to my customers for service A, B, or C, right? It's usually tied to some IP that I have. I know something, I have data, I have access to, you know, I have these supply chain relationships. is some information, right, that allows you to operate a business that competes in the market. If you actually transfer all of those services, that differentiation, to all using the same exact model, that's in the community.
46:05Rodrigo Liang:Where does the differentiation come from? And so what most companies start to realize, if you just fast forward, it's not 10 years away. It's two years away. If you fast forward and say, hey, I'm going to use all the commodity models, how do I get to charge more? And so when you find yourself in a place of low margin for all these enterprises that historically have been able to enjoy much better margins. And so what we see is the world is starting, the enterprise world is starting to come in and say, okay, how is AI going to impact me? You hear less and less about, it's about saving money. The companies do want to save money.
46:44Rodrigo Liang:And AI could save you some money. Like I could help reduce the cost of XYZ. I'm going to do some paperwork, some operational things. But what you're really starting to have people start to think about is how do I differentiate? Because you are going to have to do top line improvements, generating new businesses, generating new services, generating things that others don't have. Hard to do if you're using the exact same model that everybody else is using. So now they're starting to come in and say, okay, well, I need to train my own model. I have all this data. I know my customers better than anybody else.
47:15Rodrigo Liang:I can customize that better than others can because they don't have the data that I have. I've worked in this industry for 40 years. I know my customers. Let me train my own, and I don't want to give that to a model maker. And so you're starting to see people come in, and yes, the operational cost improvement that AI can deliver is starting to show in certain places, but more people are starting to figure, how do I create a better service, a better experience, a better use case so I can actually charge new services into the market in order to actually generate better revenues from my company instead of just saving money, when saving money is ultimately going to actually erode your top line.
47:55I was really thinking about this from your standpoint too, because you deliver such a more cost-effective and higher-performant product versus the other chips that are out there. But right now, for companies that are out there, they've been told, and I know the token maxing thing as a whole thing and now people are throttling that back a bit but like the guidance
48:16Rodrigo Liang:on everything is like throw all the money you can at ai and so at what point does that proceed and then you start to get more economical about how you're throwing money at ai yeah well i think throwing money at ai isn't really the goal right i think the goal is really is how do we actually let AI generate more business benefit to you, right? So what I would say is throw as much of your cost at AI and see if you can bring it down, right? So not spend more is take what you're spending and dump it into AI and see if you can actually generate more efficient costs. And on the other side is take what you're doing, throw it to AI and see if it opens up new services that you didn't have before that generate more value to you, right?
49:01Rodrigo Liang:And so I think that's kind of how I think the world's going to operate. Before you know exactly how you do it, people are just trying to encourage your company to actually use more to discover. To discover where can I save money, where can I generate new services. And a great example that I see in Google is in this transformation, what they emphasized was everything becomes AI first. because you don't know what we don't know. It's kind of like early days of internet. If we say, what is the ROI for email? I mean, there are companies that say, prove me ROI before you use it. Yeah. Am I typing an email?
49:44Rodrigo Liang:Should I try to figure out what the ROI of this email is? It's really hard to do. So what some companies are doing is say, look, we don't know yet, right? Just go AI first because then two, three steps down the line, you start inventing things that the world has never seen, right? So that's an incredible way of approaching a new technology from a very development-first, research-first type of company where they're creating, inventing, right? And great things come out of that. Now, the danger of that is what you said, right? If you aren't paying attention, it becomes token max. It's just like, hey, just spend, right?
50:26Rodrigo Liang:Well, that's not really the goal of just spend on AI models is how do you actually transform the business as efficiently as possible? And so what we see more companies doing now is saying, look, these are the things that we do. How do we actually use AI to actually make it more efficient, faster, simpler, cheaper? Or these are things that we sell, services that we sell. How do we broaden our offerings into the market by letting AI expand it? And so those two are much more systematic ways that we see companies actually using AI. And then what comes out of that is a calculation of, well, how many tokens does it cost?
50:59Rodrigo Liang:And et cetera, et cetera. But that's not, the goal isn't to maximize kind of token usage. The goal is to actually maximize services that you offer. And then the net result is, can I actually do that more efficiently than I was doing in the old way of doing business? I know we covered so much. And I'm so curious to see what else you have in that brain of yours. What are the biggest questions you think people should be asking today that they aren't? This world's changing very, very fast. The first thing that people need to think about is, how are you going to differentiate? Right? You've got, what is it, 8 billion people on this planet today?
51:39Rodrigo Liang:And you've got, I don't know how many, tens of thousands of hundreds of thousands of businesses out there. Right? And the fear that AI is going to actually remove jobs, I don't think that's really the natural course. I think what you'll see is people are going to use AI to create more services and compete in the market more. And if you aren't doing it, then you may run the risk of losing ground in your market. But most companies are trying to use AI to compete. And so for the first question that people should be thinking about is, how do you differentiate? Just using AI doesn't differentiate. But how do you differentiate?
52:16Rodrigo Liang:Because people want to see services they have never seen before. The World Cup is going on today. And you look at some of the players that are the most popular players out there. Two things I'll say. One, we can get robots out there running around and playing games. And I will bet that the human interest in watching machines play machines is going to be far lower than what humans are able to do, right? I mean, we can get computers to play video games for us, and there are services where people are staring 12 hours a day watching other people play video games, right? And so why? Because we want to see kind of what humans can do, and there's that piece of it.
53:02Rodrigo Liang:And so I think you're going to see, one, that I think people are going to look for services that they've never seen before, There's something that's new, something that I feel connected to, and there's a huge opportunity for people to do that. But you've got to start now, right? Instead of kind of worrying about, you know, is AI going to take away some of the jobs that we have today? And my answer is, I hope so. Some of these menial tasks nobody wants to do. Yeah. Right? I hope it does, right? Starting with, I hope I don't have to do taxes. You know, if you think about some of these things that just, you know, you don't want to do, but it should free up ourselves to create to invent to do things that other people haven't seen and now if you're uh uh the the goldie for for cape verde you go from what you know what is it 50 000 followers to like now 16 million followers overnight because you know he does something that people didn't ever expect him to do right and so i think you're going to see one is that i think for enterprises i think you need to think about how do i scale but how do i scale because the hints that you're seeing today with energy constraints, data center constraints, chip availability constraints, cost constraints, all of those things are only getting exacerbated, right?
54:17Rodrigo Liang:The world needs your service, but they need your service at scale. And so what is the ecosystem? What is the structure that allows you to kind of fast forward and say, okay, well, I'm going to do this in a way that I can actually see a path for me. at scale to compete in the global stage, which not everybody's thinking about because a lot of people are just thinking, well, how can you sprinkle a little AI to kind of save some money and get the board off my back? That could be one thing. But, you know, look, there are companies, when the internet came in, there are a lot of companies who said, you know, it's not going to last.
54:53Rodrigo Liang:Yeah, take a Circuit City, you know, your audience may know, you know, or a Blockbuster video, right? I mean, if you remember at the time, Netflix was shipping DVDs to your home. And people were saying, you mean you're taking a DVD that everybody rents for$2.99 and you swap it for a service, a$5 service for unlimited videos? You're going to destroy your business. Fast forward, it's one of the top businesses. And so if you think about kind of the embracing of the technology and think, okay, well, what can I do that I couldn't do before? And if I get fast and I get differentiated, I'm able to attack the market in a way that I couldn't.
55:36Rodrigo Liang:And then fast forward two years, I'll be the dominant player, right? And that's kind of what these folks need to be thinking about with scaling. So how do I offer something at scale that dominates? And all the other players are slow to transform, slow to adapt. I think you're going to have a hard time competing because as you see, capacity, supply, all those things are hard to get. And just like in Formula One, And if you start your last position versus pole position, boy, it's hard to catch up. It's hard to catch up in a race where everybody's going really fast. Damn, that was a great answer.
56:09So I guess on that note, as we close out, you started the company in 2017. You've scaled it up to$11 billion. This is a question I like to ask on performance. One of our sponsors is Brex and they're all about spending sport and moving faster. but on performance I like to take it a more personal way so for you as you've scaled up the company throughout your career I believe performance kind of comes down to who you surround yourself with who you're mentored by who you're surrounded by like inspired
56:38Rodrigo Liang:who are those people for you what I've learned in 32 years in the chip business is you've got to be resilient you've got to be stay in it it's never a straight line business at scale is going to come with the highest highs and sometimes the lowest lows and you've got to fight through all of it it's never a straight line and Jensen talks about this with NVIDIA and all the different things that that company went through and you know you see with some of the largest companies but being very systematic about what is it you're about what are you trying to do you can't control what the world wants at any given point in time you You can't control what the economy does.
57:21Rodrigo Liang:You can't control what the politics do, right? What you can control is your conviction around what you're building and being resilient and staying with it, right? Because great businesses don't just show up overnight. You have to keep at it, keep at it, keep at it. And this is what I'm proud of. We surrounded ourselves with great people that are incredibly hardworking, but more than that, incredibly resilient, right? and they just tackle the next challenge, the next surprise, the next thing that's got to be done with a level of consistency. And you keep going at it, and then you wake up one day, and you build something great.
58:00Amazing. Before we close out, we have to see the chip.
58:04Rodrigo Liang:Yeah, here it is. This is the latest Salmono chip. This is the SM50. It's something we announced earlier this year, but we're now actually producing some really incredible results on this. This aggregated inference that we talked about with NVIDIA chips, with some of them already used with Xeons, is just the most efficient way of actually deploying inference at scale. And we're really excited to have this out there. It's going to change the world. We're able to create very large clusters for hyperscale deployments down to very, very energy-efficient deployments for edge data centers. And so it's something that I think is going to change the inference landscape.
58:49Rodrigo Liang:And I think a lot of people are going to use this for their premium service. Amazing. Thank you so much, Rodrigo. Great. Thanks for having us. Awesome. Huge thank you to the entire RAISE team for an incredible event. And thank you to Brex, MongoDB, and Assembly AI for making this trip and series possible. If you enjoyed this conversation, you're going to love the rest of the RAISE series with Tony Kim from BlackRock, Scott Wu from Cognition, Andrew Feldman from Cerebrus, Rodrigo Yang from Salmonova, Michael Hurlston from Lumentum, CJ Desai from MongoDB, and many, many more, like our hot takes that we did at a secret location that you can find on X, YouTube, and Instagram.
59:29Subscribe to Sorcery on YouTube for more conversations with the people shaping AI and join the free newsletter. You can also do paid at Sorcery.vc for weekly insights on AI, robotics, enterprise software, consumer, semiconductors. Did I say AI? AI again. And everything that's coming next, like funding announcements and all big things in tech. Thank you. Bye.
From the publisher
Rodrigo Liang is the CEO and Co-Founder of SambaNova. The company just announced a first close on a $1B round at an $11B valuation, led by General Atlantic with T. Rowe Price and Capital Group participating.
Rodrigo has spent 32 years building chips. We covered the 101 of semiconductors and data centers: what inference is, why it is a different problem than training, and why inference will require orders of magnitude more chips than training.
He walks through SambaNova's chip lineup from SN10 to the new SN50, why a 10 kilowatt air-cooled rack changes where AI can be deployed, and why running a trillion parameter model in a single rack matters for agents, latency, and edge.
We also covered the era of premium inference, how providers measure revenue per rack, coopetition and traffic routing across Nvidia and AMD, sovereign models, the move back to on-prem, and why tokenmaxxing is the wrong goal.
Special thank you to Brex, MongoDB, & AssemblyAI for helping make this RAISE AI Summit mini-series in Paris, France happen.
Sourcery covers the people building the future across AI, hardware, and the private and public markets, subscribe for more.
Rodrigo Liang: https://x.com/rodrigoliang?s=21
Molly O’Shea: https://x.com/MollySOShea
Sourcery: https://x.com/sourceryy
𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊
YouTube: https://youtu.be/etod19D9IEg
𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒
• Brex—The modern finance platform, combining the world’s smartest corporate card with integrated expense management, banking, bill pay, & travel. https://brex.com/sourcery
• MongoDB–Millions of developers and more than 65,200+ customers across industries, including ~75% of the Fortune 100, rely on MongoDB for their most important applications. With integrated capabilities for operational data, search, real-time analytics, & AI-powered data retrieval, MongoDB helps organizations everywhere move faster, innovate more efficiently, & simplify complex architectures. https://mongodb.com/ai
• AssemblyAI–Millions of developers use AssemblyAI to power their voice ai applications and features. One API gives you access to best-in-class speech-to-text, voice agent, and speech understanding models for both pre-recorded and real-time audio. Granola, ClickUp & HeyGen are scaling with AssemblyAI - get $50 of free credits today at http://AssemblyAI.com/sourcery
𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒
(00:00) Rodrigo Liang , Co-Founder & CEO at SambaNova Systems
(00:59) SambaNova’s Series F : $1B raise at an $11 billion valuation
(03:00) The Inference problem nobody saw coming
(04:52) SambaNova's chip evolution
(07:19) Running a trillion-parameter model on a single rack
(11:00) Do $100 billion data centers actually make sense?
(14:14) What "premium inference" really means
(18:28) Speed is about to become AI's biggest price tag
(20:43) Starlink, edge computing, and AI reaching every corner of the planet
(24:27) Working alongside NVIDIA and rival chipmakers
(27:49) How customers actually measure inference performance
(32:07) The biggest bottlenecks in AI's global land grab
(35:12) Justifying the billion-dollar AI valuations
(37:53) Why SambaNova refuses to build its own cloud
(41:03) The "AI sovereignty" debate
(43:48) Data privacy fears are driving the return to on-prem AI
(47:55) How to actually get ROI out of AI spend
(51:16) The one question every business should be asking about AI
(56:09) The mentors and lessons behind a 32-year career in chips
(58:02) Unveiling SambaNova's newest chip, the SN50




