In short
Podcast Episode Notes: AI Inference: Why Speed Matters More Than You Think
Podcast Overview
- Title: The Neuron: AI Explained
- Hosts: Grant Harvey and Corey Noles
- Guest: Kwasi Ankomah, Lead AI Architect at SambaNova Systems
- Release Date: Every Tuesday
- Description: The podcast discusses the latest AI developments, trends, and research, providing insightful commentary on AI.
Episode Information
- Episode Title: AI Inference: Why Speed Matters More Than You Think
- Episode Description: Discusses the importance of AI inference speed in AI applications, featuring Kwasi Ankomah from SambaNova Systems. The episode highlights how SambaNova's chip architecture optimizes inference speed and efficiency.
Key Takeaways Importance of AI Inference
- Inference Definition: The process of using a trained model to make predictions based on new input data.
- Latency Impact: Slow inference leads to poor user experiences and increased operational costs. Acceptable latency is crucial for real-time applications, such as voice interactions.
SambaNova's Innovations
- Chip Architecture: SambaNova uses a unique three-tier memory architecture that allows for efficient data storage and processing, enabling them to run large models efficiently.
- SRAM: Fast on-chip memory.
- HPM: High-performance memory for intelligent caching.
- DDR: Large capacity for storing multiple models.
- Performance: SambaNova can run large models (e.g., DeepSeek's 670B parameter models) with significantly lower energy consumption compared to traditional GPUs.
Challenges in AI Inference
- Scaling Issues: Increased user demand leads to more inference calls, stressing hardware capabilities and causing latency issues.
- Cost Management: Inference costs can escalate rapidly with user scaling; milliseconds of latency can translate to millions in operational costs.
Future of AI Models
- Trend Towards Smaller Models: There is a growing realization that smaller, efficient models can perform adequately for many applications, reducing costs significantly.
- Agentic Applications: AI agents tend to use more tokens (10-20x) than traditional applications, increasing the need for faster inference.
Infrastructure and Observability
- Real-Time Observability: Just as databases require monitoring, AI systems also need fine-grained observability to ensure reliability and performance.
- Human Oversight: For AI applications to gain trust, a balance between autonomy and human oversight is necessary.
Discussion Highlights
- ROI Challenges: Many organizations struggle to see a return on investment from AI initiatives, often due to high inference costs and lack of proper observability.
- Open Source Models: The rise of effective open-source models presents alternatives to expensive, proprietary models, allowing businesses to optimize costs further.
Future Outlook
- AI Infrastructure Predictions: The next 12 months may see significant advancements in AI infrastructure, particularly in optimizing inference speed and reducing costs.
- Agent as Infrastructure: The concept of AI agents as a foundational component for various applications is expected to grow, emphasizing the need for efficient and reliable AI systems.
Resources
- SambaNova Cloud: [SambaNova Cloud](https://cloud.sambanova.ai)
- Solaria Speech-to-Text API: [Gladia Solaria](https://www.gladia.io/solaria)
- The Neuron Newsletter: [Subscribe](https://www.theneurondaily.com/subscribe)
Conclusion Kwasi Ankomah's insights provided a deep understanding of AI inference, the importance of speed, and the innovations at SambaNova. The conversation highlighted the need for efficient AI applications, the cost considerations of inference, and the future landscape of AI infrastructure as it evolves.
---
These notes aim to encapsulate the essence of the podcast episode while offering a structured overview for easy reference.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00All right, hello and welcome to the Neuron podcast. Today we're talking to Kwasi Ankomah. Kwasi is the lead AI architect at SambaNova Systems, where he specializes in agentic AI and solving the critical challenge of making AI models run fast enough for real-world production applications using SambaNova's revolutionary RDU chip architecture. So we thought he would be the perfect guest for the Neuron podcast. Just a quick FYI about SambaNova Systems. SambaNova builds custom chips, systems, and platforms that let organizations train and run large AI models more efficiently than with standard hardware.
0:46Hi, Kwasi. Welcome to the show. How's it going? Hi, folks. How are you doing? Hi, Grant. How are you doing? I'm doing really well and super excited to talk to you folks about AI inference and agents. So, yeah, super excited.
0:57Kwasi Ankomah:awesome we're excited to have you here it's uh it's an interesting time and uh sounds like you guys are doing some neat work yeah definitely we've been we've been kind of seeing a lot of shift in the market you know we we had this kind of huge focus on training i think everyone did about how to train these large language models and now we've kind of seen that around you know the biggest plot on that that we've got now is inference right so how do we make things How do we make inference fast? How do we make it scalable? So we've been really focusing on our architecture in speeding that up and making it more efficient and delivering these solutions to our customers.
1:34And my team really focuses on the agentic side of things, which is what I'm super excited to get into because that is showing why inference matters and all of these calls and the number of tokens is going up. And that's a really interesting element at the moment as well. So yeah, that's where we're trying to focus on at the moment. Yeah. Well, I got to ask. Okay, so let's just clarify. So very simple before we get to agents for our readers and listeners who use ChatGPT daily, maybe don't think about what's happening under the hood. So when you type a prompt into ChatGPT or any other AI and hit enter, what actually happens?
2:08Like what is inference in plain English? yeah so inference inference is coming for the from the obviously to infer so it's the it's the model going along and then making a prediction of some sort so it's taking your input and then it's basically doing the thing that large language models do which is the next token um and it's that is the actual um process of inference it goes in it runs through the model and we get an output and that output keeps going and essentially all you all all that you're doing is that we have a model that has already been kind of put on some sort of architecture. And then we are basically giving you the answer or the next token.
2:46And then, of course, as you see that stream, the next token then goes back in and we get the prediction based on the next token as well. So that, in a nutshell, it's essentially giving you, we have a model that's already been trained and we are just giving you the output of that model. Yeah.
2:59Kwasi Ankomah:Well, you know, so, you know, we always think about the idea that training is the hard part for these massive models but you and Samba Nova team are often talking about how inference and actually running them is the real challenge can you kind of explain why that is yeah so I think inference speeds kind of kind of directly dictates how the user interacts with the application right so we've all been there on you know your favorite chat that application be that what it may and when you kind of press that button to inference right so when I've talked about inference you know you're making a pass through the model and you're getting output at the end now that pass can take a long time depending on the size of the model the amount of parameters and the hardware that it's running on now if there is a big latency with a real-time application that does become an issue and we've seen that for you know if you try to run certain models on certain architecture you can have you know a time to first token that's when the first you know the time for the first token arriving at the user screen you know between kind of like 20 and 30 seconds you know if you're if you're if you're in a real world production system that is just not acceptable right that's not that most production applications don't have a latency budget of 20 to 30 seconds especially if there's a user interacting with it so i would say that the user experience has become a key factor as these things go into production how do we actually make sure that the user is having a good time and that we can scale it right like what happens then you know when we instead of having 200 users we have 2 000 users right because again you're you're doing you're running inference and you need to scale that inference to the amount of users so that that becomes a problem because usually the more the more users you have the more the hardware is under pressure in order to get the inference out and the slower the model goes again so i would say that that's probably the the the biggest the biggest one then the second is kind of the the kind of reality of like where is the cost going like so if you i always thought you know i came came from financial background and when ai first came everyone was like it's a bit like the cloud you know this is great you know no one was checking their bills you know and now you see what inference is costing you and suddenly you're like well hang on like this inference is becoming most of our cost so actually and you know milliseconds of difference actually can mean millions in operational cost like when you scale it up to kind of you know 20 million users 30 million users these getting out tokens faster is is just going to cost you less money so you know one of the things here is that inference itself is becoming an expensive game.
5:45So those would be kind of the two things, I think, from a business perspective and from a thing that people can relate to is that you're not going to have a good experience if you have slow AI and it's going to start costing you money. The last thing I'd also say is that there's now certain applications, and I'll use voice because I think it's the, you know, as certain applications, not only do they not have a latency budget, but they're like, they're almost like latency critical applications, right? So voice is one of those applications where it simply doesn't work, you know, if the latency isn't there.
6:18So I think other things you can say, oh, it works. It's just a bit stuck. Yeah, exactly that, Gory. You've got someone who's, and then you get, you know, you go get your lunch and you get a response. Now, that just doesn't work. So we've got these kind of new crop of applications that simply cannot have high latency. And that's what, you know, for me, the three key kind of differences are, if that makes sense.
6:40Kwasi Ankomah:You know, especially dealing with products, too. You know, I mean, the user just has no patience or tolerance for weight on any scale. And, you know, I think it's interesting you talked about volume, too. And as these things are scaling up, you know, you kept hearing these discussions for a long time about how the cost of intelligence was going down. And it's something I think they're never clear about is that like the cost of serving an individual unit of intelligence, for lack of a better phrase, is very cheap. But if you're dealing in product where you're serving to many, many people, those little fractions of a penny add up in a hurry.
7:21Yeah, they really do. And, you know, one of the things that we see a lot, you know, is as you started AI projects, you tended to start them on these kind of huge models, right? And, you know, they'd be like a Claude or like a GPT-4 or something like that. Now, as you say, if you run it for like a proof of concept, then this is completely fine. you know the cost but as you scale that out to many users it becomes a huge huge cost and as you're as you're serving tokens what what we're seeing again this is the kind of the the agentic stuff we'll get into is that old chat applications used to serve maybe one to two thousand tokens or that's what they needed now these new agentic era of applications are you know we're seeing like 10x in that so that's that's what i think has caused this real focus like this narrow focus on right hey come on you know oh do we need this model and you know as I'm a no but we you know we kind of work with the ability that you can host multiple models around the same thing because you might not need this expensive model for everything you might need have this cheaper model that might be able to run some of your workload as well so people are now really zeroing in on this because you can have if I have a whole product say and it's costing me let me give you illustrative example three million pounds a month to host now if you just take a slice of that and you actually put that onto like an open source model you might be able to reduce that cost by again a factor of like four or five you might be able to get it down to like 100k and now businesses are realizing that this is as they as they put more and more stuff into production this is becoming a huge thing so as you said cory like it's it's starting to really hit home and i think with most things doesn't matter if it's ai the moment it hits the balance sheet and profit and loss then people do start to get interested very quickly because then it's like, okay, hey, like, well, you know, how can we make it more efficient so that we can scale as well?
9:20Yeah.
9:20Kwasi Ankomah:Here's a common developer nightmare for you. You've finally shipped your AI agent and the transcription falls apart the second someone speaks with an accent or calls in from a noisy cafe or switches language. Does that sound familiar? Thought so. It's not just frustrating. Failed calls can kill your business. And that's exactly why Gladia built Solaria. It's the first speech-to-text model designed to handle the messy real-world stuff with sub-270 millisecond latency. More than 100 languages supported. More than 200 ,000 other developers are already building with it today. If you're serious about real-time voice AI, this is the API you want in your stack.
10:01Kwasi Ankomah:Head over to gladia.io slash slash Solaria to give it a try. That's G-L-A-D-I-A dot I-O slash S-O-L-A-R-I-A. Well, did you see that paper that came out recently? It was like from MIT, I believe, and it was kind of anecdotal, but they were talking about how 95 % of organizations aren't getting an ROI from their business. First of all, did you see that paper at all? Are you familiar with it yeah yeah do you take that to be that more of an organizational issue or is that or is the is that the inference coming home to roost like what what's your take on that i think this is this is definitely an interesting statistic right i think like with a lot of things in business if if the if the initiatives aren't kind of probably thought out or they you know people haven't thought about those costs then they then they will kind of skyrocket right because i think getting these things off the ground has been expensive especially if you've been using only a large you know only like a what we'll call a frontier level model as your only model because the cost of that will be astronomical you know maybe you're talking 20 to 40 times what you could do it on other systems so i think that definitely gone is a is a key factor into why the roi is going into into negative the second one that i think i'll kind of talk a little bit about is is that i think a lot of projects like we all when business intelligence first came and the cloud first came is they they've been run not without i would say observability and evals from day one and i think that's you know one of the principles i would say of building agents especially is that you need to have that visibility from day one i think and this is you know once you once you want to get frontier tech out i think lots of people are very keen to kind of see does it work does it work right now from day one we always talk to folks about the fact having that observability allows you to kind of think about the costs yeah you know how how can i reduce the cost from day one and also is this thing working and that's the thing and i would say the third thing is people are now realizing that there are viable alternatives right like i think there was when when the when it kind of first came there was you were talking maybe you were using uh maybe open ai you're using anthropic but now you know what we're seeing is that there's a huge demand for kind of like open source models and alternative infrastructures that can be served a lot cheaper so again if you you know if you swap around that cost base for what's available or say like our cloud you would start to see that roi definitely swing around but i you know i do think there's definitely there definitely is a um there's a need for more i would say oversight of some of these of these of these projects and i do think as we shift to kind of smaller models and more efficient models that will start to change i think this has been like almost a ramp up of how do we start ai we've all been using tokens um and not really thinking about it and now people are going to really start looking at that and there's good though i think the value has been generated so i think the business value has been generated it's just at the cost right i think that's what i think when people sometimes read the paper what they take out is that oh these projects aren't delivering value but we we know that they these things are delivering the business value the problem people are having is that how much is it costing right and i think this is what we're talking about everyone is really and this is what we we all keep talking about inference and efficiency and power how much is this costing and we are architectures like we come from the fundamental things of saying we know at scale how much it's going to cost so we've already kind of thought about that and the folks who do the hardware really kind of thought about that yeah yeah that's well yes there's a there's a ton of companies that are offering ai inference right right now right like from big cloud companies to all these startups and to even nvidia themselves in some cases right so so what what are you all doing at uh samba nova systems differently and what's your unique approach we kind of touched on it but let's like really dive in into that here like to kind of help people understand yeah for sure so i think like you know I'll talk about probably the key two areas, right?
14:14So one of them is around the chip itself, right? So we've done something, what we call three-tier memory. So that three-tier memory is basically SRAM, and that's like super fast, you know, on chip, integrated with the compute. We've got HPM, and then we've got DDR, right? So HPM is the intelligent crashing layer, and then we've got the DDR, which is like this massive capacity, like 1.5 terabytes that we've got as well. So why does this matter? It's because traditional GPU move data between memory and compute. So it's like you're constantly doing this back and forth, back and forth. So it's like having your tools in the garage when you're working in the kitchen.
14:58You have to constantly go, oh, yeah, I need a screwdriver. Now, like that, that's inefficient. and again at a small scale maybe not maybe not a huge problem at a big scale with thousands of compute um um units that does begin to add up so that that that innovation is big and the reason that's big is it allows us to store bigger models so to give you an example of this the the um the deep seek models so that would be um deep seek r1 and deep seek b3 they're 670 billion kilometers right huge models absolutely huge now a lot of providers don't serve that model they can they can't physically and we can basically be because of the way that we've architected our chip and again the folks that you know um you start the company they they they had all of this in mind when they designed the chip so that's the big thing it allows us to run very large course now the second thing that that allows us to do is it allows us to run many models so that ddr bit It allows us to store.
16:00So if you imagine that you, you know, a GPU or some of architecture could only store like this one model. And in order for you to get another model, you need another unit of compute, another kind of, you know, let's put it a node. Now, we, because of our kind of large DDR, it allows us to kind of store these other models. So you can switch. And this becomes super important for agentic applications because you might have an application that maybe uses, you know, the GPT OSS model that we're, you know, we're running at the moment. Or it might use a Lama 8B. But we can have those on the same node.
16:31So your inference and hardware cost stays flat because we can go and get the model. So that, I think, is a huge innovation that other people don't have. And then the second, and I think the big, probably the most talked about thing at the moment, is the energy consumption. The energy consumption, I think, is honestly far and away. Speed is absolutely fundamental. We know that. but if you're actually talking about what is keeping cios up at night it's that cost and that cost is coming from the fact that their energy is increasing and there's many other reasons right you know for we need to reduce the energy consumption of these chips right and one of the biggest things that we've done is that we've managed to get something like a you know i'll take the um deep seek example running on 16 sn40s which is our chip but running at you know like 10 kilowatts now to give you like the the closest you know you've got kind of like lights in video where it's like over 100 right you know like this is this is a huge huge huge difference in the in the kind of the energy efficiency so those are the two biggest innovations right because those are the two that are that are a differentiator in like how it's how do you run a model as fast as possible, but how do you run it as efficiently as possible?
17:52And you know, and how, and the third, how many models can you then serve up to the end user? So that our, our kind of hardware and our full stack approach has been tailored to answer those three questions. And so that, that's really where the three innovations are. Yeah.
18:06Kwasi Ankomah:Yeah. Could you kind of, well, we'll come back to energy because we want to talk about that. No, you're good. You're good. Well, we'll come back to energy. We'll get there. Could you kind of walk us through the developer experience? Like how does someone actually use the SambaNova platform? Is it as simple as an API call? Is it something more in depth than that? We'd love to learn some more. Yeah. Yeah, definitely. So we've got kind of like a few ways that you can, I would say, you can connect with SambaNova or you can utilize us, right? One of them is kind of straight through the SambaCloud experience.
18:41And this is kind of very similar to what other folks have, you know, open AI compatible API. You can go on, use all the models that you'd like. No real, if you're, if you're writing an application in, in some other thing, very minimal change, right? You just change the API key. Now we have lots of integrations with, you know, agentic platforms, voice platforms, all this good stuff. So again, very little lift. We manage the cloud of course, and we host all of that good stuff as well. So it's no real. And of course it's kind of like a pay as you go kind of service, similar to what you've seen yeah now we have of course different versions of this so we have like um tiered so you can kind of go up to the to more tiers and get more dedicated infrastructure then we have something that's called um like samba managed and that is basically us giving giving you the full samba nova experience so essentially it's like um you know a samba nova cloud that we can manage for you or you can manage yourself so that's kind of more of an enterprise thing so yeah if you imagine so the the first thing i said is you're using sambo nova but we've we're we're hosting the racks for you we're managing all of that stuff um and what we're saying to you now is actually if you wanted you know 10 of our racks you can actually we can have that so you have your own kind of dedicated cloud you can white label it you can do what you want with it and we will kind of manage it for you and you can have that same you can have that same configuration but you can host it on premises as well now why that's super interesting is because one of the key things that we found is that there's a lot of you know like data centers and folks who want to get into the AI business right they're like oh how do we how do we do something with this power that we've got right well how do we do something with that and some might be at the stage where they say oh you know we want to be we want to um we want to get involved but currently we can't um use enough power or we can't you know our our our specifications don't work or you know we might have too low a power draw for some of the other providers so it won't be able to power our thing so again we're we're a good fit there and what we're also seeing is you know we're a you know we've been talking to a few um customers in this space is like when you have very niche kind of edge requirements that our power draw is a real big factor because then it's like oh hey we've got this kind of old thing that we couldn't do anything with and now suddenly we can deploy ai so it's almost like ai on the edge in these weird scenarios where you wouldn't be able to put other traditional um racks and things like that so that's that's a super exciting product but i would say those are probably the two ways that people kind of um work with samba nota yeah got Got it.
21:25Got it. And then let's talk about like the agent aspect of this, because I know that's like something we all are very excited about. I guess specifically, so like how would someone use Sombanova in an agentic way? Where do you beat like NVIDIA for agentic or, you know, that's a strong framing, but like where do you think you're the best option there for the agent use for us? yeah yeah so i think for for us we kind of think about it in in a couple of stages right so the first thing is around agents use a phenomenal amount of tokens so that's the first thing they they use 10 to 20x amount of tokens and this is this is this is because maybe we've now got more models that are able to reason so we've we want to put those models um in as a let's call them a planner kind of agent that might go to some subagent so that is become very token intended yeah so that's one thing the second thing is that as because of that speed has become more important right so the way that we actually talked about it is that if you if you just do one call to the llm so as we talked about inference right so yeah that one inference that you do um that you know if that's slow maybe that you're like oh you know three to four seconds okay i'll take the hit right and then it comes back from the model i get my answer now because agents do so many calls that gets that basically that delay gets exponential so and we've seen this right like the applications i build you know we do 10 to 20 calls and if if those 10 to 20 calls are each like three seconds that gets suddenly you're sitting there waiting quite a long time for this and i an example of like a coding agent right coding agent say coding agent goes in um it takes the prompt and it thinks about something it says okay i'm going to give it to this planner agent that planner agent then goes to a sandbox who then runs the code and then speaks it out for analysis already there's like six things that have happened there each of those calls is inference so once that adds up our inference speed starts to make a big difference yeah so that's right that's probably where we've big you know heard the biggest um you know kind of praise from our clients is that something that was taking running you know let's say 150 tokens per second on nvidia or like we're running it at like 700 800 tokens as a guy who's running local models on a laptop i just can't
23:55Kwasi Ankomah:talk about how awesome 700 seconds even yeah seriously yeah and you can you can go on samba over cloud and you can see you can see our kind of you know our our token speeds and that makes a huge difference you know that's the bit you know we we just did up one of our partnerships and they actually showed a video of us versus like naked so you know they were just kind of you know couldn't believe the speed because it makes it in those real-time applications it makes a difference so that's one place i think that we're able to kind of be more um i would say i think we're able to outperform gpus in in that in that sense the second is around that kind of model coordination and model bundling and what i mean by that is you don't always need you know the same model or a huge model for certain tasks and to give you an example right if we just if we just stay with that example about the coding agent right so in in a gpu you might or other architectures you might you know use the the the frontier model which is super expensive and huge for all of those tasks now right that isn't super efficient right because and if you wanted to swap to a different model you would still have another piece of infrastructure because you can't have this kind of concept of model swapping due to the memory limitations of the chip now we because we're able to allow you to swap out models on the fly on the same amount of hardware it means that the efficiency is a lot better and the total cost of ownership especially when you have the rack is a lot cheaper so what i mean by that is let's say that you're using um that coding agent and we want our top level agent to use the the funky model because it's doing all the planning but the model that actually just goes and reads like the code and does like some note taking we don't we can have a much smaller model so to give you an example we have clients who have done this they swap that out for like a llama 8b you don't need very small model but you don't need it and that does two things right cory firstly it speeds up your application because we can run that at ludicrous tokens for a second and probably the most important the cost of inference reduces dramatically right because you are no longer paying top dollar for that premium model you're actually swapping it for and the difference in cost is it's a bank show you're talking you know 40 50 times cheaper so if you can like i was telling you if you take that slice of cost that was costing you a million and you just swap that out it really reduces it so i think that's that's kind of the yeah the second where we say that you know the the inference matters but just the fact that we can swap these models in and out and give you that choice on your hardware starts to really have that power and efficiency story that i like to talk about yeah that's interesting so is that something where the the like let's say like the agent builder themselves is like working with like like your custom software to do that on the backend or how, like this is very technical, but we don't have to spend much time on this, but I'm curious.
26:54Yeah, it's so cool. It's a fair question, right? So you've got kind of two ways of doing this, right? So you have, if you are working with us on a kind of a cloud basis, then you can like anything, you can just pick the model. But if you actually, you have this hardware yourself, you can load up what we call, you know, like a kind of essentially a bundle, you know, so you can choose what you need for your use case so then it's there then at that point right you just pick it like you would any other model everything to the to the end user once you've done that it's just it's just an it's just an open ai yeah you just sort the model name and it's that model is available to you so it's it's you know it's there ready and you can use it in your in your workflow and again we've had we've had people do incredible things with this we've had people you know like use a quen model for this at like 32 billion parameters and have like a llama and And then they might have, again, GPT OSS.
27:47And then you can have those all sitting on a chip. Now, that makes a huge difference in terms of if you take that, if you take the cost profile of that application and you compare that to a cost profile running on, say, like Claude or GPT Fart, the difference will be outstanding. You know, like all of the costs that you see here versus all the costs that you see there. So this is the start of the thing that we're starting to see people shift to this whole thing of like, hey do i need this thing and b is there a better way to do this you know using this kind of
28:20Kwasi Ankomah:this this multi-model architecture as well to anyone wondering why cool there are so many mini and nano and tiny models coming out right now this is it yeah it is it's it's it's a big one right i think there's you'll see lots of papers around like you know i think there was recent one nvidia did where they were talking about agentic uses like small small language models with fine tuning can be can be great and i think you know the new the oss 20 billion parameter one is a great great use case you know it's a there's been some models that have really stood out in this area one was the llama 8b right it was a fantastic model for fine tuning it was small enough that you could really fine tune it fantastically well for your use case and get it running and a lot of our you know a lot of our clients they've chosen to do that they've been like hey you can we have this concept to bring your own model so you can bring your own checkpoint load it onto our hardware and then of course just just use it as you normally would but um as you say the the overhead of switching models is is something that we've really really thought about and thought how can we make this as slick as possible so that people can actually run these workloads using that so i think something that the agent like you can tell the agent to do in the back end like like for this type of task use this model like uh like yeah is that the system prompt like that's what i'm like what was blowing my mind yeah so so because the way a lot of agentic frameworks run at the moment right is that you can actually name once you build the agent right so if i take something like lang graph or crew ai right two of two very popular frame and also autogen for that matter all three of those frameworks allow you to at each stage that you do that you define an agent to pass a model.
30:07So you can pass in. And so that you can pass any model in at any stage. Now, what we tend to see a lot of builders doing is that they just use, they're like, oh, I'll just use this. And they use the same model for all the things, right? But you can change that model. Now, of course, what's going on in the background is that you're still paying for that. If you change it, you still have to pay for it. What we're saying is that we want you to be able to switch those models that actually have the footprint have be super small so it's like maybe it's one or two racks and that costs you control and that's what we're trying to say to people is that the overall total cost here can be a lot cheaper um but in terms of answer your question you you can actually yeah you can you can define your agency use any models that you kind of want you can kind of connect it to any cloud it's that's open ai compatible i think all the frameworks work with that the the the shock that you might have and i think people are waking up to is the cost that comes to you after you do that yeah when you try and scale it that's that's the real kind of thing that people are struggling with it's like you know how do we how do we keep this within budget and again the thing that i think people don't talk about enough is power i think that that that is a real factor here in terms of like you in order for in order for this to actually become scalable power is actually the thing that you need to make sure that you're controlling for sure yeah we just we just saw data uh there was an economist piece and rohan paul who's like a really great uh ai twitter uh personality shared this or basically nvidia's grace blackwell chips use about 40 kilowatts for inference uh versus like 12 kilowatts for computing you can correct me if those numbers are off this is what this tweet said uh and training jumps to 80 kilowatts um so like i guess from your perspective how much how much of the energy efficiency problem is about the chips themselves versus the architecture is it both what's what's your take on that yeah i think you know there's definitely i think there's i think there's both i think the chip is massive for sure like you need to but for us we one of the um the advantage of being full stack is that we control everything so you know we have some super smart people at every level and that the efficiency of the compiler the runtime the the kind of hardware itself is massive but of course i think the architecture of the chip is definitely a big one so if you can if you can run these larger models at a smaller footprint by by default you will just manage to reduce it because the key thing that you need is like and to give you an example is how many chips do you need for model x right that's the like so how many chips do i need for deep sea because the number of chips will basically be how many how much power you need to draw to run inference for this group of people so i think you know the the more that you can reduce that then it means that one unit so for instance we are sm40l has 16 chips so our our kind of key metric is what can we fit on those 16 chips because that footprint is is you know what we call a rack so essentially or samba rack and so essentially once you figure that out you're able to i would say greatly reduce the energy and of course the full the full stack does make a huge difference but i would say making sure that the chip architecture is really was really sound is a key one as well yeah um but yeah i'm not surprised that we're seeing that and to to give you an example the the the kind of samba manage that i was talking about we're seeing so many people kind of come to us being like hey we've got this we can only get you know 20 kilowatts of energy and there's there's just you know what can we do with it and there's almost no one who can work for those no and you know we i think we're going to see more and more of that as as we go in as as we go in and and we need to use less energy anyway for various reasons we are going to see more of that where people say actually how how can i would It's not going, we want everything to go as fast as possible, but we also want to say, yes, you can go as fast as possible, but what is the cost of going that fast, right?
34:20To give an example, if you can run at 1 ,500 tokens per second, like, you know, so like some market competitive, right, super, super fast, but the amount of chips they need to do that is phenomenal, right? And I think that's where we're seeing a lot of like, you know, you see these top line numbers and I'm like, okay, but look at the, you know, tokens per kilowatt is a key metric. Like how many tokens do you get per kilowatt of energy that is being used to power those ships? Yeah, yeah.
34:47Kwasi Ankomah:That's huge. And not a metric I'd ever thought of, but you're spot on. Yeah, no, I think again, like I think some of the, the reason why, you know, agentic, people always ask, why are you connecting agents and power? and i'm like agents are using more tokens so like yes inference but they they're just scaling up a the number of models and be the number of things that you want to do with them so the next natural bottleneck of course is going to be power so that's why you know we've kind of we're really focusing on those problems and we're trying to see you know it's it's not just it's not just chips there's you know there's having the whole ecosystem makes a big difference but we can we know that that is a message that's definitely landing well and we see we see that only getting more important as as time goes on with that that's the thing i keep telling people when they have concerns about um you know like ecological concerns environmental concerns around power is that i keep telling them i said all of the incentive is on companies right now to figure out how to do this more efficiently like like there is you know the cost is astronomical the power demand is huge they can't build power they can't build 26 gigawatts worth power plants you know you know there's all the incentive in the world for this to become much more environmentally friendly over time 100 no for sure and i think there's there's all all those amazing reasons to keep to keep the power down i think you know we designing from that way up i think is is making a huge difference and we we want to make sure that our our whole our whole stack um our whole ecosystem is as is as efficient as possible for those reasons exactly yeah so uh let's maybe like zoom this out a little bit more so um you're you're kind of what i'm hearing is that like you're pretty bullish on not necessarily needing bigger and bigger models or maybe many more smarter more efficient more niche models uh or do you think it's a combo of of the two like where do you think this is all heading yeah so i think this is kind of super interesting right when i talk about the you know where i would say where the future of ai agents and ai applications in general so yeah you i do think that there is a there is a mixture of both i think the the common pattern that we're going to start to see is that you're going to feel like okay how can i make this work So let's take an example.
37:18It's one of the best OSS models at the moment. You've got GPT-4 OSS, you've got DeepSeek R1. These models are frontier level models. They cost more to host, but their levels of intelligence are great and their tokens per second are kind of manageable. And then you've got these other models that are kind of like mid-tier that work really well. Now, I think as more and more education comes through, people will start off with that top level, you know, the more frontier level stuff, just to prove out the use case, right? So like, hey, does it work? And then as they get more comfortable, and we're starting to see this already, they will then start to swap out these components.
38:00They might fine tune it themselves, or they might, quite frankly, go and find a little bit of a better model. like you know might say oh actually this you know we've we might have a a call center agent that doesn't need that we've actually got a voice model that does this at a tenth of the cost and you know we can host on your hardware let's let's do that so i i do to answer your question about i do think we're going to need both i think as yeah and um as the enterprise evolves they will start to see the need for both as well and that's what we're already starting to see as well i think most enterprises want to see now does it work you know let's let's see if we can actually make this use case work once they do that they they will then likely to be like okay how do we make this work for a for a more effective cost so i i definitely see in the complex complexity of applications so to give you an example you know let's take voice or let's just take healthcare or something you're going to need more customizable models um the question is at what point does the enterprise say okay you know this is this is good enough you know and you know this open source model is out and we can put it on and that is definitely a and that's always going to be a question right i always see this is if you work in an industry where the the bar is higher you know you might be a bit less, you might be a bit less kind of, I would say, willing to go with a open source model.
39:30But as we're seeing the gaps between open source and Frontier closing so fast that I think within maybe, you know, 18 months, you're going to start to see people being as comfortable moving to these smaller models as they are using the Frontier model. So I would say those big models, I think they're going to be there, but we're already seeing a switch to these smaller ones.
39:51Kwasi Ankomah:yeah yeah what would you say in your opinion right now is the best open model available it depends and i know is that such a cop-out answer but it's it really does and i think we like someone over we were kind of agnostic right we want to give you the best model for your use case and you know we for instance oss from it's a fantastic yeah it's a fantastic model considering its footprint 120 billion parameters it's you know it's it's it's smaller right smaller than some of its other frontier contenders and that makes that makes it more yeah exactly that makes it more accessible to customers and we're seeing that as well you people can say oh hey this is actually i can i can run this even faster on your platform which means that i can run it maybe on more even real-time applications i think that you know we've also got the quen coda models which are absolutely phenomenal you know like the full you know those are the 480 i think billion parameter model as well working really really well but i think and it's also kind of based on the open route to traffic like i think deep seat v3 has really hit a mark and also maverick as well like maverick because it's vision i know people people weren't kind of super happy with the reasoning performance but i i would also kind of you know counter that with it is a multimodal model right we've seen it from we've seen it work really well in terms of you know especially the the larger 128b variation right mixture of experts handling computer vision so again that's a very good model if you've got a vision problem like we when we set up some of these architectures we we use maverick for kind of like um predictive maintenance and stuff like that it does a really really good job um but i think the one the one that i think people has captured people's imagination and they they're able to because they're kind of just comparing it to frontier right has been probably v3 right v3 or r1 because r1 it's kind of you know it's reasoning expertise but v3 is just it's just a very strong and consistent model and i think you know it works to the point where you know if you did a kind of a blind test you'd be you'd struggle to see the difference in its output between especially the new 3.1 that's just come out as well um and you know 0528 those they're they're really pushing to you know they're really pushing at the door of the frontier models saying hey we this is we are pitching ourselves directly against you yeah um so it's a hard one to answer because they are you know they're but those are the ones to me that stand out and i think you know from what we're seeing from our clients and what i'm seeing when building agents those are the ones those are the models that really really do give kind of stellar stellar performance is samba nova ever going to create your own models do have you considered that or have you created them um and or are you just staying up focused on infrastructure um what do you think yeah we're kind of focusing on the kind of inference side of things and that but you know there's so many amazing model model providers and you know it's we want to focus our energy on making them as efficient as possible and making sure they run as fast as possible yeah we'll like we'll stay kind of on the on the infrastructure and um optimization side of things i think it's a smarter business move i really do you know when you consider what goes into building models and uh and that you're chasing it's not a simple task light years ahead already you know or or like anthropics uh whole thing where dariel almede was basically like yeah you know it's basically each model has its own business so it's like yeah maybe you train a model for 200 million and that's profitable but then you have to go make another business trade a model for a billion and then you have to make that profitable yeah it's it's it's very very super cost and as you said it needs needs a phenomenal team and what i'm really what really we're really bullish about is just the amount of quality in the open source you know like the labs that are coming out you know miss like mistral deep sea llama like you've got phenomenal open source models that are coming out.
43:56And I do think it's going to be even more in the next 12 to 18 months. I'll be surprised if we don't see other entrants coming into the market. I agree. That's exciting for us because we're just like, keep it coming because we're agnostic and we just want our customers to be able to have the best choice. And I'm super excited with someone who loves open source to see every day I read a new paper and someone's just brought out a model. I'm like, it's the the the buoyancy and the kind of enthusiasm and the open source is amazing i'm
44:23Kwasi Ankomah:scrolling through the hugging face model releases every day and you're seeing seeing what's new and what's dropping and what people are doing with with ones that exist and uh there's always something it's crazy and it's it's getting there's always something it's yeah and what's great about that is you know if you think about use cases right this is where i was if you think about use cases this becomes so important you've got you've got models that are coming out you know that are that are training for specific use cases now you know people coming out with this or specific language and this is such a big thing you know one of the things that we talked about i think is we have we're working with you know a leading voice provider who kind of picked us because of our of our of our latency budget and it's amazing to see you know the custom models they're building and what other people are building just for voice.
45:16And I just think in the next 12 to 18 months, you're going to see, oh, I want to go and build a agent to do case management in fraud risk software. Now that seems super niche. And I'm sure there won't be a model that does exactly that. But as we go on, you'll start to see these models narrow down so that there'll be a day when you go onto Hugging Face and you'll be like, oh, I need this one. Cool. you know and that that is super super powerful because that again will will allow people to just be like hey i can use this model and it works for my use case as well as anything else great it's sort of like the dream of the custom gpt marketplace but i just like custom open source models everywhere yeah 100 yeah yeah they perform super well as well yeah they do looking ahead 12 months what should our audience be watching out for in terms of ai infrastructure like what are the signals that this like inference bottleneck is getting solved like is it just samba nova crushing it like what what do you think is um yeah what do you think the next 12 months we should be looking for you know there's a few things people should be looking out for i think they should be looking out for the almost like the what we talked about the kind of training inference right so like people are like we're now optimizing for for inference and i think the companies who are largely going to be able to figure that out, we'll be able to kind of own that next generation of AI applications.
46:41Like we, of course, we think, you know, training is still going to have a big part, but we also see this new kind of resolved interest as AI grows and more and more people just run inference, right? You know, you have companies that were not, you know, that didn't deploy any AI applications in 2024 who have now started to deploy AI applications. And that, you know, if that keeps happening as it is, that market's going to go even further right and then the the next is kind of like the agents as infrastructure right like you know this is this is something that i think about a lot like how it's not just like ai systems right ai systems are useful and they're going to be wrong but i'm talking about something that's actually going to go like something that's scalable as a as a as a database and a web server you know as an agent like and that is gonna be again people who figure that out that's a bit what i spend a lot of my time thinking about that like too much time right how do i how do i scale this to be something that is as reliable and as consistent as a normal ai application that's where a lot of that is yeah do you think that's a memory uh problem or like Like what, I mean, we jokingly ask like both hard questions that are like trillion dollar questions, but like, do you think that, like if you could solve it, you're probably in the money, but do you think that's like a memory problem that makes them more consistent?
48:08What do you think is the true thing there? It's a great question, right? There's a couple of things that we need to think about. Firstly, it's the memory is a big one. How do we ensure that the agent is learning over time, right? That's like we need to make it more flexible in its learning. And that in itself is a whole area. What memory do you attach it to? How much do you keep? And also where it gets to a point where the agent has too many memory and too many tools and gets confused. So those are probably big things in terms of memory. But one of the things that I think about a lot is in terms of the kind of the tool calling and the guardrails around that right like one one of the big things with agents was let's give them all the tools right and now we've got some point where we can give them all the tools but as anyone who's built agents in production will find you give your agent all the tools it might have a problem picking which tools it's got you all should use and people figuring out that balance for their between building like a generalist agent that does everything to like an agent that basically manages your call center is going to be a really big shift like how the people who were and that always is the business expertise right what what are we trying to do and then the third thing i think is i would say just almost world-class real-time observability the same way that you do have with database if you think about how much you can observe about a database like it's it's crazy like i used to you know anything any transaction at the row level that all of the index speeds, every little thing about that database is measured.
49:54And I think in agents, we're at the stage where people are like, oh, we measure the answer, the end-to-end. But as you can see, everything in that step, to give you an example of our code calling agent, whether or not it calls the agent, does it call it with the right tools? What was the result of the tools? Like there is so much that people need to start digging into to ensure that it is super, super reliable. And I think that people who figure out that almost, you know, nano level of a real-time observation is going to be super, super huge. Then I think the last thing is just the element of humans in the loop, right?
50:34That's going to be, the key one is around kind of almost like trust, right? Like every, we all talk about, you know, We talk about agents being able to work autonomously, but we know that there's going to be a big behavioral shift. If you want something as agent as infrastructure, there needs to be a step where there's a... Because, again, we all know that web servers and databases have database admins. They don't just sit there on their own. There's very specialized people who oversee the database. And I think in order to get that level of trust, you're going to have to be comfortable being like, okay can can the agent return control to a human in case of any failure and things like that so i think those are kind of you know like these are the key things that honestly how do we scale up like you know like again can it is it able to have memory and improve you know be that reinforcement learning memory is it able to call all the right tools and can we get to a place where you are confident in that human in the loop infrastructure because that that's the difference now and i think the people the companies by the way who have done it really well have done those three things
Read the full transcript
51:42Kwasi Ankomah:they have really well really well well quasi thank you so much for joining us today this has been super interesting uh learning more about thank you no it's been a pleasure learning more about samba nova what you guys are doing and i'm anxious to keep up and follow it because this is really really intriguing and you're feeling a feeling a large niche that's it's very important if if our readers wanted to learn more about SambaNova, keep up with you. What's the best way to do that? Yeah, that's a very good question. So we would always say to people, head to cloud.sambaNova.ai to get started. We've got some credits on there.
52:19You can go in, you can see the fast inference for yourself. We'll give you all the information about how to connect to our models and our ecosystems. We've got all of our integrations in there. So that's definitely the best place to kind of keep up with what we're doing and of course go to samro.ai if you want to learn a little bit more about the company and the other products that we've got if you want to keep up with me personally you can follow me on linkedin um i'm always always talking about agents to anyone who will
52:44Kwasi Ankomah:listen i just cased you down i love it i love it well to those of you watching who haven't uh please take a minute to like subscribe to the channel we really appreciate you taking the time to learn to listen and hit us up with questions in the comments if you have them we'll do our best to to answer or get answers where we can't. And that's it for today, folks. Farewell for now, humans.
From the publisher
Everyone's talking about the AI datacenter boom right now. Billion dollar deals here, hundred billion dollar deals there. Well, why do data centers matter? It turns out, AI inference (actually calling the AI and running it) is the hidden bottleneck slowing down every AI application you use (and new stuff yet to be released).
In this episode, Kwasi Ankomah from SambaNova Systems explains why running AI models efficiently matters more than you think, how their revolutionary chip architecture delivers 700+ tokens per second, and why AI agents are about to make this problem 10x worse.
💡 This episode is sponsored by Gladia's Solaria - the speech-to-text API built for real-world voice AI. With sub-270ms latency, 100+ languages supported, and 94% accuracy even in noisy environments, it's the backbone powering voice agents that actually work. Learn more at gladia.io/solaria
🔗 Key Links:
• SambaNova Cloud: https://cloud.sambanova.ai
• Check out Solaria speech to text API: https://www.gladia.io/solaria
• Subscribe to The Neuron newsletter: https://theneuron.ai
🎯 What You'll Learn:
• Why inference speed matters more than model size
• How SambaNova runs massive models on 90% less power
• Why AI agents use 10-20x more tokens
• The best open source models right now
• What to watch for in AI infrastructure
➤ CHAPTERS
Timecode - Chapter Title
0:00 - Intro
2:14 - What is AI Inference?
3:19 - Why Inference is the Real Challenge
9:18 - A message from our sponsor, Gladia Solaria
10:16 - The 95% ROI Problem Discussion
13:47 - SambaNova's Revolutionary Chip Architecture
15:19 - Running DeepSeek's 670B Parameter Models
18:11 - Developer Experience & Platform
21:26 - AI Agents and the Token Explosion
24:33 - Model Swapping and Cost Optimization
31:30 - Energy Efficiency 10kW vs 100kW
36:13 - Future of AI Models Bigger vs Smaller
39:24 - Best Open Source Models Right Now
46:01 - AI Infrastructure Next 12 Months
47:09 - Agents as Infrastructure
50:28 - Human-in-the-Loop and Trust
52:55 - Closing and Resources
Article Written by: Grant Harvey
Hosted by: Corey Noles and Grant Harvey
Guest: Kwasi Ankomah
Published by: Manique Santos
Edited by: Adrian Vallinan
