In short
The TWIML AI Podcast - Episode #757 Summary: Scaling Agentic Inference Across Heterogeneous Compute with Zain Asgar
Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington interviews Zain Asgar, co-founder and CEO of Gimlet Labs. They explore the challenges and solutions related to heterogeneous AI inference across diverse hardware, particularly in the context of agentic AI systems that consume significantly more tokens compared to traditional language model applications.
Key Points
Background of Zain Asgar
- Co-founder and CEO of Gimlet Labs.
- Adjunct professor of computer science at Stanford University.
- Previous roles include general manager at New Relic and experience at NVIDIA focusing on efficient compute in large-scale clusters.
The Vision of Gimlet Labs
- Gimlet focuses on making AI workloads at least 10 times more efficient.
- The company addresses the challenge posed by rising token consumption in agentic AI applications.
- Initial goals included optimizing models for smaller hardware devices before shifting focus to large-scale data centers for greater impact.
Heterogeneous Inference
- Current industry practices often rely on high-end GPUs, which Zain argues is unsustainable.
- Gimlet proposes disaggregating workloads across a mix of hardware types—ranging from cutting-edge GPUs to older CPUs—to optimize cost and performance.
- The architecture involves a "three-layer cake":
- Workload Disaggregation - Breaking down workloads into granular components for optimization.
- Compilation Layer - Mapping models to specific hardware targets for enhanced efficiency.
- Kernel Optimization System - An innovative system leveraging LLMs to autonomously rewrite and optimize compute kernels.
Networking and Performance Challenges
- The complexity of networking in heterogeneous environments raises issues related to cost of compute, memory bandwidth, and latency.
- Potential trade-offs between numerical precision and application accuracy are critical considerations.
Real-Time Optimization
- Discusses the dual approach of fast-path deployment and ongoing post-deployment profiling to adjust workload allocations.
- Emphasizes the importance of observability for making informed decisions about resource utilization.
Future Directions and Market Dynamics
- Zain highlights the expectation of a sustainable shift toward heterogeneous systems, moving away from supercomputer-like integrations.
- Anticipates a future where AI applications can efficiently use a mix of high-performance and commodity hardware, especially in data centers.
Customer Deployment Insights
- Gimlet is primarily targeting data center customers with a focus on self-hosted environments rather than general-purpose cloud offerings.
- The company’s technology can handle various workloads and optimize costs, demonstrating significant advantages in cost-per-token metrics.
Conclusion The conversation ends on a note about Gimlet's upcoming product launches aimed at developers, offering tools to orchestrate large-scale agentic workloads efficiently. Zain expresses excitement about empowering developers to run models in a more cost-effective and faster manner.
Key Takeaways
- The need for sustainable AI workloads and cost-efficiency drives Gimlet's innovations in heterogeneous compute.
- A structured approach to workload disaggregation, compilation, and kernel optimization can yield significant performance improvements across diverse hardware.
- Observability and real-time adjustments are vital for optimizing performance in complex environments.
- The future of AI infrastructure may lean towards leveraging a mix of high-end and commodity hardware to meet diverse application needs.
For more detailed information, the complete show notes for this episode can be found at [the TWIML AI website](https://twimlai.com/go/757).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This podcast is sponsored by Google.
0:33Join developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and more than 75 other supporting companies to build the open tool stack for multi-agent software and trusted agent identity on Agency. Agency, which I recently discussed on the podcast in my interview with Vijoy Pandey, is now an open source Linux foundation project where you can help create the protocols, specs and tools that power next gen AI infrastructure. Visit agency.org to learn more and join the build. That's A-G-N-T-C-Y dot O-R-G. If you take a look at training hardware, it's kind of gone the way of like building supercomputers, right?
1:20Like, you know, people don't talk about building machines anymore. They're like, here's my entire rack, right? This starts looking like, you know, what Cray was doing. So in some ways, you know, you could be like, oh, we kind of regressed back to the supercomputer era. And I don't know if I use that word positively, right? We're building these like fully vertically integrated systems. I'm not sure that's the route for inference. I think inference is much better served as a large-scale work cloud where you can utilize a bunch of relatively commodity hardware and be able to scale out efficiently.
1:58All right, everyone. Welcome to another episode of the Twin Wall AI podcast. I'm your host, Sam Sherrington. Today, I'm joined by Zane Asgar. Zane is co-founder and CEO at Gimlet Labs and an adjunct professor of computer science at Stanford University. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Zane, welcome to the podcast. Hi, Sam. Thanks for having me here. Super excited to be here. Excited to have you on the show and looking forward to digging into our conversation. We'll be talking about the work you're doing around heterogeneous inference for agentic systems.
2:32to get us going. I'd love to have you share a little bit about your background. As I mentioned, I'm co-founder and CEO of Gimlet and also adjunct faculty of computer science at Stanford. Prior to this, I was a general manager at New Relic through an acquisition of my previous startup, Pixie. And actually, you know, a bunch of people from Pixie are now at Gimlet as well. I was an EIR benchmark capital where the idea for Pixie came out of. I was in Google research and spent a lot of time at NVIDIA. So kind of focus on like efficient compute and being able to orchestrate and run compute efficiently on large scale clusters.
3:06And where did the idea for Gimlet come from? What's your what are you going after there? So when we started Gimlet a couple of years ago, we had a we had a focus on like, how do we actually make AI workloads, you know, at least 10 times more efficient. Right. And part of the part of the challenge over here has been that you've seen this like huge explosion in AI, AI workloads, especially around agentic AI, where you're consuming like 10x more more tokens. And really, if you want to be able to keep this somewhat sustainable, you need to have these big leaps on improvements. So that was our original focus with Gimlet.
3:38And when we started off, we were really thinking about how do we get models to run on things like your laptop and, you know, Raspberry Pis or whatever, right? Like small scale hardware and how do we get like the best efficiency? Exactly. Any kind of edge device. But one of the things we kind of realized is that, you know, we built up our stack to work on this very, very heterogeneous system, right? Because like, you know, two MacBooks look very different than, you know, typical Windows laptops running like Intel or AMD CPUs, which look very different than like an iPhone. And we got really good at running models efficiently.
4:10So we kind of realized that this technology that we've built up is actually pretty applicable to data center scale systems. And, you know, there's this large, large scale problem. And so we could have a much larger impact if we actually improve that system. and our team decided to start going after that space instead of directly targeting edge devices. Partly because we think that the edge market still needs a couple of years to really materialize. Whereas the data center market is crazy hot right now. Exactly, exactly. And I think the fact that we can basically orchestrate and run stuff across many different hardware to optimize like the unit economics and also get best in class performance is like a big, you know, big benefit over there.
4:53One of the things that you guys talk about is specializing in compute for agentic AI workloads. Does that just mean LLMs and multimodal models? Or is there something specific about agentic that you think needs to be considered in the hardware? So one of the things, you know, we started to build for is like, how do you run like entire applications, right? Like today people run agentic frameworks. those frameworks typically call out to different apis to run models which lands a meaning you spend a bunch of time doing all these api calls and orchestrating api calls we kind of wanted to take the approach of like well what if you could avoid all the round trips and put everything in a single system and agentic systems as a whole are very heterogeneous right because there's a bunch of like you know cpu compute happening there is a bunch of you know things like database calls etc etc and there's a lot of like model llm model execution and if we can better orchestrate all of those and understand when and where things are going to be run, we can both improve the efficiency and the latency of the system, but also be able to make optimizations within the models themselves because we know how they're going to get used.
5:57Got it. And what kind of optimizations are you typically looking to make? Yeah. So one of the things we do, one of the biggest things we do from an orchestration perspective is just like right sizing the hardware, right? Today, you know, general consensus is that people will go buy like the highest hardware they can, which, you know, is basically a B200. And you'll basically maximize how many ever B200s you have to run your model workloads. One of the things we do is actually do pretty fine-grained partitioning of all your models and including your entire, like, you know, we think about an agent as an entire like data flow graph.
6:31So how do you actually like break up the data flow graph? How do you break up the models so you can assign the most performance critical pieces to the highest end hardware and then offload the less performance critical pieces to hardware that will be better utilized by running that workload? What you're describing reminds me of a conversation that I had with Lin Chow at Fireworks about kind of the way that they're approaching distributing inference across their cloud environment. I imagine that, you know, anyone who's dealing with inference at large scale is going to be thinking about how to do that kind of distribution.
7:05Does the heterogeneity add a particularly, you know, interesting element to it? So once you have like any kind of large scale, so people typically talk about, you know, things like pre-fill, decode, disaggregation, where you run pre-fill on one set of nodes, decode on another set of nodes. When you started looking at heterogeneity, especially if you look on heterogeneity across vendors, right? This problem exists both within a vendor and even across vendors. You start having to think about a lot more of the trade-offs of like, you know, what is my cost of compute? What is my cost of memory bandwidth?
7:35What is my cost of memory capacity? And then, you know, it's basically this giant optimization problem of like figuring out where to put which workload in order to optimize the cost per token. And is that something that you are thinking about on a, like a real-time basis, like a runtime basis? Or is this more of a deploy time consideration where it is, you know, when you're deploying an agent, it gets kind of introspected and you determine what should go where at that deploy time? It's a little bit of both because when you first deploy it, you'll probably have a good estimate of where things are.
8:11But it's pretty, most of these workloads are complicated enough that it's pretty hard to know exactly where things are going to land. So there is a whole bunch of profiling that goes on after the fact. So like observability and stuff kicks in. And then you get a better idea of how to change the allocations in order to improve the performance. So in our system, we actually have two paths. There's the fast path where we just try to get something up and running. And then there is another path which tries to move the workloads around to optimize the utilization of the hardware. And then on an API perspective, there's a whole bunch of like routing decisions that need to get made because, you know, you need to know where the data is available, which cache, what's not cached.
8:48And, you know, where the models are loaded, you might be offloading models in the CPU memory and things like that. So you need to know like the cost dynamics there on a real time basis. I'm imagining a kind of an optimization problem or a control loop where you've got like a Kubernetes cluster with a bunch of different stuff in it. you coming from New Relic, you mentioned observability, you've got a bunch of monitoring types of tools and you're kind of mapping all that to some kind of costs, factors and using that to moving things around in that kind of environment. Is that kind of close to the way you have mapped things out?
9:23Yeah, that's actually not far from how our thing works, right? We run on top of, we run completely on top of Kubernetes and, you know, we actually use some pretty cool new stuff like DRA and Kubernetes to really help like orchestrate some of these workloads. Oh, what's DRA? It's the resource allocator, dynamic resource allocator. So basically it allows you to like, you know, be like, I'll have this many slices of capacity available. Like how do I actually route to? Oh, nice. Yeah. So we can actually like, you know, instead of like thinking of a GPU as like a full resource, we could be like, oh, here's like a quarter GPU worth of work.
9:56Oh, interesting. I'm trying to remember the name of the company that used to do Run AI. Yeah, they were acquired by Intel, I think? NVIDIA. NVIDIA, yeah. And so now that technology is more readily available? Correct, but it's really not supported on GPU hardware per se. So one of the things we do is actually really make it easy to run workloads on GPUs and other accelerators while partitioning it into these segments. But we rely on Kubernetes to help orchestrate it because we don't want to rebuild everything, right? So we want to rebuild the pieces that matter and then rely on the high-level frameworks to provide us the hooks necessary.
10:31So does Kubernetes DRA already support GPU partitioning or is it a broad framework and you have to plug in the GPU bits? It's a pretty broad framework. It's more just for doing like dynamic resource allocation. So we basically put in the GPU bits over there. And DRA itself is actually not a mint GA in Kubernetes yet. Okay. It's supposed to get released in the next version, I believe. Okay. So the challenge that really comes in, I think, when you deal with heterogeneous stuff is, you know, there's kind of like two things. Orchestrating on heterogeneous hardware is more challenging because the networking fabrics and stuff could be pretty different.
11:11So you have to figure out like how, you know, these, you know, because people are typically running very high speed network fabrics or like 200 gigabit, 400 gigabit, 800 gigabit Ethernet. Sometimes they're running like, you know, InfiniBand or something more proprietary, like even EnvyLink across like a rack. and you have to kind of figure out how you can actually orchestrate across these different fabrics. And every single hardware has, you know, a different programming model, right? They're not, you know, most of them, obviously NVIDIA uses KUDO, but there was a whole bunch of other factors if you're targeting like Intel and AMD hardware.
11:42Yeah, talk a little bit more about why that networking heterogeneity causes challenges. I think, you know, folks that have been in the space long enough, think back to the OSI seven layer stack. And at some point, we should be abstracted from all that underlying detail. But that is not the case here. I think we're getting there slowly, right? Like, for example, you know, at Gimlet, most of our stuff runs on top of this thing called Rocky, which is RDMA over a converged Ethernet. Right. RDMA was like doing remote DMA. Yeah, remote DMA. Yeah. Yeah. Remote direct memory access. So it basically tries to emulate like memory semantics over a network interface.
12:23and you really need stuff like this because you're like oh i want to move data between two different accelerators right and the fastest way to do it is to be like oh here's this chunk of memory i'm just going to send it over to to this other machine um the the challenge you typically run into with this type of stuff is you know um you don't usually want to traverse the cpu while doing it um right because you don't want to have to have the cpu do like a you know a call to the GPU, capture the stuff in the CPU memory and then copy it out. So there are things like GPU direct that actually allow you to do this and are pretty mature in the NVIDIA ecosystem.
13:02And they're usually relatively mature if you go within like a single vendor, right? Like if you want to say, I want to copy from AMD to AMD, it's like relatively straightforward. The problem you run into is like when you want to copy from one machine to another, they may not be able to transparently do it. So there is a whole bunch of like, you know, systems engineering that happens over there. Even if you're trying to transfer caches and stuff, they may not even be in the same formats, which means I need to apply some transposes and move data around in order to actually get it to properly map to other hardware.
13:34And all of this is in a spot where milliseconds matter or maybe even microseconds matter. Gimlet operates Gimlet Cloud, but it also sounds like a big part of the heterogeneity that you encounter has to do with disparate customer environments. What's the split that you tend to see between usage on your cloud versus usage in customer environments? Is there a focus? Is one more of a focus to another? How do you think about that? Gimlet, you know, as of today, we haven't publicly launched our cloud product. We have like a handful of like early access customers on our actual cloud product. Almost all of our sales and we do, you know, we do about eight figures of revenue right now have been through data center, data center customers and partners on the semiconductor side.
14:27And over there, it's mostly self-hosted. Eight figures is not bad for our company that launched two weeks ago. Yeah. Well, you know, we've been working on it for a little while under the radar. Talk a little bit about the licensing model. Like, is that primarily software licensing? Is that also hardware provision? Is that, you know, how does that all work? So today we've been mostly, you know, going to like, you know, some of the countries and data center side partners, and that's been a software license. And it's usually restricted to some number of nodes or there's an additional percentage licensing fee based on how much hardware you utilize.
15:01And then on the cloud side, when we were planning to launch our product in Q1. That'll be much more of like a usage-based pricing. So we've talked a little bit about, you know, some of the technical challenges, networking and heterogeneity. Like how does that lead? Those all sound like, you know, areas for compromise, but you're also targeting, you know, you mentioned 10X performance as kind of an initial vision. you mentioned better performance and better cost per token like how do you uh overcome you know those those challenges those compromises and get to something that's actually better so we think about you know at the highest levels we think about our stack kind of as this like three-layer cake right so you give us like the agent you know agent graph which is basically a blow and as of today we can basically absorb you know things like you know lang chain or even people have written like python code around like hudden face or whatever right like we can absorb that stuff into our system directly we don't have like our own like agentic interface or something we expect people to come in with their own um uh own workflows so we want to meet developers where they're at and do developers need to like call some libraries if they're like in a pure python non-lang chain type of environment or yeah when you import gem lead we do a lot of um a lot of our stuff relies on like doing tracing through the python code um so it is currently only usable through python that might change in the future uh but as of today you know it's basically stuff written around python or around like other other libraries where we can generate a graph out of it so it might be analogous to how you might use a langsmith or something like that where you're decorating functions that call out to LLMs, that kind of thing?
16:48Right. Except on our side, we expect people to just use whatever system they're using and we intercept the calls and figure out how to map into our system. Okay. So we don't have like our own like decorators. Okay. Got it. Got it. And so you're intercepting the calls at the Python function level. and is the implication that you're, you know, then potentially, you know, the natural kind of place for that to run is like on a local, you know, GPU, et cetera, but you're kind of, you know, remoting that out to wherever it needs to be running? Correct. So when you upload it into our system, we'll basically run through it and be like, okay, here's the graph that we know.
17:30And then we'll separate out the models to be able to go run on GPU resources. and you know we know like for example the important model from like hugging face or something we'll know exactly what model it is and bring it into our system and it's all pretty pretty automated like our goal is you know five minutes from like hey i'll have this like thing that runs locally to i want the scalable in the cloud yeah and and are you limited to specific models to publish models or if you've developed some custom model can that be used as well Yeah, so our whole system is designed to be able to run on any type of model.
18:06We actually support more than LLMs in our stack as well. We have support for several different types of models. And so that kind of goes into how we architected the system. So once we have the agent graph, right? The first layer that we do is we think about as workload disaggregation, which is how do we actually figure out what are the granular components of this workload and how to split it up in some meaningful way. And this is actually where, you know, the vast majority of benefits come in because knowing what hardware is available and how to split it up, you can go ahead and start like packing these things in pretty tightly into the different resources.
18:41And this is also where the cost modeling comes in. You start thinking about, you know, what is my cost of compute on all these different hardware? What's my cost of memory bandwidth? What is my cost of memory capacity? And you start figuring out what are the critical resources for each one of these sub pieces of the workload. And then you can then allocate it to the lowest cost resource right and of course all this is within some sla bound so you know like like very technical terms like this lines up as some like big convex optimization problem right with a bunch of different assumptions and you try to just like optimize this down and figure out here's the optimal allocation um in real life it's very unlikely you'll be able to get the optimal allocation because the hardware resources you want may not be available so you'll just try to you know find the best fit um where you deviate from the optimal um the next layer down is that we have our compilation system which will basically take all these workloads and lower it to the target hardware um and when you do that this is where we do all the optimization around like okay well this stuff runs on like a cpu this stuff runs on a gpu and we'll you know go and compile that stuff down the challenge you typically run into with any kind of you know generic compiler framework is that usually land up with like, at least in the AI space, you land up with like relatively unoptimized like kernels.
19:54So we have a framework underneath it where we actually autonomously rewrite your kernels using LLMs in the background to try to generate like more optimized code to run on like every, you know, like to run most efficiently on AMD hardware or Intel hardware or NVIDIA hardware. You said you're using LLMs to do the compilation? So for the first layer in the model compiler, layer. We used MLIR plus Torch MLIR and LLBM to do the compilation. And then the layer below that, we have like automatic kernel synthesis, which is basically like take all the compute code that's running and how do we generate better compute code.
20:34And that's actually done using LLMs. And did you have to post-train a model to do that or fine-tune? It's all post-training. So can you talk a little bit about that process? Yeah, yeah. So actually there, for reference for the readers, for the watchers, sorry, there are actually a few blog posts on our blog that talk about this in lots and lots of detail. But at a high level, it's like a multi-agent system, right? So it's like a Gimlet-hosted multi-agent system in our own stack. But basically what it does is there's a part, there's like a supervisor-based model which tries to generate new kernels, right?
21:13So you give it your PyTorch input. In our case, it's usually PyTorch. That's the only code we really, really optimize. So you give us the PyTorch input, and then we can give it reference code if we have it available. So, for example, we may have CUDA reference code available, right? Because there may be a CUDA version of that kernel available. If not, it doesn't matter. It's not required. It's a hardware in the loop system, which means that the first agent will actually go and generate a whole bunch of candidate kernels. we will then go run this candidate kernel on the target hardware um and we'll do all the profiling and correctness checking uh because typically you know when you generate new kernels there's there's always challenges with correctness so you need to go run all the tests make sure it's correct we capture the profiling data and then there's another part that basically looks at the profiling data and says okay try these new techniques to optimize the kernels further right and then the first generation part again will go regenerate new kernels based on the you know optimization criteria that it thinks it needs to do and then you kind of repeat this loop a few times and it converges um and if you get a faster kernel you use it if not you just fall back to the previous implementation interesting and is the correctness testing that you're doing is it um you know does it run or is it you know somehow faithful to the original uh the original pytorch code i guess i'm I'm imagining that you could produce valid kernels that run, but somehow have slight deviations from the original algorithm.
22:50Yeah. So what we do over there is that we actually run through a bunch of numerical tests to make sure that the kernels are numerically equivalent. The caveat over here is that... Doing some kind of sampling of the input space and using that to determine. okay um the challenge over here is that it's basically impossible well maybe that's a strong word but it's very very hard to generate a completely numerical equivalent kernel because of the way floating point math works right um and uh you know this is actually part of the challenge because you look at this problem you're like oh we can just prove that they're equivalent right but it's actually not that easy because you know if a times b is not the same thing as b times a is certain things start looking really really weird right and that's where you are about floating point math.
23:37And because of that small precision issues, it actually makes this problem pretty difficult. So we're actually working on some research
23:47to show what is the impact? How do we understand what the impact is of accumulating these inaccuracies in the loop? Does that imply that some type of quantization or quantized models or something like that, it would be more easily optimized because you don't have the floating point issues or are they two different areas of floating point? Well, I think a lot of people quantize into like a lower precision floating point now, like FP4 or whatever. So like ultimately you still have those issues, but the good thing about FP4 is there are not that many possible values. So it's actually a little bit easier to verify.
24:30Obviously the higher the precision, the harder it is for us to verify correctness and knowing the exact bounds um so when we do the kernel generation we typically what we do typically is that we cache all the kernel gens offline for like a whole list of models and then verify them and make sure to put them back in um so we try not to do it in real time because we're still afraid of like potential numerical errors in the system yeah it it strikes me that like if you were talking to me as a customer and explain this and like i'm not sure that i want an llm in my you know compilation and optimization step just given hallucinations and and all that kind of stuff and you know given that you you know that they're difficult to verify like that seems like that would uh that would discourage me do you find that like well so normally we tell people that we can generate them optimized kernels and people are pretty excited because they hear about this new technology right but like I said you know we tell people hey don't worry about it because we have basically run it offline and made sure that these kernels are correct because we're not ready to run it like you know the third layer of cake doesn't run into in the loop right now and so the argument then is that the the sampling based verification you know you feel like there's sufficient coverage and you can flesh out any issues.
25:56Yeah. And at the end of the day, we basically, you know, once we are running a set of models, like we always verify, you know, some set of models, right, and make sure they're correct. And you can just get the performance scores on the model and see if they still match. And so we got as far as the second layer of the cake, did we get the third layer of the cake yet? The model compiler itself is fully, you know, LL compiler based, right? So it'll never generate code that's not accurate. it we did get to the third layer which is the kernel optimization that happens automatically um and to be honest the way i would think about this is that you know what would happen before is that we would have our compiler our compiler would you know spit out a bunch of code and then our team would go by and be like oh you know we think we can write this to be even faster right um the problem is that you know it's easy to do this when you're when you target like one hardware like one piece of hardware it becomes like untenable our team's not very large right It becomes untenable when you start thinking about targeting hardware from three different vendors and multiple generations of hardware in each vendor.
26:53So then you start thinking more creatively. And we're like, oh, look, these are the steps that I take to optimize the kernel. I write a more optimized kernel, run some benchmarks on it, figure out why it's not as performant as I think it should be, make some changes, keep doing this until I get a good kernel. So we're like, hey, why not build an agent that does this? And if you like the results, we keep it. if we don't like the results, then we will go write it manually. So ultimately, you know, while we talk about these three layers of the cake, the third layer of the cake is really how do we cut down the amount of work we need to do in order to get a functioning system at large scale.
27:28And over time, I think we'll start thinking about like, how do you apply AI to systems as a whole? Right. But I think we're not, I think you're right in the sense that if you go start telling people all these decisions are getting made by LLMs, they're probably gonna get a little upset. so you know we don't really run the llms in the live loop today but we do see that changing over time as as you know better guardrails and better testing infrastructure comes online what are some examples of ways that you would see them as being useful so i think there's an argument to be made for example that even in our planning system where we figure out how we're going to orchestrate these workloads we could have like an llm help in the planning right like Like right now we have all these rules and all these kind of rules and heuristics.
28:11Yeah, exactly. You can start thinking about there being some kind of like, you know, like a tool, like an MCP tool or something that says, oh, I will tell you exactly what the performance of this hardware is going to be for this workload. And you can build this out on like a much more structure, like, you know, LOM goes and examines all these things and then, you know, figures out how to plan this stuff out. We aren't that sophisticated yet, right? Most of the planning is done using just like, you know, heuristics and optimization criteria. but it'd be pretty cool to see like you know systems that can actually use all the live information to to do this over time yeah i'm thinking about the the uh like tco advantages that you're that you're targeting like better cost cost per token like have you do you have a sense for like how much of that comes from each of these layers of the cake Like how much of it, you know, if we just disaggregated the workloads and kind of put them to the, you know, the most efficient or the best cost hardware for that workload, like that would be 60 % of the benefit.
Read the full transcript
29:12And then compilation is 30 and kernel is 10. Or is it like, you know, the other way around, like, you know, 20, 40, 40, something like that. yeah so i think it really really depends on what hardware you're targeting right because if you're targeting something like um like say like you know at the end of the day you're thinking about just you just think about the model work that's easier easier to think about and that's the vast majority to compute if you think about something like that and you're trying to run on top of h100s right the code is like very very well optimized at this point which means that at the kernel level you're just not going to get all that much right because it's been optimized for for years at this point um and like ultimately you know we can talk about how machines might be able to do a better job but you know hundreds of people hammering on something for years will probably get you pretty close to the best answer there um and you know i think if you take a look at stuff like h100 we see like you know single digit percentage improvements right at the kernel level at the lower levels yeah at the lower level because it's been so hammered out and it's something that you can figure out like statically like oh i have to run this you know deep seek model here are all the optimizations i need to do in order to make this performant on an h100 um interestingly if you start looking at even like an rtx 6000 right which i think is like a b40 uh or a b200 you'll start seeing like much more significant gains to be had over there um you know we have seen 20 30 40 improvements in performance in many cases uh and that's because it's much much it's it's a lot less explored than the hopper optimizations have been um if we start taking a look at you know things like kernels on you know mac hardware kernels on you know amd and intel we also see like significantly more improvements like you know sometimes even like over 2x uh and that's partly because like some of these frameworks like apple for example they don't have a properly working like torch compile equipment so we get a lot of benefits because you're comparing against like relatively unoptimized code.
31:15So we've been talking about Intel AMD and NVIDIA, but you also support deployment on the M processors? Yeah, we have always, you know, we've, like I said, we started off as an edge company and we have a soft spot for edge hardware. So, you know, we always have dreams that will eventually have a fully distributed cloud where we can run stuff in the cloud and run stuff on end hardware and have it orchestrated between both of them over time. So we play around with those dreams all the time. Nice. And do you have anyone running on Mac hardware kind of in the data center? Because that's something that people are doing.
31:51I'm sure people are doing it because I think like, you know, probably Apple and other folks like that will be incentivized to do that. But we don't we don't work on that. Doing that today on data center side. When you say you don't work on it, mean you don't have anyone doing it or the product doesn't support it yet? A product could run on Mac hardware because we do support that, but we haven't done it at data center scale ourselves. Only on end user side things. Okay. Okay. But yeah, you're a kernel area, you know, somewhere between single digit percentages to maybe like, you know, a 2x improvement exists over there based on what hardware you're looking at.
32:23Could be even more if you're running on very unoptimized or new hardware. Right. The, you know, the compiler side, it's mostly around just like doing better fusion, figuring out what all this stuff is. probably at this point, if you compare against best in class, maybe you're looking at single digit percentages, maybe 20, 30 % over there. And some of that is just because we're trying to build, the main challenge of the compiler is not the optimizations, it's just being able to apply these optimizations across different vendors hardware. And the last layer of the stack, I think is where the vast majority of benefits exist.
33:03A lot of times you hear people say things like, like, oh, my GPUs are only like 30 % utilized, right? And what that ultimately means is that you're wasting like two thirds of your GPU capacity. So there is a fair amount of like optimization work around just more efficiently utilizing the hardware you have available. And that's where I think the vast majority of the benefits are. Do you come across interest in applying head over gen 80 for like training scale or is it primarily interested? Like, is there, like, if so, like, what are the structural, if so or if not, like, are there structural reasons why folks are or are not trying to do that?
33:46I think on the training side, when people talk about heterogeneity, it's usually still heterogeneity in a homogenous sense where they're like, okay, I'll have an NVIDIA cluster or I'll have an AMD cluster and they don't ever try to mix the resources. And part of the reason is, like if you take a look at training hardware it's kind of gone the way of like building supercomputers right like you know people don't talk about building machines anymore they're like here's my entire rack right this starts looking like you know what cray was doing like 40 years ago or something where they're like oh here's like an entire rack of machines so in some ways you know you could be like oh we kind of regressed back to the supercomputer era and i don't know if i use that word positively right we're building this like fully vertically integrated systems and i'm not sure that's the route for inference.
34:27I think inference is much better served as a large-scale work cloud where you can utilize a bunch of relatively commodity hardware and be able to scale out efficiently. And I think the only way this becomes scalable and sustainable is that you can interoperate easily with different hardware. So I do think that there's going to be some divergence over time. And I think you're going to start seeing this. Even NVIDIA, for example, announced their CPX, which is their context processor. It's a way that heterogeneity making its way into the work cloud, even within a single vendor, single generation.
34:57Yeah, but there are certainly vendors that are pushing the kind of supercomputer profile, you know, rack scale solutions for inference, thinking of folks like Cerebris, thinking of folks like Qualcomm, Saminova, you know, the same ideas apply, but it sounds like you're seeing a lot more heterogeneity on the inference side. Yeah, I just think that, you know, as systems tend to scale out and mature, they tend to become more disaggregated over time and not more aggregated. And maybe over time, all of the new expensive hardware goes first to training and then inference gets the hand-me-downs. Right, and I think in most cases, I'll be fine, right?
35:44Like, you know, you might need some of the highest-end hardware for the most performance-critical work, even on the inference side, but then you can use a lot of the older hardware to then do the inference. Can you talk a little bit about, just to make the conversation more concrete, like specific workloads or use cases or customer experiences? Yeah, so we've deployed on a bunch of different heterogeneous hardware that's deployed. I think there are some numbers that you've probably publicly seen from us on our blogs and maybe even from some like other companies. For example, we showed that if you utilize um intel's uh gaudi processors which are you know relatively old at this point um and mix that with like a b200 or mix it with an h uh h100 or h200 you can actually get pretty significant tco benefits um and the reason for that is you know you're basically if you if you think about the different work clouds you're you're essentially looking at what is the cost to compute and the cost of memory bandwidth and the gaudi has like extremely low cost of memory bandwidth because it's an hbm chip that's relatively low priced um so if you can start to think about how to utilize that difference you can actually get pretty significant uh cost cost benefits on that um and because it's so much you know relatively so much cheaper you can use a lot more of them and essentially make up for all the perf differences you have um and we can pack the most performance critical workload on like a b200 so from a user perspective they can't really tell but the cost is actually substantially lower.
37:22Have you talked to the folks at Qualcomm? They have their AI inference systems. They just announced the AI 250 and AI 200 systems. Seems like one of the things that they are talking about is inserting themselves into these heterogeneous environments. and I would imagine that something like what you're offering might facilitate that. Yeah, I think it would be useful to them. We're just in like very early conversations with them, to be honest. But I think it makes a lot of sense. Like one of the things that, you know, everyone thinks about is that most of these data centers and most of the customers like, you know, have a very high preference to buy large scale NVIDIA hardware, right?
38:13Because that's the default. Everything's going to work on it. So the reason heterogeneity is interesting from a business perspective to other semi-companies is that it's an easy way for them to go to their customers and be like, hey, if you buy a few of these on the side, we can show you the benefit of heterogeneity. And then you don't have to wholesale make large bets on... You don't have to rip and replace. Yeah. Exactly. Yeah. And in particular, it makes the story easier if you're competing against potentially depreciated older NVIDIA capacity, right? It makes it easy for you to say, oh, I can utilize, you know, all my old NVIDIA capacity plus some of the newer capacity from a different vendor and then see what the changes are over there.
38:55Your typical customer deployment is like clearly heterogeneous in terms of infrastructure. are they you know similarly more heterogeneous in terms of use cases and applications or are they you know uh more homogenous like they they are you know running this to support a single application or a single you know agent or set of agents at a large scale it's mostly mostly the latter uh the applications we typically see that are running in a lot of the data center stuff that we target they're relatively homogenous because they're usually utilized by like one or two or three customers um most of the deployments that we've seen have like basically heavy hitter customers and then you know to serve like a large scale they're going to have a whole bunch of really small workloads but for the most part the vast majority of the workloads is are coming from like a couple of different players okay so it's not necessarily something that for example like uh a CSP cloud service provider might create a general purpose inference cluster using your stuff.
40:03It's more, you know, I have this workload and I need a place to run it and try to drive down cost per token. I mean, they could, a CSP could use our stuff. But these days we've mostly been targeting people who have either their own data centers or like, you know, think about like sovereign clouds and other Neo clouds. But have you seen a lot of, talk a little bit about what you've seen in the sovereign cloud space. Like I hear that, you know, thrown around a lot, not a lot of use case examples, you know, in part because of the, you know, the sovereign nature of it, perhaps. You know, what are you seeing there?
40:41Yeah, so I think even amongst the sovereign clouds and neoclouds, there's kind of, there's a few different classes over there. So there's some people who have like pretty good uses of their clouds. And there's some people that are building cloud capacity, but it's not entirely clear what applications are running on there just yet. And a lot of it, I think, is a software gap. We try to target people who are actually using their stuff because it gives us better feedback on how to improve our things. There's also a class of like, say, the sovereign clouds who are building data centers up, you know, even in the US.
41:19They're not like US sovereign clouds, but they're external sovereign cloud companies that are building capacity up in the US. I think they have like a lot of real workloads because they're basically selling that capacity to other customers. So those are interesting people for us to work with because they have like actual usages. And then one of the other things that kind of drives sovereign clouds is because of all the export restrictions and everything, they're pretty limited by what hardware they can get, which typically makes their data centers like heterogeneous by almost by design because they're just getting like whatever they can find.
41:47So that actually is a pretty interesting area for us just because we can enable the usage of all that hardware. You know, along the lines of like talking about Mac or the M series chips, talking about Qualcomm, do you anticipate that your focus will primarily remain in VD, Intel, AMD or are you anticipating like an extreme level of heterogeneity where you're supporting, you know there's kind of a long tail to it yeah so what i want to say is on you know we have support for all these different vendors but the vast majority of deployments that you know we go into or look into are still you know mostly nvidia nvidia shops um and even within nvidia the heterogeneity is like pretty interesting right across different accelerators there's different cost economics and stuff so we can still utilize that and orchestrate things um we there are a few people we're talking to that want to have data centers with extreme heterogeneity and these are mostly on the sovereign cloud side and that's because of the issue that i said they have so much capacity they can get from each one so like okay we're going to put all of them in here to get the maximum capacity um but for the most part what we find out is that people have like two vendors in their data center um and that's partly because you know we're solving one layer of the stack but you know as you start thinking about all this heterogeneity just doing the system management and everything across a lot of different vendors is a lot of work.
43:16So people tend not to have more than like, you know, two, maybe three at the most vendors and their actual data centers. And is the heterogeneity that you're seeing at like the, at the rack scale or the, you know, rack of racks type of scale, or is it, you know, even more distributed than this? Like many years ago, a startup that I worked at was really kind of on the cutting edge of like distributed systems. And like we would go in and like stand up these environments with like literally like, you know, a cast off of this, like some, you know, Dell DL 240s, like a little bit of that, a little bit of that.
44:01like very, very like, you know, part of the idea was that you could, you know, create enterprise grade, you know, SLAs on commodity compute. And, you know, that idea has surfaced in a lot of different ways in the industry over time. are you seeing like that level of, or are you anticipating that level of heterogeneity or is it more like I'll have my rack of H100s next to my rack of slightly newer, older gear? Yeah, normally it's a little bit further away than on a single rack because usually single rack people have the same type of machine on. Especially with the vertical integration that happens on all these machines.
44:49machines um i i guess maybe um if we take a step back there's kind of there's pretty much like laws of physics limits on how many machines and capacity you can have per rack right if you go to like a generic data center you know you can maybe have like 20 kilowatts a rack and you can't really stack like a ton of machines so you'll actually see a lot of data centers that run gpus that aren't optimized for gpu workloads typically have like you know just a couple of machines in each rack because you're going to exceed the power budget of that rack. If you go past like, you know, 25, 30 kilowatts, you need to start having, you know, back of rack cooling or like a rear door heat exchanger is what they call it.
45:28And that usually requires some level of outfitting, right? So they have, it's liquid cooling, but liquid cooling going to a radiator in the back of the rack. And then if you really start, you know, going to, you know, 100 kilowatt plus racks, you start going into like direct pitch of cooling, right? which is what a lot of the latest greatest data centers would use. And once you get to that point, you're basically building like a very rack level system. So typically, you know, if you're like not super sophisticated, you have like one or two machines a rack because that's all you can do. And if you're very sophisticated, you have all these things packed in super tightly, but then now it's very vertically integrated.
46:04So either way, you're not seeing a lot of heterogeneity on a single rack because there's just not that much space one way or the other. Yeah, essentially the GPU and the cooling and power requirements kind of take you out of that kind of super heterogeneous, like stack up a bunch of commodity stuff. Like it tends to be more, got it. Okay. And that might change, right? As like lower power systems and, you know, more optimized systems come off. But the industry trend so far has been to try to increase. And maybe smaller models as well. And maybe smaller models. yeah but the industry trend at least on the hardware side uh if you look at nvidia's road map is that they're actually putting more and more things on the rack right like i think their next generation rack which is kyber is going to be like 600 kilowatts right it's like it's like almost an entire data center and like and like a single rack yeah any thoughts on like from where you are today you again you've been at this for a couple years recently launched though a lot of early success like how you see things evolving over the next year Yeah, so we are super excited about, sorry, we're super excited about getting our, you know, agent cloud launch, which will be more of a developer facing products.
47:20Like so far, we've been going mostly to like data centers and et cetera. But, you know, I think our heart is really into like going into a developer facing product, working with developers to make running like these models and stuff like much cheaper and easier and faster. Actually, as you started saying that, I was going to say, like, if I'm a developer and you're not, you know, offering me a framework, do I care much beyond like the, you know, the cost per token and the feeds and speeds of the performance ultimately? Like, what are the qualitative reasons why I'm going to care about running on your cloud as opposed to someone else's?
47:56Yeah, so I think the areas that, you know, will have a lot of capabilities around are one is like orchestrating all the large scale agentic workloads, including like asynchronous work and stuff. That's been a big area where people have asked us for like, oh, I want to make a whole bunch of batch model calls and stuff by a bunch of like asynchronously running agents. Like we can support that pretty easily in Gimlet. So that's actually a big, big part of it. The other thing is obviously, you know, the cost economics and performance and just like the observability around how these workloads are running.
48:25Well, Zane, thanks so much for jumping on. And I feel like this was a bit of a rapid fire chat about what you're up to. But I learned a lot and it's interesting stuff and looking forward to keep tabs on it. Thank you. Thanks again for taking a dive. Yeah, thanks so much.
48:54you
From the publisher
In this episode, Zain Asgar, co-founder and CEO of Gimlet Labs, joins us to discuss the heterogeneous AI inference across diverse hardware. Zain argues that the current industry standard of running all AI workloads on high-end GPUs is unsustainable for agents, which consume significantly more tokens than traditional LLM applications. We explore Gimlet’s approach to heterogeneous inference, which involves disaggregating workloads across a mix of hardware—from H100s to older GPUs and CPUs—to optimize unit economics without sacrificing performance. We dive into their "three-layer cake" architecture: workload disaggregation, a compilation layer that maps models to specific hardware targets, and a novel system that uses LLMs to autonomously rewrite and optimize compute kernels. Finally, we discuss the complexities of networking in heterogeneous environments, the trade-offs between numerical precision and application accuracy, and the future of hardware-aware scheduling.
The complete show notes for this episode can be found at https://twimlai.com/go/757.




