Everything you need to run Mission Critical Inference (ft. DeepSeek v3 + SGLang)

19 Jan 2025 · 1 h

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Everything You Need to Run Mission Critical Inference (ft. DeepSeek v3 + SGLang)

Podcast Details

  • Title: Latent Space: The AI Engineer Podcast
  • Episode: Everything You Need to Run Mission Critical Inference (ft. DeepSeek v3 + SGLang)
  • Date: January 2025
  • Guests: Amir Haghighat and Yineng Zhang from Baseten

Key Highlights

  • Model Launch: DeepSeek v3, a highly anticipated model, has been released and is currently one of the best open-source LLMs, positioned 7th in the LM Arena leaderboard.
  • Trends in AI: A surge in the release of large open weights models from Chinese labs, highlighting the competitive landscape in AI model development.
  • Inference Challenges: Deep learning models of this scale present unique serving challenges that require specialized infrastructure.

DeepSeek v3

  • Specifications:
  • Contains 671 billion parameters.
  • Requires significant resources (e.g., using H200 GPUs for optimal performance due to high memory requirements).
  • Performance: DeepSeek v3 is optimized for high throughput and low latency inference under mission-critical conditions.

Three Pillars of Mission Critical Inference

  1. Performance at the Model Level:
  2. Focuses on the efficiency of a single model running on a GPU.
  3. Utilizes techniques like speculative decoding for improved throughput.
  1. Horizontal Scaling:
  2. Addresses the need to scale models across multiple replicas and regions.
  3. Emphasizes that traditional tools like Kubernetes alone are insufficient for handling mission-critical workloads.
  1. Developer Experience for Compound AI Systems:
  2. Importance of a robust developer experience for managing complex workflows that involve multiple models.
  3. Emphasizes the need for tools that facilitate low-latency multi-step inference.

SGLang

  • Overview: A new framework designed for efficient model serving with superior performance and usability compared to existing frameworks.
  • Key Features:
  • Radix Attention: Optimizes KVCache for better performance.
  • Concentrated Decoding: Transforms decoding processes for structured outputs.
  • Speculative Execution: Allows for more responsive control flows in models.

Insights from Guests

  • Amir Haghighat: Co-Founder of Baseten, highlighted the importance of performance optimization and infrastructure for serving large models.
  • Yineng Zhang: Lead Software Engineer, discussed the technical challenges of deploying and optimizing DeepSeek v3 and the advantages of using SGLang for high-performance inference.

Challenges & Solutions in AI Inference

  • Scalability Issues: Traditional methods often fall short in meeting the demands of high-traffic applications; innovative multi-region strategies are necessary.
  • Quantization and Performance: Challenges in serving large models at scale, particularly with FP8 quantization, which many current frameworks do not fully support.

Future Trends

  • A focus on model fine-tuning and RLHF (Reinforcement Learning from Human Feedback) as methods to enhance model performance in specialized domains, e.g., healthcare.
  • Observations on how developers can better leverage frameworks like SGLang for custom requirements and complex workflows.

Conclusion This episode provided deep insights into the evolving landscape of AI model development and deployment, emphasizing the critical need for robust performance, scalability, and usability in mission-critical AI inference systems. The discussion around DeepSeek v3 and SGLang highlighted both the potential and the current challenges faced by AI engineers today.

For more detailed notes and resource links, visit the [Latent Space website](https://latent.space).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome back. tokens of data, including synthetic reasoning data distilled from DeepSeq R1. Right now, on the LM Arena leaderboard, DeepSeq V3 is rated the 7th best model in the world with a score of 1319, right under the full O1 model, Gemini 2, and 4O latest, and above O1 Mini, Grok 2, Gemini 1.5 Pro, and Claude 3.5 Sonnet. This makes it the best open weights model in the world in January 2025. There has been a big recent trend in Chinese labs, releasing very large open weights models, with Tencent releasing HunYuan large in November and Hyluo releasing Minimax text this January, both over 400B in size.

1:14However, these extra-large language models are very difficult to serve. Base10 was the first of the inference NeoCloud startups to get DeepSeq v3 online because of their H200 clusters, their close collaboration with the DeepSeq team and early support of SGLang, a new VLLM alternative out of UC Berkeley that is also used at frontier labs like XAI. Each H200 has 141 GB of VRAM with 4.8 terabytes per second of bandwidth, meaning that you can use eight H200s in a node to inference DeepSeq V3 in FP8, taking into account KVCache needs. We have been close to base 10 since Sarah Guo introduced Amir Hagihat to SWIX, and they supported the very first Latent Space Demo Day in San Francisco, which was effectively the trial run for the podcast you're listening to right now.

2:12Since then, Philip Kiley has also led a well-attended workshop on TensorRT LLM at the 2024 World's Fair. We worked with him to get two of their best representatives, Amir and LED model performance engineer Yining Zhang, to discuss DeepSeek, SG Lang, and everything they have learned running mission-critical inference workloads at scale for some of the largest AI products in the world. Spoiler! Amir thinks there are three pillars of mission-critical inference workloads, and we spend quite some time discussing what you need for each of them. In other news, invites are now rolling out for the second AI Engineer Summit in New York City from February 20th to 22nd.

2:56We are bringing back the surprisingly successful AI leadership track from World's Fair, and the AI engineering track is now wholly focused on agents at work. If you are building agents in 2025, this is the single best conference of the year. We are curating all attendees and will sell out after we announce speakers this coming week from DeepMind, Anthropic, OpenAI, Meta, Jane Street, Bloomberg, BlackRock, LinkedIn, and more. Look for more sponsor and attendee information at apply.ai.engineer and see you there. Watch out and take care.

3:39Hey, everyone. Welcome back to the Lead in Space podcast. our first recording of 2025. I'm Alessio, partner and CTO of Decibel Partners, and I'm joined by my co-host, Spix, founder of SmallAI. Hey, and today we are here with a special double guest episode with Amir. Oh my God, I don't know your last hug yet. That's close enough. That is good. I thought it was close to try to go. That's really good. And Yining Zhang from Base10. Welcome. Thank you. Amir, we've met before. You're a co-founder of Base10, which is one of the leading sort of LLM inference platforms. I don't know how you, what do you consider yourself?

4:17That sounds fun. And you are a lead software engineer on the model performance team. And you guys recently shipped DeepSeek V3 as one of the many models that you do host. You also are very involved in SGLang. And that was actually one of the reasons we're discussing an episode with you, even before DeepSeek V3 dropped as a Christmas present to everybody. So we can take this a number of directions, But I think one thing we wanted to just get off the bat on was to start with DeepSeq and maybe and then we'll work our way backwards back to SGLang. But DeepSeq is, you know, more kind of more recent.

4:50Why are people so interested? What's the history of like, I guess, like DeepSeq in general from your perspective? Yeah, because DeepSeq V3, I think, is currently considered the leading open source LLMs based on the benchmark results and the chat area results. and it's so big you know it's 671 billion parameter MOE and I think it's a game changer for the open source AI so everyone is interested in this model yeah one of the interesting things is like they are bootstrapped like you know very private small lab like they have a lot less fewer resources than than others but also it's also interesting like that it's just open weights For some reason, the Chinese labs are much better than the American labs at sharing open weights.

5:43And that's obviously beneficial for Base 10. It is in your incentive to serve these models at all times. What are the unique challenges that you face offering something this large? Yeah, I think because the model is very large. And if we use something like H100, we cannot serve this model. because, you know, even we use the H100 8 cards, it should be 640 gigabytes memory. So DeepSig V3 model, it has 671 billion weights. So even use the FP8 Precision, you need, I think, 71 gigabytes for the weights. And you also need extra memory for the KVCache. So it's not possible to run that on H100. That's why we chose H200 to run that model or use Smarty node to run that model.

6:38Yeah, it's very challenging. And I think another one is that DeepSync 3, the way it's released was the FP8 pre-session. So if you want to run it, you should support that kernel because I think the default is the BF16 and even larger. But if you want to run the FP8, you need to support the quantization. So I think currently even the Tensor RTLM doesn't support the FP8. So if we want to implement that feature, we should do some feature development for that. The last challenging part is that if you want to do some debugging or some performance benchmarking, it's very hard. Why? Because the model is so large and the loading time is so long.

7:30Yeah. So it makes it more complicated for the developer to do some debug. Is it complicated or just slow? I mean, you've just only mentioned loading time, but like... Yeah, loading time is slow. Yeah, you're right. It's not complicated. It's just, you know, just go for more coffee. Okay, okay. Okay. Can you maybe just give people a quick rundown of all the models you support on Base 10, how it compares just on size, just to, you know, people hear 671 gigabytes, but like, is that a lot more than other models? And then you mentioned the BF16 FBA, what's kind of like the usual that you see? And do you see any variation based on model size or anything like that?

8:10I think at the base 10, something like LAMA 70B is more common. I think LAMA 3 has released the 405 billion weights, but I think there are just a few users use that. So before the DeepSync Fee 3, I think we haven't encountered that issue for that so large weights. I think DeepSync Fee 3 is the first one so big model that we should use H200 or use H100 multi-nodes. And was that because of performance or why did people not use the 405B llama? I think the, you know, what I hear from people is like, you know, the performance gains of like the 405B at inference times are like not worth it, you know?

8:56So the 70B is kind of the sweet spot. Who are the people that use V3? Are people that were using maybe the Lama 70B model and just want better performance? Are people that are just experimenting? I think that's kind of the question that people always have. There's always a lot of excitement around open source models, but then maybe the question is like, what are they really good for? I can answer this just observationally. The interest that we have seen, and some of this is running in production, some of it is just at the interest level, is generally not coming from folks who are trying to upgrade from a certain open source model to DeepSync v3.

9:33We're seeing it generally speaking from folks who are coming from Cloud and are doing so either because, and here's, I'm going to give you a list of reasons. and generally the reasons are a combination, a certain combination of these. In no particular order, either it is they're being rate limited or the price is too high, or they have certain latency requirements or time-to-first token requirements for their use case that cloud cannot hit, or they want to have full control over the model as opposed to running it behind an API where the model underneath them can potentially change. And a couple of other reasons, but generally it's a combination of those.

10:17You mentioned the speed and some of these things. Do customers want to change the hardware? Also, you're referring to H2O as the default thing. Do people come to you and it's like, hey, I'd rather use a smaller system and can I have worse performance? Or how do you work with customers on that? Generally, people come with certain requirements around the latency, the throughput, the cost. Generally, they're not coming in saying, I want this particular GPU skew. At least like as we go up market and we're talking to, you know, foundation model companies, for instance, the things that are top of mind for them are those requirements, not a particular GPU skew.

10:56We're doing the different GPU skews not because we want to, you know, offer, oh, look, we have H200s, look at us. We have MIGD DH100s, look at us. It's not that. It's really because those are the tools to achieve a certain kind of time-to-first token for certain types of models or certain kind of throughput and scale or a certain kind of price per million token or what have you, per million images, depending on the modality. And that's the reason that we're talking GPU skews. I wanted to pick up a little bit on this FP8 thing. It seems like, you know, I think Noam Shazir started talking about sort of training natively quantized.

11:41And I think that's what DeepSeek seems to have done, at least they said in their paper. Is this a trend? Like, is the community settling on one form, one sort of numeric that everyone knows about? Tell us more about what you're seeing here in terms of like the training trends, those sort of model trends in the, I guess. So I think a lot of companies as well, like together, they'll also release like quantized versions of the Lama models, right? Like for like turbo or light inference or, you know, like just based on different levels of speed. Like, do you do anything there in terms of quantizing the models that you serve?

12:13I'll let Yining answer the sort of patterns around, you know, using FP8 in training. But I want to draw one distinction that I think gets to the latter part of your question, Sean, which is that unlike companies like, you know, you mentioned Together or, you know, Fireworks or Replicate, Base10 doesn't provide a shared inference endpoint for the popular open source models. That is a product that we don't have on purpose. Those work really well. The shared inference endpoints for open source models work really well for situations where the user is saying, hey, let me just call a certain popular model behind an API, pay by the token.

12:53That is not our average customer or the median customer. Our customers generally have their own custom models, very custom workflows, strict requirements around latency and time to first token, and can't deal with noisy neighbor problems. Like, oh, the API is slow because some other customer has been calling it a lot. Other requirements around infrastructure flexibility, around regions, for latency reasons, for compliance reasons. That's the side of the inference market that we capture. And at base 10, when you deploy a model, whether it's your own custom weights or an open source model, you get dedicated inference, dedicated resources, dedicated inference.

13:36And so when it comes to the quantization question, where it matters is that we would never quantize the model behind the user's back and say, look at us, there's a faster Lama 70B that has been somehow quantized to be faster and cheaper. our customers are coming to us with those requirements that I mentioned, but in particular, when it comes to model quality, they have strict requirements. They would not be okay with us touching the weights, if you will. We have done things like speculative decoding with a couple of different ways, but all of those things, those methods guarantee the output is unchanged as opposed to quantization.

14:13So when it comes to quantization, we have built tooling that allows are users to quantize their models for the ones that we're working more hands-on through our forward deploy engineering team. We're working with them on evals as well to ensure that the quantized models are meeting their requirements. However, this is all very much in conjunction with the engineers that are customers as opposed to us doing it behind the scenes. Yeah, FP8 training is very interesting. And I think the DeepSeq team is the first one to use FP8 training for the large model. I think before that maybe LinE.ai sorry 01.ai they used the FP8 training and others, I think most of them used the BF16 training and it's the game changer.

15:06And for us because the FP8 kernel should be implemented inference it used the blockwise FP8. And currently, even you use something like Kuda or Kublas, you cannot support that. So you should use something like Tweeten to implement the kernel, or you should use something like Catalyst to implement that kernel. So I think that's the challenging part. My theory is that this will pick up in terms of the models that people release. Like increasingly, it'll not be BF16. There was a bit of quantization. I was trying to look for the paper while you were speaking but i couldn't find it but there's a there's sort of this like ablation of quantizations paper that was out last year that showed that there is benefits to quantizing and sort of natively trading all the way until like like six bit and then like even smaller than that and maybe going too far yeah i'm not sure if you you know what paper i'm talking about but yeah there's there's an interesting trend for sure yeah i i think even they use the fp8 quantization, the benchmark result is very good, such as something like GSM8K.

16:13The score is nearly 94.6. It's so high, you know. I think it's higher than every other open source LAM, even the LAMA 400, 05 billion. So I'm going to move on a little bit in terms of like one other notable detail, and then, you know, we don't have to speak too much about DeepSea because obviously we don't know that much unless you work on the team. But another trend that they have is the fine-grained MOE. I think that there is this question about whether or not MOEs will be more of a thing. Basically, this time last year, Mick Strau was sort of kicking off a bit of an MOE trend with 7B, 8B by 22B.

16:55And then the rest of the year, no MOEs, basically. So is this discovery of fine-grained MOE is going to be a relevant trend for this year? Yeah, I think so. Because as far as I know, some companies such as Baidu or Baidu Dance, their internal dominant AOM, they use the MOE architecture. And their weights, I think, are similar to the DeepSeek MOE model. So I think after this new year, the MOE inference optimization will be very essential and so important. At the same time, why hasn't the big labs done it? I think, well, Lama 4 or 5B is dense. I think Grok is also a dense model. You can correct me if I'm wrong.

17:45And yeah, it's just a weird counter trend, I think. this time last year I was like kind of writing my recap and I was like all right MOEs like seem like they're going to be trending and then they did not trend anyway so it's just just a note that I would flag out there but I generally agree like it seems like FineGrain MOE is working out and I would definitely see definitely want to see more people adopting it I went to Jeff Dean's session in Europe's and he also mentioned as well that I think one of the Gemini I think Gemini 1.5 Pro as an MOE, which I don't think we knew before that. Yeah, yeah.

18:17So the reason why Lama open-source MOE model, because I think they tried to train our MOE model, but they failed. So that's why they didn't open-source MOE model for Lama series. Why are the causes of failure? Like, why do MOEs fail? Like, I think this is another thing that people are talking about, right? Like the failures of 3.5 Opus, the failures of GPC5, like, you know, it's a thing that people are sort of rumoring. Yeah, because if you want to train model, for the training staff, you need some benchmark or some score. But for the MOE model, the benchmark score is even lower than the dance model.

18:55So in that way, or in that case, they think the MOE model is worse than the dance model. So they didn't release that MOE model. Okay, one more thing, I guess, maybe more commercially relevant. D6 API pricing is very competitive. How do you decide pricing in this kind of landscape with open models? Yeah, so it goes back to the use cases that we serve. And again, going back to the fact that we don't have shared inference endpoints for the different models. And so our pricing is never per token. Customers, like I said, generally come with their own custom models or open source models, but with strict requirements around certain things.

19:37latency and time-to-first token requirements or security and compliance requirements or a particular scale that they're looking for without running into noisy neighbor problems, things like that. And the way that we price has always been based on consumption, based on consumption of resources. And that takes one of two shapes. One is the shape where things are running inside of our infrastructure. And we're running, by the way, on top of multiple different public clouds, many different regions within those. Then we charge them based on the hardware that the resources that they're using. The second shape that it takes is that our customer brings their own cloud.

20:15And we're seeing this more and more where a customer has big committed resources inside of their AWS VPC or GCP or what have you. And in that world, we also have a consumption model. Of course, the price is very different because they're using their own resources, but we are managing those resources for them. An example of that that we're seeing more recently is the fact that we've had to build multi-cloud capabilities for us so that we can have a single model horizontally replicate across different regions and even different clouds. And more and more as we go up market, our customers have their own cloud commits.

20:58They are also multi-cloud in order to be able to get good prices and good capacity. And it's unreasonable to expect every one of them to build the same multi-cloud capabilities that we have built. And so then they take advantage of what we have built and use all of the different cloud resources that they have as a holistic unit and have their models at inference time horizontally scaled across those and even optionally overflow to our cloud when they start running out of committed resources. All of that has a consumption pricing model to it. Can we talk about what it takes to actually run your service?

21:34So we had episodes with, you know, replicate model of our works. We always like to ask this question, obviously, since you're not the model maker, all the secret sauce is in how you actually run the model. I know you also have Truss, which is your more developer-led SDK. Can you maybe quickly run people through how do you go from taking the DeepSeq V3 weights to like actually run it? What goes on behind the scenes? and then we can talk about SG-Lang in depth a little more. Yeah, totally. So we have, like you said, we have Trust, which is our open source model packaging and deployment library. Trust works with different frameworks underneath it, has very native and deep support for TensorRTLM.

22:17Somewhat as an accident of history, we happened to have access to TRTLM before it was announced, contributed back to it, and we still do, pushed it to its limits and had to go beyond it in certain areas as well. So for example, you know, the Triton inference server, we've had to build our own version of that for performance and reliability reasons. But we invested in it heavily because it tended to be for the use cases that we were seeing from our customers, it tended to be the best framework to handle the latency and throughput requirements that we were seeing. In particular, when it comes to the kernels that they come with.

22:55I'm yet to see folks do better than what NVIDIA can do when it comes to Kruda kernels. However, trust is not tied to TensorRTLM. For example, for the DeepSync example that you mentioned, it's working with SGLang, which is really cool to see. And we will be investing more and more on SGLang, especially as the developer experience is just so much better than TensorRTLM. We've built a lot around TensorRTLM productize them too, to make it easier to work with, but still SGLang has been a joy to work with. Another trend that is really promising, and I learned this from the SGLang folks, is that the TensorR TLM folks have promised to modularize a lot of TRTLM so that other frameworks like SGLang can grab certain parts of it and build on top of it.

23:45And so as a user, you don't have to go all in on one framework versus another. You can really pick and choose based on the requirements that you have. And that's really been our approach as well. We have customers on Base 10 that are using TensorFlow and we have ones that are using VL and we have a growing number that are using SGLang2. It's not about really tying yourself to one versus another. It's about using the best of the bunch, depending on the requirements of the customer and for their inference workloads. How did you think about designing the framework? So Replicate, also a cog, which was kind of more tied to Docker.

24:25What were maybe some of the design decisions that you had? And how do you think that's changing, especially as the models change and like the run times change? Yeah, totally. So we started Trust, gosh, like four or five years ago. And at the time, the sort of principle that we had in mind was let's make sure that easy things are easy, but hard things are possible. And so an example of easy things being easy is that, you know, think of it as a very simple, you know, you have a model. What do you need to do to serve it? Well, you need to load it up and then you need to write the code for the inference path.

25:03And, you know, Trust actually had, you know, hooks for these two things. And so you could just write two functions and voila, your model was being served at least as a single unit. We can talk about the horizontal scaling part separately. That's a whole different topic. And so we did well when it comes to easy things being easy. I think we struggled with the hard things being possible in the early days. Hard things, example of hard things are cases where we're seeing where more and more our customers have their own custom models. Custom models that sometimes they fine tune, sometimes they've pre-trained.

25:35You know, we now have six or seven foundation model companies as customers who are sophisticated enough to pre-train their own models. And they're trusting us with the inference layer. That's not a situation of, hey, here's two functions, good luck. I have to have much deeper integrations with them. And so that's where we started rethinking some of the abstractions of trust over time to allow for those custom use cases. And that has been successful. Another place where we didn't think about at first, but became important, was seeing more and more use cases where the customer was saying, I can serve my models on base 10 using trust, fine.

26:13but my use case is not just call the model, get the response and run with it. I actually have a multi-step inference workload. So an example of that is the company Bland AI with their AI phone calls. To make an AI phone call happen, you need to transcribe what the human said, a couple of LLM calls to figure out what to say back, and then text-to-speech to actually have the end-to-end workflow working. Now, you can have these be three separate models, three separate deployments, But think about what happens is that you have to call the first model, wait for the response, call the second model, wait for the response.

26:48All of that network back and forth is killing you. The latency is becoming too high. That's not something that we had designed for initially. And so that's when we came out with Trust Chains, which is the devx for building these multi-step, multi-model inference workloads, but doing so in a very low latency way. So that instead of you orchestrating all of these calls and incurring all of that network latency, you're actually making one call and these models are actually talking to each other. They run independently on their own hardware, with their own autoscaling behavior, but the data from one to the next step is being actually streamed.

27:27And that way, going back to the AI phone call use case, you can get sub 400 millisecond latency AI phone calls that actually feel very realistic. And those are all models hosted on base 10, or do you also do a change? Those have to be models hosted on base 10. If one of those steps is not hosted on base 10, then you're still incurring a massive latency on the network side. Yeah. And then just to maybe tie this into SGLang, how do you kind of think about the hidden magic? Should people know that you use SGLang? Should people care, especially for the people building the models? Does it matter to them that you use a certain model runtime or do they not care?

28:08Everything just goes through the baseline platform the same. Yeah. Should we talk about it? Yes, 100%. We want to be the transparent provider. I don't want to say, oh, just give us your model and voila, magic and trust our magic. I want that magic to be very transparent to our customers. That has worked really well for us. And you really need that, especially when you're onboarding foundation model companies. And they're not going to just turn a blind eye on how things are run underneath the hood. When it comes to customers caring about what's happening underneath, they do. But more than caring about this framework versus that, they care about how the final output.

Read the full transcript

28:50In other words, is the quality the same or somehow something has changed underneath the hood and the model isn't actually producing the same quality? How is the latency? And especially for certain use cases, what is the time to first token? And is that sustained? What is the P95 of that? What is the P99 of that? How well does it handle throughput? When you start getting a massive burst of traffic, does it still sustain those P95s and P99s? How do I make sure that the security of the data being sent into the model is guaranteed? How do I make sure compliance is guaranteed for HIPAA use cases? And how do I make sure that the data remains within a certain geo for compliance reasons or for latency reasons?

29:38And so those are the concerns that the customers are coming to us with. Less so about, hey, here's my model. Make sure you run it with your TLM or make sure you run it with SGLang. Can you maybe give us an overview of all the different frameworks that people might use? So you have SGLAN, TRT-LLM, VLLM. Those are kind of like maybe the open source research ones. And then some of the other commercial companies are building some of their own stuff. But what's the state of the art today? Maybe like the top three most popular. And then we can talk about why SGLAN came to be and what makes it different and some of the performance boosts that you get.

30:15Okay. Yeah. I think for the common use case, maybe not the DeepSea Cafe 3, for the common use case, I think SGLAN's performance is better than FLM, and its usability is better than TENSR-TLM. So when users care about the performance and the usability, I think they will choose SGLAN. And for the DeepSeq Phi 3 case, because we do a lot of optimization in SGLAN, something like DeepSeq Phi 2, they proposed attention parent named MLA, multi-latent attention. And I think SGLAN is the only framework to support that. Maybe Lite-ALM and TRT-ALM also support, but PLM doesn't support. And also in SGLAN version 0.4, we also support the DP attention for DeepSeq.

31:08And in the latest SGLAN release, we also support the blockwise FP8 kernel. And that kernel was adopted and copied by PLM later. So I think we have done a lot of optimization for DeepSeq. That's why SD-Line is the recommended engine by the DeepSeq team. And maybe one thing to point out, and I think this is important, is that the framework that you choose is part of the equation for running mission-critical inference workloads, but it's only a part of it. So maybe I can draw this out just based on the experience, based on what I've seen in the market as to what it takes to run mission critical inference workloads in production.

31:53I think it takes three things. And each of them individually is necessary but not sufficient. One is performance at the model level. So in this case, how fast are you running this one model running on a single GPU, let's say. The framework that you use there can matter. The techniques that you use there can matter. The MLA technique, for example, that Ine mentioned, or the CUDA kernels that are being used. But there's also techniques being used at a higher level, things like speculative decoding with draft models or with Medusa heads. And these are implemented in the different frameworks, or you can even implement it yourself, but they're not necessarily tied to a single framework.

32:37But using speculative decoding gives you massive upside when it comes to being able to handle high throughput. But that's not enough. Invariably, that one model running on a single GPU, let's say, is going to get too much traffic that it cannot handle. And at that point, you need to horizontally scale it. That's not an ML problem. That's not a PyTorch problem. That's an infrastructure problem. How quickly do you go from a single replica of that model to five to 10 to 100? And so that's the second pillar that is necessary for running these machine critical inference workloads. And what does it take to do that?

33:16It takes, some people are like, oh, you just need Kubernetes. And Kubernetes has an autoscaler and that just works. That doesn't work for these kinds of machine critical inference workloads. And you end up catching yourself wanting to bit by bit rebuild those infrastructure pieces from scratch. This has been our experience. And then going even a layer beyond that, Kubernetes runs in a single cluster. It's a single cluster. It's a single region tied to a single region. And when it comes to inference workloads and needing GPUs, more and more we're seeing this, that you cannot meet the demand inside of a single region, a single cloud's single region.

34:00In other words, a single model might want to horizontally scale up to 200 replicas, each of which is, let's say, two H100s or four H100s or even a full node. You run into limits of the capacity inside of that one region. And what we had to build to get around that was the ability to have a single model have replicas across different regions. So there are models on Base 10 today that have 50 replicas in GCP East and 80 replicas in AWS West and Oracle in London, etc. And that was a big investment that we had to make. The final one is wrapping the power of the first two pillars in a very good developer experience.

34:44To be able to afford certain workflows like the ones that I mentioned around, you know, multi-step, you know, multi-model inference workloads. Because more and more we're seeing that the market is moving towards those, that the needs are generally in these sort of more complex workflows. So these are the three pillars that it takes to run mission-critical inference workloads. And the choice of the framework, the serving framework, is really a part of the first pillar. And that's something that I'm seeing in the market that people who are somewhat new to it, they're like, well, VLM equals equals production.

35:21That's what it takes to run inference workloads. And in practice, that is not true. And I wanted to call that out. I agree with Amir because I think it's open source libraries such as VLM, SGLang, LATLM, or TENSRTM. They only provide a library. They don't provide a product solution. Yeah. Can we maybe talk about some of the SGLang unique things? I read through the paper. It sounds like some of the main use cases, like when you have very large batches, which makes sense for your use case. And also kind of longer context. You know, what was the decision buying creating the framework? which I think is like around one year old.

35:58I think the paper came out December 2023, something like that. So it's still fairly new compared to some of the other ones. And then maybe what were some things that you had to change, you know, as you built it or any fun stories? Yeah, yeah, yeah. I think last year, oh, not last year, sorry. At 2023, maybe August, at that time, Viamin and Ying want to create the SGLN maybe for the front end, something like LLM program. They want to solve that problem. And at 2024, January, they support something like Redix cache. It's a prefix caching technology. I think SG-Line is the first framework that supports prefix cache.

36:39And at February, they also support a concentrated decoding and support some jump forward. So at that time, it's a no for the language generator, not the inference backend. And at 2024, July or June or July, we want to make SGLAM a fully functionality ALM inference engine is just equivalent with the VLM or with the TensorFlow ALM. So at that time, we published a blog compared with other frameworks. And its performance is amazing. At that time, I think its performance is maybe three times, its throughput is three times than VLM. So after that, VLM also do some refactor to make it faster. And at September and December, we continue to release new versions for SG-Line.

37:36Yeah, we support some DeepSeq optimization such as MLA optimization, DPR tension optimization. And we also support the serial overhead, CPU schedule. Also, we support something like SGLN router for the cache aware load balance. Yeah, we deliver so many features. We just build and ship. And I think why Liamin and Ying want to create a new framework rather than use the existing solutions such as VLM or TensorFlow RT. because at that time, you know, at that time for the VRM, I think it's easy to use, but its performance, maybe it's not good. Some design, I think it's not okay. Yeah, maybe the code is a little messy.

38:21And if you want to extend some new feature on top of that, it's a little hard. And TensorFlow RTM, I think it's blazing fast. Its performance is so good, but it's not easy to do some secondary development. If you want to add some new feature, it's a little hard. So just think about, oh, how can we create a new framework? It can achieve the good performance. Also, it's easy to develop, to maintain. So that's why they create the SGLAN project. Let's run through maybe the three main techniques behind SGLAN. So the first one is a radix attention, which focuses on KVCache. when you think about a model that is, you know, as large as DeepSeq v3, especially, like having better KVCache reutilization is great.

39:08Can you just talk a bit about that performance impact? Yeah, Redix cache, I think it's the technology of the prefix caching. And it is a special case for something like block size is one, you know, for VM or for other frameworks, they use something block size 32. And SGLAN use the block size one. I think if you use the block size one, you can make the cash hit rate higher than other frameworks. I think that's the many benefits. And for your case specifically, how does that change when you have like a base 10 type use case where you do not have a share endpoint versus like, you know, is this less helpful for GPU clouds to do one model for like many people that have like very different use cases versus like when you have just one endpoint for one customer?

39:58I'm sure to have a system prompt that a lot of models share and things like that. Anything you want to mention there? Yeah, we've seen this be massively helpful for the reason that you mentioned. There is a certain sort of finite number of prompts or at least prompt prefixes that are being used per customer. And what we've seen is that prefix caching and the fin techniques to make that better has been massively helpful. However, we still had to build on top of that. The example there is that you have a model with dozens of replicas, each of which has its own state of KVCache. A new request comes in.

40:42And what we used to do back in the day was that that request would be randomly assigned to one of these replicas. But the better way to do it is that knowing the state of KVCache in these different replicas, trying to decide which one it should go to. One of the parameters that you need to consider, there are other parameters to consider around the size of the queue at each of the replicas and the location of each replicas, depending on how geo-aware you want to be. But adding that additional consideration around KV cache-aware load balancing was something that we saw improve latency quite a bit for our customers.

41:23And then the second part, which was maybe the harder one to understand as a practitioner, which is this idea of like turning some of the decoding process into a finite state machine instead of a more open-ended. When you're using, especially for structure outputs, can you maybe explain what that means? And I would love to learn too. So maybe this is an opportunity for everybody to better understand how you think about going from a normal kind of like token by token decoding to having a more, I wouldn't say precompiled, but like pre-understanding of what the paths are going to be. I think SGLAN supports concentrated decoding, and it also supports jump forward.

42:02And we use something like outline or the X grammar to do something like change the comfort of the schema from JSON to the FSM, the state machine. And we can use the state machine to control the output, something like the output maybe to the JSON mode or something like. It should be obey some rule. So in that case, because the output should obey some rule, so you can skip some token, something like you should decode four times, but you should obey that rule. So you can get that token in advance. You can just use one preview to replace the full decoding, for example. So that's why you can jump forward.

42:48I guess the question is like, why doesn't everybody do that? When I was reading, I was like, this just sounds better, especially both for accuracy. You know, you're kind of constraining for structure output as well. You can do faster decoding. Are there downsides to it as well? I think maintain jump forward is a little hard. Yeah. At later, we support something like CPU overlap. I think in the overlap model, we even make it compatible with the jump forwarding. Because just if you want to maintain the jump forward with other features, it will be more complicated. So I think we only use the fault setting.

43:25We disable it by default. But if you want to enable it, you can just use some arguments to enable that. But it's a little hard to maintain, especially compatible with other optimization features. Just as a side note, you mentioned Xgrammar, which I never heard about. And I looked up the GitHub repo. It's actually from MLC, which we talked to TQ, I think, a while ago. Any comparisons between Xgrammar and Outlines? Is there a trend in this world or is it mostly settled science? To be honest, I prefer ex-grandma, you know. Okay, yeah, tell us. MLC AI is founded by Tianqi. Both Tianqi and his ex-grandma's other Yixing Dong, they were the students graduated from Shanghai Jiao Tong University.

44:13And the creator of SG Lang, Lian Ming and Ying, They also graduated from Shanghai's Jiao Tong University. Oh my God. Is it the Berkeley of China? Yeah, you're right. And I think Xgrammar's performance is better than the Outlines. And also in the TensorFlow RTLM, the latest release, TensorFlow RTLM also integrate Xgrammar as the backend for the constructed coding. Okay, this is new to us. We had Remy from Outlines speak at my past conference, but I wasn't even aware of Xgrammar being a thing. But yeah, I mean, structured output is something that a lot of people care about. We had OpenAI talked about their structured output implementation, and there's a lot of interest in making sure that there are no trade-offs.

45:01I think there's a little bit of FUD around how maybe the models are dumber when you use structured output instead of the sort of base, sort of next token generation, but I don't think it's significant that much. Yeah. We can talk about the last one, which I don't know if it's as relevant for Base 10, which is the third technique of SG Language is API speculative execution, which seems to be only for API-only models. Oh, yeah. I think it's the front-end feature. It's not the back-end. Yeah. Something like you have some control flow for the LM task. Yeah. Such as you have the one request to get a result and you just continue to another call.

45:41And for this case, you can use the SGLAN front-end language to describe the control flow. It will make it easier to control that pattern. Okay, awesome. Tracing this human path, I'm pretty sure I know the answer, but is there a reason big projects like Grok, like XAI, also use SGLAN? Yeah, yeah, yeah, right. Is it just the same people? Yeah, yeah, yeah, right. Right. Lian Min and Yin are the XAI member of the technical staff. I mean, it makes sense. I wonder if it, you know, what's the impetus for SG Lang to kind of break containment? It seems like VLM obviously has the, it's one year older. It has more community pool.

46:24I wonder how this will shake out. I don't really know. But, you know, you said it's a library. You said VLM's library of SG Lang is much, much more comprehensive. I mean, do people care? Maybe it's like when you're serving models at scale, then you start really prioritizing the sort of performance that SG-Lang offers. I think if you care about the performance, maybe TensorFlow RTM is the best solution for now, especially for the latency sensitive scenery, TensorFlow RTM doing well. But if you also want to implement some feature by yourself or do some optimization by yourself, you want to customize the framework.

47:06I think SGLAN is a good option. And VLM, I think VLM's community support is very nice because it was used by so many users and it has so many GitHub stars. And, you know, SGLAN, when I participated in the SGLAN team at July, it has only 2 ,000 stars. And right now it has more than seven stars. Yeah, I think it also grows so fast. Anything that people should look forward to that's on the roadmap for Xilang that people should be aware about? Yeah, yeah, yeah. And we post the roadmap in the issue and we ping that issue. Also, we have bi-weekly meeting to think with the community about our progress, our plan, which feature do we want to implement in this quarter, something like that.

47:55And we also co-host some meetups, something like the first meetup we co-host with the MLCLM and FlashInfer. And we also participated in some hackathons, something like CamoAI hackathons. We do the presentation about SG-Line. I just saw it now. It sounds like actually that there's, you know, we mentioned Eagle and mentioned Medusa. I think Amir mentioned Medusa, but Eagle is also part of that cabal of speculative coding techniques. It looks like now you support it. Yeah, we already support it. And I think in the open source implementation, something like VAM, SGLAR, and other frameworks, I think it's the SOTA performance.

48:35And currently, even use the TensRTM, it only supports Ego1, not Ego2. And one thing to note about speculative decoding, different versions of it, is that the framework supporting it is one thing, but you will have to do the job of the training of the draft models or the additional heads. And a lot of the benefits will come from how good you are at the training aspect in terms of the data that you use to train the draft model to essentially distill the target model or mimic its behavior so that you can have a very high rate of acceptance. it's the kind of throughput improvement that you get are ultimately dependent on how good of a job you do at training the draft model in the case of the draft target model mechanism.

49:31So that's another thing that's like, hey, does the framework support it? Or you can just turn on speculative decoding with a flag. That's not the case. There's more that goes to it. One more side thing on training i also noticed that with opni offering fine-tuning for o1 and all these things i think people are also very interested in sort of rl trainers is what they what you have here it looks like you're supporting hugging face tl and open rl hf do you think that this will become something that a lot of people are demanding like the sort of general feel of rl for lms relatively abandons i think up till like the end of last year basically yeah i i think so i don't know it's like it's One of those things where maybe people have to wait for a base model that has some layer looping or some other sort of friendly architecture for reasoning instead of just pure RL on LLMs.

50:26Because I think so far people have not really exploited RLHF as much in the wild. I mean, correct me if I'm wrong. Yeah, so I can give you some examples of when we've seen at work. Again, this is generally done by our customers before they come to us for inference. But there's examples like in the healthcare world, fine-tuning models for understanding medical jargon, like for Whisper, for instance, a version of Whisper that I can actually understand medical jargon. So that's not an LLM use case. In the LLM use case, staying in the healthcare space, the models that can do medical document extractions and do a very good job compared to even state of the arts because of the data that the company had gathered through human in the loop.

51:17Are those going to go away? The need for those are going to go away because there's a model that can do reasoning and do a very good job at it. I don't know. My intuition says, yes. Will it be cost effective? That's the question that I have. In the short term, no. In the long term, maybe. But I haven't seen the need for the more traditional fine-tuning actually go down. In fact, we see that quite a bit right now in the market. The question for us is, do we want to address that market knowing that the entire market might go away one day? And my general answer to that is, let's solve today's problems.

51:57Even if they're not around in two years, you will learn a lot along the way by onboarding customers that have today's problems. And you'd learn from them about tomorrow's problems and you will build ahead for them. Hang on. Why do you think fine tuning might go away? Because like you said, there's going to be models with complex reasoning capabilities that can actually figure it out in a few shot kind of way without needing a large data set to fine tune the model with. That's what some people are saying. I really have trouble believing that that'll be the case. Much more believe that it's just easier to change your prompts rather than actually do full fine tunes or even parameter efficient fine tunes.

52:45For sure. Is there anything else that we haven't touched on that you wish the people ask you more about because it is something that is very interesting from your point of view of what you're seeing among your community. Yeah. When we released the DeepSeek Feast 3 support, we have some community user, something like a cursor. Do you know cursor? I think it's very popular. Of course. Of course. I use them every day. When I type code. inside of my IDE, my terminal, it actually opens cursor instead of VS Code. I feel very bad for VS Code. And when we released the DeepSync Fee Story support, some employees from the Curso team also very interested in our implementation and reach out and ask some questions from us.

53:31So I think as SGLAN grows faster and the features optimization, we iterate so fast. And I think there will be more users from different companies, from different teams to use it. Honestly, I would go back to what I emphasized earlier, which was that I wish more people asked about what it takes to run mission-critical inference workloads. Because I see this in the market sometimes that they're like, well, I can just use VLLM and that puts my model behind an API and that is production. But really, it takes three pillars that all need to be there. One is performance at the model level. That is where frameworks that we talked about today really help you with, but you still have to guide them.

54:24Like when it comes to speculative decoding, yes, they support it, but who's going to train or fine tune the draft model or the Medusa heads, or who's going to ensure that the reliability of the VLM server that you see it in production, like there are crashes, how do you recover from those without affecting production traffic? But by itself, that's not enough because invariably that one model running on a set of hardware is going to get too much traffic that it cannot handle. And at that point, you need to horizontally scale it. And that's not an ML problem. That's not a PyTorch problem. That is an infrastructure problem to ensure that you can horizontally scale up your model extremely fast to meet your P90, P99 latency requirements.

55:15And to ensure that you're not running out of capacity in a single region that that model lives, you end up having to scale that model across different regions and across different clouds even to ensure that that model is not being starved of resources in the one place that it lives. So that's an area of investment that we started investing in some time ago and really paid off this past year. And the third pillar is enablement of workflows. Workflows such as the sort of AI phone call example that I told you about, the ones that require multi-step, multi-model inference, but in a very low latency way.

55:58That's the third pillar that really allows the developers to be able to use the power of the first two pillars and then be able to combine them when especially you need multiple models for your workflows and doing so in a reliable way, repeatable way, and a low latency way. And those are the three pillars, honestly, that we have been investing a lot on, some of which we started investing in three years ago, and it really started paying off a year ago. So it takes quite a bit of build to get it to the point where you're truly running folks, customers' mission-critical inference workloads. What do I mean by mission-critical inference workloads?

56:41Inference where if inference is slow or down, the main product of our customer is slow or down. So they really care about it. They have strict requirements around latency, around being able to support large throughput, about being able to do so in a way that, you know, other customers' usage doesn't affect the SLAs that they are getting and dealing with noisy neighbor problems. An inference done in a compliance way, whether it's HIPAA or certain SOC requirements. And also inference done in a geo-aware kind of way, both for compliance reasons and also for latency reasons, where you forward the traffic has an impact on latency in situations where 50 milliseconds really matter, 100 milliseconds really matter.

57:29And we're seeing more and more of those use cases. Well, one way I would recommend doing that is kind of a manifesto type of thing. I'm sure you know Heroku's 12-factor app. I've seen that, yes, yes. That's a good idea, actually. Yeah, maybe even put it on a separate property than Base 10 and just go like, here's what we think, you know, mission critical a application should be and you know have some thought leadership there and flesh it out and you know see if the market takes it on as a mission and obviously you will be best prepared to to serve that market as well i've also seen this done very well with enterprise ready.io i think it used to be done by i think it's called gravitational or replicated one of those from uh replicated yeah yeah these kinds of things when you have a list of requirements when like you're like look everybody needs this okay like write them up and then you know put a little bit of marketing on it split it out from the main company brand that tends to work very well yeah good idea cool well thanks thanks so much for your time i think this is a really good dive into both base 10 and sg lang and a little bit of deep stick v3 which people are very interested in i'm trying to talk to them as well because obviously they're uh they're a fascinating lab and but you know i think you guys you guys are doing a lot to make it accessible for everyone so thank you so much And as a, just to give based on some street cred, you know, they were one of the first sponsors for Latent Space events.

58:50And Amir brought a hundred croissants at our Latent Space hackathon in 2023. So yeah, I just want to bring that up. I always tell, I saw Phil and at AWS reInvent and I told them that was one of the first events that we really did. And one of the turning points of this industry, as far as community goes, in my mind, you know, everybody, everybody, everybody was there. The croissants? No, no, not the croissant. The event itself. Yeah, entire companies launched. Yeah, you were, I mean, you know, like Natter from Brev was there and did the, with Joseph from Roboflow, they did like the prompt battle thing.

59:23Like Harrison was a judge and it's like Jerry from Lomindex was there. Like kind of like everybody that is kind of now breaking out. If you look at the graph that Jensen put on the screen at CES with some of the companies they work with, a lot of them were at that event. So yeah, thanks for staying involved with us, Amir. and I'm sure we'll do more together. And thank you guys. Many more years to come, for sure. Thank you for taking the time today. Good to see you both. I'll see you, Sean.

From the publisher

Sponsorships and applications for the AI Engineer Summit in NYC are live! (Speaker CFPs have closed) If you are building AI agents or leading teams of AI Engineers, this will be the single highest-signal conference of the year for you.

Right after Christmas, the Chinese Whale Bros ended 2024 by dropping the last big model launch of the year: DeepSeek v3. Right now on LM Arena, DeepSeek v3 has a score of 1319, right under the full o1 model, Gemini 2, and 4o latest. This makes it the best open weights model in the world in January 2025.

There has been a big recent trend in Chinese labs releasing very large open weights models, with TenCent releasing Hunyuan-Large in November and Hailuo releasing MiniMax-Text this week, both over 400B in size. However these extra-large language models are very difficult to serve.

Baseten was the first of the Inference neocloud startups to get DeepSeek V3 online, because of their H200 clusters, their close collaboration with the DeepSeek team and early support of SGLang, a relatively new VLLM alternative that is also used at frontier labs like X.ai. Each H200 has 141 GB of VRAM with 4.8 TB per second of bandwidth, meaning that you can use 8 H200's in a node to inference DeepSeek v3 in FP8, taking into account KV Cache needs.

We have been close to Baseten since Sarah Guo introduced Amir Haghighat to swyx, and they supported the very first Latent Space Demo Day in San Francisco, which was effectively the trial run for swyx and Alessio to work together!

Since then, Philip Kiely also led a well attended workshop on TensorRT LLM at the 2024 World's Fair.

We worked with him to get two of their best representatives, Amir and Lead Model Performance Engineer Yineng Zhang, to discuss DeepSeek, SGLang, and everything they have learned running Mission Critical Inference workloads at scale for some of the largest AI products in the world.

The Three Pillars of Mission Critical Inference

We initially planned to focus the conversation on SGLang, but Amir and Yineng were quick to correct us that the choice of inference framework is only the simplest, first choice of 3 things you need for production inference at scale:

“I think it takes three things, and each of them individually is necessary but not sufficient:

* Performance at the model level: how fast are you running this one model running on a single GPU, let's say. The framework that you use there can, can matter. The techniques that you use there can matter. The MLA technique, for example, that Yineng mentioned, or the CUDA kernels that are being used. But there's also techniques being used at a higher level, things like speculative decoding with draft models or with Medusa heads. And these are implemented in the different frameworks, or you can even implement it yourself, but they're not necessarily tied to a single framework. But using speculative decoding gets you massive upside when it comes to being able to handle high throughput. But that's not enough. Invariably, that one model running on a single GPU, let's say, is going to get too much traffic that it cannot handle.

* Horizontal scaling at the cluster/region level: And at that point, you need to horizontally scale it. That's not an ML problem. That's not a PyTorch problem. That's an infrastructure problem. How quickly do you go from, a single replica of that model to 5, to 10, to 100. And so that's the second, that's the second pillar that is necessary for running these machine critical inference workloads.

And what does it take to do that? It takes, some people are like, Oh, You just need Kubernetes and Kubernetes has an autoscaler and that just works. That doesn't work for, for these kinds of mission critical inference workloads. And you end up catching yourself wanting to bit by bit to rebuild those infrastructure pieces from scratch. This has been our experience.

* And then going even a layer beyond that, Kubernetes runs in a single. cluster. It's a single cluster. It's a single region tied to a single region. And when it comes to inference workloads and needing GPUs more and more, you know, we're seeing this that you cannot meet the demand inside of a single region. A single cloud's a single region. In other words, a single model might want to horizontally scale up to 200 replicas, each of which is, let's say, 2H100s or 4H100s or even a full node, you run into limits of the capacity inside of that one region. And what we had to build to get around that was the ability to have a single model have replicas across different regions. So, you know, there are models on Baseten today that have 50 replicas in GCP East and, 80 replicas in AWS West and Oracle in London, etc.

* Developer experience for Compound AI Systems: The final one is wrapping the power of the first two pillars in a very good developer experience to be able to afford certain workflows like the ones that I mentioned, around multi step, multi model inference workloads, because more and more we're seeing that the market is moving towards those that the needs are generally in these sort of more complex workflows.

We think they said it very well.

Show Notes

* Amir Haghighat, Co-Founder, Baseten

* Yineng Zhang, Lead Software Engineer, Model Performance, Baseten

Full YouTube Episode

Please like and subscribe!

Timestamps

* 00:00 Introduction and Latest AI Model Launch

* 00:11 DeepSeek v3: Specifications and Achievements

* 03:10 Latent Space Podcast: Special Guests Introduction

* 04:12 DeepSeek v3: Technical Insights

* 11:14 Quantization and Model Performance

* 16:19 MOE Models: Trends and Challenges

* 18:53 Baseten's Inference Service and Pricing

* 31:13 Optimization for DeepSeek

* 31:45 Three Pillars of Mission Critical Inference Workloads

* 32:39 Scaling Beyond Single GPU

* 33:09 Challenges with Kubernetes and Infrastructure

* 33:40 Multi-Region Scaling Solutions

* 35:34 SG Lang: A New Framework

* 38:52 Key Techniques Behind SG Lang

* 48:27 Speculative Decoding and Performance

* 49:54 Future of Fine-Tuning and RLHF

* 01:00:28 Baseten's V3 and Industry Trends

Baseten’s previous TensorRT LLM workshop:



Get full access to Latent.Space at www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
Everything you need to run Mission Critical Inference (ft. DeepSeek v3 + SGLang)Latent Space: The AI Engineer Podcast · 1 h
Listen in VO